Chinese Flash Models Criticized for Over-Reasoning Latency

oran_ge · x · 2026-08-23

Developers reported that Chinese 'flash' tier models, such as DeepSeek and Kimi, suffer from severe latency issues due to 'over-reasoning.' Some complex tasks trigger up to ten minutes of processing time before output. The author speculates this may result from excessive post-training exceeding the model's parameter capacity.

Original post →

More from Models

Models channel →