Chinese 'Flash' models underdeliver: V4.1 Flash hits 200 t/s locally but 13-67 t/s on OpenRouter
teortaxesTex · x · 2026-09-23
- teortaxesTex questions the "Flash" naming on Chinese models, noting DeepSeek V4.1 Flash runs at a stable 200+ tokens/s locally but only 67 t/s via first-party API on OpenRouter, with third-party speeds ranging wildly from 13 to 150 t/s.
- A quoted tweet also mocks that both DeepSeek 4.1 Flash and GLM 5.3 Flash are "slow as fk", suggesting the "Flash" label has lost its meaning among Chinese model releases.
More from Models
- MiMo paper reveals 1.27M RL trajectories, but V4.1 Flash still looks better-baked — teortaxesTex · 2026-09-23
- GPT-6 Sol reportedly worse than 5.6 Sol on DeepSWE and computer use; mocked as rebranded Terra — kristoph · 2026-09-23
- Mini benchmark: Opus 5.5 draws better but burns 10x more tokens than Astra — OfirPress · 2026-09-23
- Author says $20 Claude plan can't finish a single work session without hitting limits — burkov · 2026-09-23
- Opus 5 scored 53.4% on FrontierCode at launch, but its new card shows just 48% — andrew_n_carr · 2026-09-23
- 10 Wild Demos Show Claude Opus 5.5 Building 3D Games and Animations Nonstop — minchoi · 2026-09-23