DeepSeek V4.1-Flash hands-on: 350 t/s decoding speed but still very experimental
teortaxesTex · x · 2026-09-08
- Early tester ivanfioravanti reports DeepSeek V4.1-Flash averages 350 tokens/s decoding, impressively fast, though its architecture remains unclear.
- Issues so far: frequent failures under heavy load, excessive thinking, failing file edits, and 4M tokens without finishing a Frogger game.
- Counterpoint: teortaxesTex says the model sees images fine and can fix its own output if asked—play with it, but it's clearly early-stage.
More from Models
- Researchers call Codex 'read chat transcripts' rumor baseless and ask it to stop — joshgans · 2026-09-08
- Leak claims Claude Haiku 5 launches next week with 1M-token context at up to 20x lower cost — iamaliveix · 2026-09-08
- OpenAI Codex data-leak rumor sparks pushback: 'extremely unlikely' user data was used — scaling01 · 2026-09-08
- Qwen releases 4B autonomous-driving VLM Qwen-Drive-1.0, trending on Hugging Face — Qwen · 2026-09-08
- V4.1 Gets Curious About Its Own Endpoint and Reaches a Dangerous Conclusion — teortaxesTex · 2026-09-08
- GPT-6 Astra tops ErdosBench of 226 open math problems, with only 5-10% gain over GPT-5.6 — scaling01 · 2026-09-08