DeepSeek V4.1 Flash Launches on Together AI, Beating GPT-5.6 Sol at a Third of the Cost
DeepSeek V4.1 Flash is now available on Together AI, claiming agentic benchmark wins over GPT-5.6 Sol and V4 Pro at one-third the per-task cost. Built on a 552B+196B MoE with 890B KV-cache compression, it reportedly beats a 1.6T-parameter model with about a third of the parameters.
2026-09-12 ~ 2026-09-13 · 3 related posts
- DeepSeek V4.1 Flash tech report: KV cache compression lets 552B beat 1.6T — 机器之心 · 2026-09-12
- DeepSeek V4.1 Flash lands on Together AI at one-third the cost per task — togethercompute · 2026-09-13
- DeepSeek V4.1 Flash: 890 bytes/token KV cache, 552B MoE beats V4 Pro on agentic tasks — togethercompute · 2026-09-13