DeepSeek-V4.1-Flash lands on Fireworks: 552B MoE for coding and agents at 1/40th claimed cost
lqiao · x · 2026-09-11
- Fireworks has onboarded DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with a 552B-parameter backbone, native image+text input, and up to 1M token context.
- Its Causal Encoder-Decoder architecture activates only 8B parameters at prefill and 16B at decode, cutting KV cache footprint to roughly a quarter of DeepSeek-V4-Flash — pitched as a workhorse for coding, cybersecurity, and agentic workloads.
- Fireworks claims it outperforms Opus 5 and GPT 5.6 Sol on DeepSWE, CyberGym, and Automation Bench at 1/40th the cost (vendor figures). Serverless pricing: $0.22 / $0.007 cached / $0.66 per 1M tokens (input/cached/output); on-demand dedicated deployment also available. Metadata shows function calling and fine-tuning not yet supported.
More from Infra
- L3Harris says fine-tuned open-source models beat frontier AI in 48 hours at 95% lower cost — eliano · 2026-09-11
- Qwen3-TTS 1.7B hits 1.6x real-time voice cloning on CPU via llama.cpp — alexcovo_eth · 2026-09-11
- NVIDIA ships NVFP4-quantized Qwen3.8-27B, trending on Hugging Face — nvidia · 2026-09-11
- Qwen 3.8 125B-A6B runs 15.3% faster on Mac via speculative decoding on mlx.fast — TheMoonMidas · 2026-09-11
- Olam Labs CEO: only compute and data remain as bottlenecks to AGI — garrytan · 2026-09-11
- Open-sourced inference acceleration for structure-based models ships benchmarked and documented — AllThingsApx · 2026-09-11