Fireworks: DeepSeek-V4.1-Flash Matches DeepSWE Quality at 1/15th the Cost

lqiao · x · 2026-09-16

Fireworks AI benchmarked its new DeepSeek-V4.1-Flash, claiming a new Pareto frontier on quality vs. cost. Key insight: running DeepSWE, input tokens outnumbered output 174:1, 99.6% were cache hits, and those hits made up 60% of the bill — netting $0.43/task vs $6.52 at equal quality. Agent cost optimization hinges on input caching, not output pricing.

Original post →

More from Infra

Infra channel →