Programmer Ditches Top Models: Cost per Task Varies 20x, Speed Is the New Agentic AI Bottleneck
ievkz · reddit · 2026-09-18
A programmer found he no longer needs the smartest, priciest models. Key metric: cost per completed task — GPT-5.6 Luna (max) costs $0.18 vs $3.26 for GPT-6 Astra (max), nearly 20x apart (AA Intelligence Index: 38 vs 53). Switching to Luna's cheapest tier ended his rate-limit frustrations on a $20/month plan. The new bottleneck is speed: most models on DeepInfra/OpenRouter run 25-50 tokens/s (5-15 min per Codex task), while DeepSeek-V4.1-Flash via its own API hits 300 tokens/s with a higher intelligence index of 40. His take: the race for smarter models is ending, the agentic race is now about speed, and Chinese models currently lead — while OpenAI's upcoming data centers may push thousands of tokens/s. Context-window race is over too; its growth is what spawned coding agents.
Related event: Dev ditches top model, finds nearly 20x cost gap per task(2 posts)→
More from coding & agent
- Open-source Nautilo lets your AI agent DM coworkers and fetch feedback for you — Dan_Jeffries1 · 2026-09-20
- Anthropic's Head of Product Drops a 28-Minute Masterclass on Agents in Production — ifioknkem · 2026-09-20
- Teknium: Jev can't compact context well — Hermes summarizes 95% of it away — Teknium · 2026-09-20
- HarnessRouter: routing agent harnesses instead of models, a fresh infra idea — daniel_mac8 · 2026-09-20
- MCP tool naming: short generic verbs vs explicit prefixes for LLM tool selection — skvark · 2026-09-20
- GameToMac launched 10 days ago and already runs AoE IV, CS2 and Diablo IV on Apple Silicon — nickbaumann_ · 2026-09-20