Burning 200M tokens for 30M useful ones: what maxing out models teaches you
gaganghotra_ · x · 2026-09-26
After maxing out 89M tokens in a day and nearly 100M the next, the author concludes Grok 4.5 is actually pretty good and that you only truly understand a model by maxing it out rather than reading benchmarks — since you inevitably burn 200M tokens to extract 30M genuinely useful code.
More from coding & agent
- ChatGPT5.5 with a tool harness cracks Sokoban: 68 tool calls in 3 minutes — tak3sh8 · 2026-09-26
- Open-source APK Reverse Skill Hits 1.4k Stars, Lets Claude Code Analyze Android Apps — lxfater · 2026-09-26
- ThoughtDAG adds Jev-based context scoring: 391ms vs 24.8s for judging recalled excerpts — Lopsided_Scarcity979 · 2026-09-26
- Theo slams OpenRouter for using Jev: a non-reasoning classifier that can't gauge task complexity — intellectronica · 2026-09-26
- How Do You Optimize Your MCP Toolset So Agents Actually Use It Well? — onehundredemoji69 · 2026-09-26
- I Was Going to Ship 155 MCP Tools. The Token Math Said 10. — Difficult_Coffee_713 · 2026-09-26