Qwen 27B Coding Run Consumes 919M Input Tokens, Caching Cuts Costs
TheZachMueller · x · 2026-08-25
Developer stats for DeepSWE runs on Qwen2.5 27B show 919M input tokens and 8M output tokens across 10,580 agent turns. While the input volume looks high, the author notes that KV Cache allows reuse across turns, making it far cheaper than processing fresh tokens. In comparison, the most inefficient config (Claude Code high thinking) took nearly 5 days on an RTX 6000.
More from coding & agent
- Building a layer to freeze intent and verify authority inheritance for long-lived agents — tallmetommy · 2026-08-25
- Building great evals part 8: The discriminatory property of evals — realmadhuguru · 2026-08-25
- Run agent in a daily cron loop with goals logged to local SQLite — curious_vii · 2026-08-25
- Hark Building Models to Use Computers Better Than Humans — adcock_brett · 2026-08-25
- 12 Java Projects to Stand Out: From Payments to AI-Backends — _jaydeepkarale · 2026-08-25
- Agents Force Rethinking Systems of Record — hwchase17 · 2026-08-25