I tested 20+ ways to make a cheap coding model act like an expensive one
KangarooAnxious9394 · reddit · 2026-10-06
The author ran pre-registered experiments (protocols committed to git before runs) on real repo commits, using Haiku as the cheap agent and Sonnet as the strong model, testing 20+ methods to make a cheap model perform like an expensive one.
What worked:
- A stronger model that only speaks up when the agent repeats mistakes: +7 successes in 63 at 1.3x cost (always-on advisor got +8 but cost 3.5x).
- Running the agent's change and reporting facts beat giving advice: 35/42 vs 32/42, formatting regressions dropped 10 → 0.
- Capturing user preferences in their own words and carrying them forward raised compliance from 40% to 90% (15/15 vs 0/15).
What didn't: memory of code knowledge, generic checklists, rules learned from git history, model routing, clarifying questions; none raised success for the strong model (45/45 with or without).
Cheapest per solved task: Haiku + "conscience" $1.22 vs Sonnet alone $1.41. Paper, protocols, failures and tool are public (source-available, non-commercial licence; works with Claude Code, Codex and OMP).
More from coding & agent
- Matt Shumer builds Hogwarts live on Spawn with AI in a multiplayer world — mattshumer_ · 2026-10-06
- Ruff author shares build-speed trick: one crate per file for faster compiles — charliermarsh · 2026-10-06
- Notable MCP critic converts, now wants an official Gmail MCP — zeeg · 2026-10-06
- Developer Runs 1,200+ Blender Modeling Sessions to Blind-Rank LLMs and Agent Harnesses — Izolight · 2026-10-06
- Software-in-the-loop: self-supervised scaling of terminal environments lifts Terminal-Bench 2 to 53.56% — Zhongzhi Li · 2026-10-06
- Real product photos plus LTX-2.5 rotation videos and Codex-coded UI rebuild a brand demo — OpenAIDevs · 2026-10-06