From useless to good code in 4 years: hillclimbing gave small LLMs a 100x boost
menhguin · x · 2026-09-07
menhguin argues what LLMs can do — even badly — matters: in 2022 they were cost-inefficient beyond labeling/copywriting but could write basic code; by 2026, 100x improvements from hillclimbing tasks mean small models write very good code. Quoting merve, he adds that token-maxxing tasks solvable by classic CV stacks in under 200 LoC with real-time latency is the wrong approach.
Related event: Developer charts 100x gains for small models and plunging AI video costs(2 posts)→
More from Models
- User notes Astra follows strict skill rules less reliably than Sol — petergyang · 2026-09-07
- Astra still generates fake unrelated criticisms when fact-checking, user finds — AndyMasley · 2026-09-07
- MLP paper shows neurons turn monosemantic in clustered regression, challenging global subspace view — burkov · 2026-09-07
- Early user: OpenAI's Astra is faster, more token-efficient and higher quality on hard tasks — Yamapama · 2026-09-07
- Gemini 4 Deep Think checkpoint tested: 85k-token SVG generations — Ryoiki-Tokuiten · 2026-09-07
- Gemini's agentic video understanding cuts tokens by 88% and costs by 66% — patloeber · 2026-09-07