From useless to good code in 4 years: hillclimbing gave small LLMs a 100x boost

menhguin · x · 2026-09-07

menhguin argues what LLMs can do — even badly — matters: in 2022 they were cost-inefficient beyond labeling/copywriting but could write basic code; by 2026, 100x improvements from hillclimbing tasks mean small models write very good code. Quoting merve, he adds that token-maxxing tasks solvable by classic CV stacks in under 200 LoC with real-time latency is the wrong approach.

Related event: Developer charts 100x gains for small models and plunging AI video costs(2 posts)→

Original post →

More from Models

Models channel →