RL on custom search harnesses may beat the “one big model” idea

shangbinfeng · x · 2026-08-04

The post argues that if you believe in a “one big model” future, you should try RL on a custom search harness and task: the curve climbs quickly as the harness and task are made more specific.

The implication is that agentic systems may matter more than monolithic models for real work, because the bottleneck is often the environment, feedback loop, and task design rather than raw model scale.

Original post →

More from coding & agent

coding & agent channel →