OpenAI 'exploration maxxing' vs Anthropic 'exploitation maxxing': a pass@k theory

scaling01 · x · 2026-08-21

The author extends his thesis that OpenAI is exploration-maxxing while Anthropic is exploitation-maxxing, with observations:

The quoted thread conjectures: as RL compute tends to infinity, a reasoning model's pass@1 should approach the base model's pass@inf (lifting the whole curve); the pass@inf gap under fixed compute budgets comes from RL environment quality.

Related event: Debate Over pass@k: OpenAI's Exploration vs Anthropic's Exploitation(2 posts)→

Original post →

More from Models

Models channel →