OpenAI alum Jerry Tworek says the future of AI research is test-time compute
agihouse_org · x · 2026-07-27
This article summarizes a technical talk by Jerry Tworek, OpenAI’s former VP of Research and a co-creator of HumanEval, now CEO of Core Automation.
Core points
- Code became the first serious LLM vertical because correctness is easy to verify: programs run, tests pass or fail.
- HumanEval mattered because it gave the field a clean, checkable progress signal.
- Jerry’s “the era of evals is done” comment is not anti-evaluation; it is a warning that static benchmarks get optimized once they become important.
Why benchmarks break down
Once a benchmark matters, labs train to it, tune for it, and saturate it. Scores then reflect benchmark optimization as much as capability.
The next frontier: automated AI labs
Jerry’s current thesis is that future progress may come from systems that:
- propose ideas,
- run experiments,
- use tools,
- and improve via real-world feedback loops.
The slides shown in the post emphasize autoresearch as scaling test-time compute, with future directions including diversity/exploration, better research-performing models, and possible test-time training.
More from Companies & People
- SpaceXAI joins NVIDIA’s Open Secure AI Alliance member list — XFreeze · 2026-07-27
- NVIDIA backs open models, says frontier AI needs both closed and open weights — Diyi_Yang · 2026-07-27
- Atlas Discovery exits stealth with a drug-response prediction platform — garrytan · 2026-07-27
- AMD and Anthropic strike a 2GW MI450 deal with up to $5B in AMD investment — thione · 2026-07-27
- Cognition acquires Poke maker in a low-nine-figure deal to boost Devin — thione · 2026-07-27
- Jensen Huang says models should not be gatekept for a chosen few — Hesamation · 2026-07-27