OpenAI alum Jerry Tworek says the future of AI research is test-time compute

agihouse_org · x · 2026-07-27

This article summarizes a technical talk by Jerry Tworek, OpenAI’s former VP of Research and a co-creator of HumanEval, now CEO of Core Automation.

Core points

Why benchmarks break down

Once a benchmark matters, labs train to it, tune for it, and saturate it. Scores then reflect benchmark optimization as much as capability.

The next frontier: automated AI labs

Jerry’s current thesis is that future progress may come from systems that:

The slides shown in the post emphasize autoresearch as scaling test-time compute, with future directions including diversity/exploration, better research-performing models, and possible test-time training.

Original post →

More from Companies & People

Companies & People channel →