Jerry Tworek: 'The Era of Evals Is Done' – Insights on Codex, RL, and Automated AI Labs
agihouse_org · x · 2026-07-23
In an interview at AGI House, former OpenAI VP of Research Jerry Tworek discusses why he believes 'the era of evals is done.' He covers Codex, HumanEval, Copilot, RL shaping reasoning, tool use (99% systems, 1% algorithms), test-time compute scaling, and building automated AI research labs.
Related event: Ex-OpenAI VP Jerry Tworek Says Era of AI Evaluation is Over(2 posts)→
More from AGI Musings
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- 'Hallucination' Is a Category Error: Naming AI 'Intelligence' Limits Our Imagination — Genaforvena · 2026-09-11
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11