Core Automation founders say transformer limits, not scale, are now the bottleneck
Training Data (Sequoia) · rss · 2026-07-29
A Sequoia interview with Jerry Tworek and Rohan Anil argues that the next bottleneck for smarter systems may be architecture, not scale.
The central thesis
- Jerry Tworek led reasoning at OpenAI and came to believe scaling reinforcement learning was a path toward AGI.
- Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer.
- At Core Automation, they now argue transformers have taken models as far as they can go.
Why they think the current stack is stuck
- The missing capability is continual learning: models that adapt at test time.
- In-context learning taps out quickly; the example given is Codex needing compaction after about 20 minutes.
- Fine-tuning risks catastrophic forgetting.
- They argue pre-training and RL should be optimized end-to-end, and that transformers waste computation inefficiently.
The company angle
They also explain why frontier labs are unlikely to chase alternatives while locked in the coding-agent race, and say their own goal is to build the world’s most automated lab. One concrete step is automating kernel generation, which they describe as a place where frontier models still lose to highly skilled humans.
More from AGI Musings
- AI/ML should be treated as a discipline for studying change, not constants — burny_tech · 2026-07-29
- A repost asks how “pacing” frontier AI would work in practice — TheTuringPost · 2026-07-29
- AI labs are publicly backing both open weights and tighter frontier oversight — Full_Tangelo_7450 · 2026-07-29
- OpenAI study cited in podcast says 20% of ChatGPT chats are education-related — 3scorciav · 2026-07-29
- AI won’t kill us, but systems that replace judgment may make us worse at thinking — AryHHAry · 2026-07-29
- AI slowdown looks unlikely even as OpenAI and Anthropic staff back pacing — iruletheworldmo · 2026-07-29