Latent Space Podcast: MIT's Alex Zhang on RLMs, Agent Swarms and Harness Design
Latent Space · rss · 2026-10-02
Latent Space interviews MIT PhD student Alex Zhang, the force behind Recursive Language Models (RLMs) and GPU Mode/KernelBench.
Key points:
- GPU kernel automation: from the CUDA Mode Discord to KernelBench; AI-written kernels still leave room for human expertise — one expert insight can replace enormous brute-force token search.
- RLMs explained: context offloading, code execution, recursive subagents, and shared memory; an RLM-based harness was first to solve ARC-AGI-3, ahead of OpenAI.
- Harnesses as compositional generalizers: Claude Code, Codex, and Pi are structurally more similar than they look; the model you query in the future may secretly be an invisible swarm of agents.
- OpenAI's 10,000-agent experiment: 130B output tokens and $40M-equivalent problem solving — though much of an agent swarm may be wasted search, and convergence remains hard.
- Also covered: GEV and non-autoregressive LMs, Kimi vs OpenAI multi-agent approaches, Sakana AI's open-ended research, capability overhang, speculative programmatic tool calling, and Neuralese.
The episode also covers research taste: why PhD students should take bets that initially look trivial, weird, or pointless.
More from AGI Musings
- Stanford public classes: 2,727 volunteer teachers, a steady 1:10 ratio for six years — chrispiech · 2026-10-02
- Jacob Sansbury essay: assume a few years of normalcy left, SaaS and productivity tools are dead ideas — TAbrodi · 2026-10-02
- ICPN keynote asks whether consciousness makes a difference for artificial minds — burny_tech · 2026-10-02
- Former OpenAI research VP rejects doomsday narrative: 'bring AI benefits as fast as possible' — robleclerc · 2026-10-02
- WSJ: Prompt Language and Fawning AI-Speak Are Bleeding Into Real-World Meetings — gaganghotra_ · 2026-10-02
- Christian Szegedy revisits 2019 interview: his 'crazy' timelines for AI math and code — ChrSzegedy · 2026-10-02