YC Paper Club: Same Model Weights Score 30% on ARC-AGI but 95% With a Better Harness
Y Combinator · youtube · 2026-09-07
Y Combinator's Paper Club convenes researchers and founders for a deep dive on agent harnesses, arguing they are real research, not mere prompt engineering. The anchor claim: identical model weights score 30% on ARC-AGI but 95% with a better harness, so harnesses should be as expressive as possible.
Key segments:
- Self-improving harnesses: Seth Karten presents Prime Agent, a self-improving RLM harness; discussions of context as an L1/L2/L3 cache, a Turing-machine-to-von-Neumann framing, agent messaging, and results on ARC-AGI, Emulator Bench, and GPU kernels.
- Local personal AI: Jon Saad-Falcon demos OpenJarvis, a personal AI on personal devices — how far behind local models are, five primitives of a personal AI stack, letting cloud models optimize the local stack, and a claimed 800x cost reduction versus the cloud.
- YC's work agent: Josh France and Regan Bell present QM, YC's internal agent harness for every employee — history of YC's internal agents, OpenClaw and a fleet of 50 agents, pulling the brain out of the sandbox, letting agents choose their own sandbox and model, the "grind tool" for goal budgets, and why agents don't understand social context.
More from AGI Musings
- New AGI scale: days since we found a task humans can do but AI can't — currently 0 — mariofilhoml · 2026-09-08
- Population Ethics Consistency Test, built with Claude, forces you to bite bullets — lxrjl · 2026-09-08
- Jensen Huang says AGI is already here — sourdub · 2026-09-08
- Gary Marcus amplifies a fiery Terence Tao take on AI — GaryMarcus · 2026-09-08
- Garrison Lovely's 'Obsolete' Book on AI's Trillion-Dollar Labor-Replacement Race Due Sept 2026 — GarrisonLovely · 2026-09-08
- Math community clashes over whether AI-generated results count — RexDouglass · 2026-09-08