YC Harnesses Deep Dive: Same Weights Jump From 30% to 95% on ARC-AGI
ycombinator · x · 2026-09-07
YC convened frontier researchers and founders for a deep dive into agent harnesses, pushing back on the view that harnesses are 'just prompt engineering': the same model weights scoring 30% on ARC-AGI hit 95% with a better harness.
Key topics:
- History and self-improvement: a five-minute history of harnesses and self-improving designs; Seth Karten demos Prime Agent, a self-improving RLM harness.
- Context as cache: treating context as L1/L2/L3 cache, the Turing machine to von Neumann analogy, and inter-agent messaging.
- Benchmarks: ARC-AGI results under different harnesses plus Emulator Bench and GPU kernel generation.
- OpenJarvis local AI: Jon Saad Falcon's five primitives of a personal AI stack, cloud models optimizing the local stack, claimed 800x cheaper than the cloud.
- YC's internal agents (QM): building an agent for every employee, a fleet of 50 agents via OpenClaw, letting agents pick their own sandbox and model, the 'grind tool' for goal budgets, and lessons on agents missing social context.
More from coding & agent
- Design Docs Are All You Need: DeepMind's library regenerates all code from NL docs — omarsar0 · 2026-09-07
- Agents are only powerful if they connect to the systems you already use — shensi · 2026-09-07
- Teknium's agent-driven cleanup sheds 375,000 lines from Hermes Agent codebase — Teknium · 2026-09-07
- Cursor makes self-hosted cloud agents generally available for enterprise networks — thione · 2026-09-07
- Anthropic Open-Sources Claude Commerce Agents, Shopping and Merchant Agent Blueprints — thione · 2026-09-07
- Tactical programming is dead: Kent C. Dodds says fall in love with problem-solving instead — mattpocockuk · 2026-09-07