Same Weights, 30% to 95% on ARC-AGI: YC's Deep Dive Into Why Harnesses Are Real Research
ycombinator · x · 2026-09-07
Harnesses are often dismissed as mere scaffolding—prompt engineering, not real research. YC's Paper Club argues the opposite: the same model weights score 30% on ARC-AGI with a weak harness and 95% with a better one.
YC gathered frontier researchers and founders for a deep dive, with Seth Karten's talk pushing beyond naive prompting toward an "agentic OS"—using eval-driven harness design to unlock more of a model's underlying capabilities via persistent computation, memory, and agent-to-agent communication.
Highlights:
- François Chollet on why harnesses matter
- Accidentally building an auto-researcher
- A five-minute history of harnesses
- Self-improving harnesses
- What YC learned building an agent for every employee
More from coding & agent
- AI-generated PCBs plus desktop fabrication: from scrap board to working circuit in an hour — debreuil · 2026-09-08
- AI model Astra wins 40 of 43 AtCoder contests; Opus 5 dip hints at training contamination — teortaxesTex · 2026-09-08
- Open-source BYOK playground forces LLMs to output legal moves across five board games — wilsonye · 2026-09-08
- Astra Transcribes a Piano Part From YouTube, Then Builds a Playable App — danshipper · 2026-09-08
- Prompt Injection Attacks on AI Agents Up 340% in 2026; Fixes Must Be Architectural, Not Prompting — Thionne_WTZ · 2026-09-08
- Dev proposes open protocol for agent tool discovery at $0.05 per request — kleffew94 · 2026-09-08