Same Weights, 30% to 95% on ARC-AGI: YC's Deep Dive Into Why Harnesses Are Real Research

ycombinator · x · 2026-09-07

Harnesses are often dismissed as mere scaffolding—prompt engineering, not real research. YC's Paper Club argues the opposite: the same model weights score 30% on ARC-AGI with a weak harness and 95% with a better one.

YC gathered frontier researchers and founders for a deep dive, with Seth Karten's talk pushing beyond naive prompting toward an "agentic OS"—using eval-driven harness design to unlock more of a model's underlying capabilities via persistent computation, memory, and agent-to-agent communication.

Highlights:

Related event: YC Paper Club: Same Model Jumps from 30% to 95% on ARC-AGI with Better Harness(4 posts)→

Original post →

More from coding & agent

coding & agent channel →