DAIR.AI compiles 21-paper Harness Engineering collection tracing the loop from GPT-2 to self-rewriting agent harnesses
omarsar0 · x · 2026-09-09
DAIR.AI has published a curated Harness Engineering paper collection (21 papers), compiled from the YC Paper Club "Harness Edition" session. The central thesis: a harness is everything between the model weights and the world — the loop, the context it assembles, tools and skills it can reach for, sub-agents it can spawn, and increasingly the code of the harness itself. The same weight file can score 30% or 95% on the same benchmark depending only on what surrounds it. The collection traces the arc from the 2019 bare while-not-EOS sampling loop (GPT-2), through few-shot prompting making the context window the first lever a system designer can pull (GPT-3), up to 2026 harnesses that rewrite themselves. DAIR.AI also launched a new course, Vibe Coding AI Apps with Claude Code.
More from coding & agent
- Are agent swarms more legible than single agents? Safety researchers debate — jd_pressman · 2026-09-09
- Idea: AI coding tools should pitch app ideas from your unused credits and build them — thisiskp_ · 2026-09-09
- AI Can Solve Navier-Stokes but Can't Write a Maintainable TypeScript Abstraction — burny_tech · 2026-09-09
- Open-source MCP server compresses agent context: 93.8% recall vs 51.9% baseline — Beginning_Note_5385 · 2026-09-09
- How one engineering org tripled output in 18 months by removing handoffs, not just AI — rseroter · 2026-09-09
- 14B open model on one RTX 4090 matches hosted frontier model on text-to-SQL; mnemiq open-sourced — ycombinator · 2026-09-09