Training inside the harness lifts Qwen3-14B from 22.2% to 54.8% on Spider 2.0-SQLite
omarsar0 · x · 2026-10-11
The HarnessSQL paper shows that training a model inside the same execution harness it uses at deployment can more than double performance: Qwen3-14B jumps from 22.2% to 54.8% on Spider 2.0-SQLite, and Qwen3-8B from 15.5% to 45.2%, with both transferring to BIRD-Interact and LiveSQLBench.
Method:
- Traditional text-to-SQL training produces one static query, while deployed database agents inspect schemas, run probe queries, and revise—harness behavior that only appears at inference time
- HarnessSQL builds isolated executable databases with hidden answer checks; teachers run inside the target harness, only verified trajectories are kept for SFT, followed by execution-reward RL
Takeaway: train/deploy harness consistency is a critical variable for agent performance.
More from coding & agent
- Free Photoshop clone surfaces online with control ports built for AI agents — aakashgupta · 2026-10-11
- Use Claude and Codex to build a 'design lab' that centralizes design iteration across projects — RileyRalmuto · 2026-10-11
- gemini-cli PR fixes ACP session/load responding before history replay completes — Nodge · 2026-10-11
- Three agents burned $47.35 overnight: uncapped loop instructions are a billing trap — aakashgupta · 2026-10-11
- Stop Calling Yourself 'Just a Vibe Coder': Learn Enough That Agents Can't Fool You — brandon_galang · 2026-10-11
- Agents Don't Read Your Policy Docs: Gravitee demos gateway-level AI governance — AI Engineer · 2026-10-11