Training inside the harness lifts Qwen3-14B from 22.2% to 54.8% on Spider 2.0-SQLite

omarsar0 · x · 2026-10-11

The HarnessSQL paper shows that training a model inside the same execution harness it uses at deployment can more than double performance: Qwen3-14B jumps from 22.2% to 54.8% on Spider 2.0-SQLite, and Qwen3-8B from 15.5% to 45.2%, with both transferring to BIRD-Interact and LiveSQLBench.

Method:

Takeaway: train/deploy harness consistency is a critical variable for agent performance.

Original post →

More from coding & agent

coding & agent channel →