Text-to-SQL agent pitfalls: separate critic raises cost 53% with zero accuracy gain
renatyv · reddit · 2026-10-09
The author shares hands-on benchmark results (498-task suite, pi AI agent + OpenRouter + GPT-5.6 Luna) on building a text-to-SQL agent.
What didn't work:
- Forcing LIMIT 5/100 on exploratory queries — limits leaked into final answers
- Capping the agent at 15 queries — not enough to explore the DB, let alone write good SQL
- A separate critic-reviewer in a fresh context: 307/498 vs 306 with inline self-review, no measurable accuracy gain, but +53% cost and +58% latency
What worked:
- Batching multiple queries per tool call — saves tokens, calls, and latency
- Low effort mode: only 1.2% fewer correct answers at 18% lower cost and 23% faster
- Feeding database summaries from db-snooper to cut exploratory queries
- reasoning=off + database summaries: same quality, significantly faster time-to-answer
Two blog posts detail the harness pitfalls and the reasoning/critic/schema-link decisions.
More from coding & agent
- Procedural infinite city written 100% by Opus 5.5 runs in-browser via WebGPU — jason_mayes · 2026-10-09
- Voyager launches as an open harness for AI-driven video, graphics and games — jgooten · 2026-10-09
- TermGrade: 1k open-source executable terminal-agent RL environments with full training recipe — maximelabonne · 2026-10-09
- Step 5 Preview free in Cline for a week, beats Kimi K3 and GLM-5.3 on DeepSWE — StepFun_ai · 2026-10-09
- Open-source idea-to-launch prompt turns vague ideas into full go-to-market plans — Andrew0_0 · 2026-10-09
- Thrixel launches AI 3D creation engine that lets coding agents build editable, interactive 3D worlds — RanaHanocka · 2026-10-09