RL post-training Qwen 27B as a Cypher agent lifts graph-query accuracy 6.2 points for $119
sophiamyang · x · 2026-10-11
Adithya Giridharan RL post-trained Qwen 3.8 27B on Fireworks' serverless training API to answer questions over a 650K-node Neo4j graph by writing Cypher queries.
Key facts:
- The task remodeled the BIRD codebasecommunity dataset as a property graph (649,846 nodes, 1,380,394 relationships); the agent can call a runcypher tool up to 8 times before answering.
- Training used only machine-generated questions; evaluation was on 186 unseen human-written questions from the BIRD benchmark.
- Strict accuracy rose from 0.598 to 0.660 (+6.2 points, 95% interval 3.1–9.4) against an untrained adapter, clearing a pre-registered threshold.
- Notably, the gain came entirely from greater reliability on questions it could already sometimes solve — it never solved a previously unsolved question. The strict/lenient scorer gap also shows how answer formatting affects scores.
- The whole experiment cost about $119 with no GPUs to provision.
More from coding & agent
- macOS 27 recreates Vista's permission-dialog nightmare — Okta pushes XAA protocol for AI agent auth — yenkel · 2026-10-11
- Meta & CMU's IdeaScientist uses RL agents for cross-domain research ideation, lifting novelty from 36.3% to 67.0% — ZeYanjie · 2026-10-11
- Agents on a trading MCP server backtest 6x more than humans but almost never deploy — QuanTradin · 2026-10-11
- A notes MCP server that lets Claude Code, Codex and Gemini share one notebook — lovegrover · 2026-10-11
- Google's TabFM: a zero-shot foundation model that predicts table data in one forward pass — Prompt Engineering · 2026-10-11
- O'Reilly 'Agent Memory' book enters early release with first two chapters live — danielrock · 2026-10-11