ReFigBench paper: same model scores swing on identical tasks across Claude Code and Codex harnesses
omarsar0 · x · 2026-09-23
A new paper quantifies how much the harness shapes coding-agent results. GPT-5.5 was run inside both Claude Code and Codex on the same 1,000 tasks: with a specialized PowerPoint workflow, it improved in one harness and got worse in the other — and the harness changed scores even with identical prompts.
The paper also introduces ReFigBench, a benchmark asking coding agents to rebuild real arXiv overview figures as editable PowerPoint slides that preserve text, layout, and connections. It spans ten configurations across the GPT, Claude, MiMo, and MiniMax families, scored via artifact checks, two families of LLM judges, and blinded human comparisons.
Key finding: perception remains the main bottleneck. The specialized workflow dropped native connectors in every configuration, yet human judges still preferred its renderings in most matchups — evidence that harness design and workflow engineering materially change what you get out of the same model.
More from coding & agent
- One vanilla JS shader, no assets: AI agent renders a golden-hour sea-cliff lighthouse — NathanWilbanks_ · 2026-09-23
- Fixing agent memory: InfoWorld argues AI memory must be inspectable to be trustworthy — rseroter · 2026-09-23
- 27k-tool x402/MPP catalog released, no longer gatekept — MurrLincoln · 2026-09-23
- ML author Burkov slams OpenCode as 'one huge bug' — burkov · 2026-09-23
- Dev joining Runway shares the homemade AI bootcamp resources he used to get up to speed — tlakomy · 2026-09-23
- Claude Opus 5.5 single-shots a riso-style train journey in ~45 minutes — mshort3 · 2026-09-23