ReFigBench paper: same model scores swing on identical tasks across Claude Code and Codex harnesses

omarsar0 · x · 2026-09-23

A new paper quantifies how much the harness shapes coding-agent results. GPT-5.5 was run inside both Claude Code and Codex on the same 1,000 tasks: with a specialized PowerPoint workflow, it improved in one harness and got worse in the other — and the harness changed scores even with identical prompts.

The paper also introduces ReFigBench, a benchmark asking coding agents to rebuild real arXiv overview figures as editable PowerPoint slides that preserve text, layout, and connections. It spans ten configurations across the GPT, Claude, MiMo, and MiniMax families, scored via artifact checks, two families of LLM judges, and blinded human comparisons.

Key finding: perception remains the main bottleneck. The specialized workflow dropped native connectors in every configuration, yet human judges still preferred its renderings in most matchups — evidence that harness design and workflow engineering materially change what you get out of the same model.

Original post →

More from coding & agent

coding & agent channel →