6,840 CABRA tasks stay near-perfect for agents until code merging tests understanding

Code Understanding is a Bottleneck for Coding Agents

Nishant Balepur, Kiran Tomlinson, Tobias Schnabel

cs.SE, cs.AI, cs.PL

2026-10-07

CABRA builds 6,840 call-graph tasks. Eight bare LLMs fade as size grows; six agents stay near-perfect with grep and scripts, then lose accuracy when merging divergent class logic.

What problem this solves

Repository benchmarks such as SWE-bench turn GitHub issues into edit-and-test tasks, and difficulty is usually read off the size of the diff. Line count is a poor ruler. Codebase, task type, and edit site all move together, so a long patch and a missed call chain are tangled in the same failure.

Deng et al. link low accuracy on SWE-Bench Pro to larger edits. The traces have another face: accuracy is also low when agents read more, analyze dependencies more, and test more. CABRA pulls those apart by generating programs from scratch and stretching one axis at a time.

Method

CABRA, the Coding Ability Blueprint for Rigorous Agent evaluation, sets the task type first and the size schedule second. Issues are not mined from a live repository.

A program is a directed acyclic call graph. Nodes are functions, edges are calls, and bodies are random arithmetic, with string and array variants so the result is not an artifact of floats. The graph splits into Cinit, left untouched, and Cedit, which must change. Edges between the blocks are drawn with probability 0.3, except when the task is dead-code removal. Five tasks follow Fowler's refactoring catalogue: delete functions main can never reach; thread a new parameter down the call chain; thread a new return value up from a sink; lift a repeated slow call to the lowest common ancestor; or, under DRY, move a repeated get to the first common descendant.

The size n is 5, 10, 25, 50, 100, or 200, and only one axis moves. Function traversal sets edited functions to n and holds the untouched set at 25. Function search sets distractors to n and holds the edited set at 25, a needle-in-a-haystack laid out as code. Runtime resolution wraps identical blocks in if check() and allows an edit only on the branch that actually runs. check is a Collatz-style loop that takes n steps to resolve. Instruction following appends n distinct rules, such as asserts, exact comments, and temporary variables. Dead-code removal cannot use that axis, because deletion is the only legal edit. The last two axes fix both function sets at 10. Where a reasoning effort control exists, it is set to medium.

The grid is 6,840 tasks: three code types, five tasks (four on the instruction axis), four axes, six sizes, 20 instances each. Scoring is black-box. Every function is compared with the reference on 50 random inputs, at least 500 tests in all. Dead functions must be gone, the live branch must be the one edited, and every extra rule must be met. The score is 1 only when all of that passes.

Eight tool-free LLMs write a thought block and then rewrite the whole file: GPT-5.4 Mini, GPT-5.3 Codex, GPT-5, GPT-5.5, Grok-4.1 Fast, DeepSeek-V3.2, GPT-OSS 120B, and Mistral Large 3. Six agents share GitHub Copilot's public harness 1.0.76: GPT-5.5, GPT-5.4 Mini, Gemini-3.5 Flash, Gemini-3.1 Pro, Sonnet-4.6, and Opus-4.7. Each turn is a ReAct step, one bash call, then the tool output. The paper puts the bill near $25k in the main text and at $28,262 in the appendix total.

Results

Figure 2 plots mean accuracy against n. All eight bare LLMs fall as n grows. On traversal, search, and runtime the six agents stay near perfect more often than not. GPT-5.4 Mini falls with no tools, then beats tool-free GPT-5.5 and Grok-4.1 Fast on non-instruction tasks once n passes 100. Instruction following is the only base axis where agents clearly drop: at n = 200 every agent except Opus-4.7 declines.

Agents hand the scale to tools. Traversal moves parameters with an AST script, search greps for names, and runtime execs check. The script need not grow with n. Prose reasoning from the bare model does.

Calls split into read 22%, analyze 14%, search 11%, edit 36%, test 8%, and other 9%. Qwen3-4B labels them. Twenty human labels per class agree on 94% of CABRA traces and 89% of SWE-bench traces. In Figure 3, read and analyze calls and tokens rise with n on three of the four axes. Edit calls rise with n only under instruction following. Runtime token use stays roughly flat. 48% of reads and 40% of analyses sit in the first quarter of a trace. 69% of reads and 56% of analyses occur before the first edit. A search is followed by an edit 12% of the time and by a read 19% of the time.

The same labels on mini-swe-agent traces for SWE-bench Verified, from Claude-4.5 Opus and Sonnet-4.5, Gemini-3 Pro and Flash, GPT-5.2, and GPT-5 Mini, show why the synthetic grid exists. Human time buckets run from L1 (under 15 minutes) to L4 (at least 4 hours). Lines edited, reading, analysis, and testing rise together. Point-biserial correlations with accuracy are all weak. A negative r means a larger trace signal tracks failure. Understanding calls, read plus analyze, still rank first.

Metricr with accuracyReference
Understanding calls-0.200read + analyze
All tool tokens-0.178
Edit tokens-0.177
Read calls-0.171
Lines edited-0.159
Edit calls-0.151weakest of the eight

Section 5 adds an equivalence merge. Two classes have ℓ lines each and differ in behavior on d lines. Shared logic moves into a parent, each child may keep at most d differing lines, and behavior still has to pass the same black-box tests. Renames, reordered commutative lines, split or merged statements, and distributive rewrites make equivalent code look different. ℓ runs from 5 to 200, with 50 tasks at each size. Edits grow roughly in line with ℓ, while the line combinations an agent might have to compare grow quadratically in the worst case. Agent accuracy falls as ℓ grows, and understanding tools dwarf edits. The text does not give per-point percentages.

Behavior is preserved in over 90% of cases. Most misses leave too many lines in the child: a one-line difference was never located. Accuracy falls as d moves through 1, 3, and 5. At ℓ = 50, multi-line decomposition and distributive multiplication hurt, Gemini especially, and three of the six agents struggle with multi-line gaps. Stripping out variable renaming helps the most, which points to alignment by surface names. Many bare LLMs in the appendix never clear 0.5. Opus-4.7 nearly saturates at 50 lines and still struggles with multi-line gaps and renames at 200. Telling the agent up front that exactly one line differs lifts accuracy.

Why it matters

Lines edited is a weak primary knob for coding-agent difficulty. CABRA separates tasks tools make easy from tasks that still fail once the understanding load grows. Grep flattens search. Instruction count and cross-class merging do not.

A smaller model that composes tools can beat a larger bare model on these controlled items. The calls worth spending are the ones that read the code about to change.

Synthetic items do not replace SWE-bench. The controlled axis attributes the error. The repository issue checks whether that attribution survives real code.

Limitations

The programs are random arithmetic, strings, and arrays. One call-graph family, five refactorings, and a single harness are thin ground for other languages and tool stacks.

On SWE-bench Verified every correlation is small. Understanding calls at -0.200 versus lines edited at -0.159 is a ranking, and a weak predictor either way. Human difficulty levels move reading, analysis, edits, and tests together, so the repository data still does not identify a cause. Token totals are a rough proxy, because a later call inherits reasoning from earlier ones.

Tools can skip the ability an axis was built to measure. Once grep solves function search, that axis mostly measures whether the agent writes a script. The labeler is Qwen3-4B, and agreement falls from 94% on synthetic traces to 89% on SWE-bench traces.

11% of traces touched a shared directory, and under 0.2% looked at another session. Sandboxed reruns kept the same trends. The headline numbers still come from the original batch.

Terms

Source

What people are saying

Related papers

All paper explainers