Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale
Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini
cs.SE, cs.AI
2026-10-06
FlowAgent repairs Google pre-submit test failures with a fine-tuned Gemini 2.5 Pro ReAct loop. Of 195 fixes, 67.18% were correct; developers applied 28,554 of 295,508 changes.
After a Google developer creates a change in Critique, TAP runs the affected tests across the monorepo. The median gap from that failure notification to the next edit is 27.48 minutes. A quarter of developers are already editing by 6.72 minutes, and the fastest tenth by 2.37 minutes. Once the author leaves to dig through logs, the context of that change is gone.
SWE-Agent, AutoCodeRover, Passerine, and Meta's Engineering Agent repair failures after submit, offline, with a ReAct loop and no deadline tied to the next human edit. FlowAgent runs in the pre-submit outer loop, outside the IDE: CI is already red, and the patch has not reached trunk. The paper presents it as the first industrial pre-submit repair system built for that latency, accepted at ASE 2026.
A failure hits pre-execution filters before any model call. Build errors, tests that were already red in the repo, known flaky tests, and changes made by automated tools are skipped. So are changes the developer has already edited again, so a suggestion is not aimed at a stale diff. Images, binaries, and changes over 100 files or 1,000 lines are skipped too. The rules come from internal practice and early user feedback, to cut hopeless attempts and noisy suggestions.
What remains enters a sandbox. An orchestrator reads the test logs, the change description, and the diff, then starts a generate-and-validate loop. The model can read, edit, and delete files, search code, build, run tests, and read logs. The cap is 100 tool calls, to limit runtime, and 30 minutes of wall clock, set against the 27.48-minute median before a developer edits. The model is an internal Gemini 2.5 Pro fine-tuned on Google code, with temperature 0.1 so actions stay mostly deterministic, and top-p 0.95.
If the loop emits a diff, a deterministic check follows. The agent must have actually run the tests on its last cycle. Otherwise the system runs them again. A model saying the tests pass is not evidence. One more model call then reads the whole trajectory and compresses the root cause and the fix into a short UI summary, ideally one or two sentences, at most five.
Two post-execution filters run before anything is shown. If the developer edited the change while the agent was working, or already submitted it, the fix is dropped. Some of those later edits touch unrelated files, so the patch might still apply, but it can also conflict. Those suggestions are discarded outright.
The fix is attached to the existing TAP failure, not posted as another finding. A change already carries hundreds of analyses. The UI leads with the root-cause summary, then a button to preview and apply inside Critique, plus a link into Cider.
Three engineers, each with at least five years of experience, reviewed 195 real TAP failures from 37 teams. They judged whether the patch matched the intent of the change, then met to resolve disagreements. They marked 131 fixes correct, 67.18%. The other 64 were wrong, or made the tests pass without matching intent.
From October 2025 every TAP failure entered this pipeline.
| Stage | Changes | Note |
| Failing tests | 2,567,829 | all |
| Dropped before execution | 1,785,955 | no model call |
| Repair attempted | 781,874 | 30.45% kept, 36,479 authors |
| Patch produced | 421,818 | 53.95% of attempts |
| Dropped after execution | 126,310 | developer already edited or submitted |
| Shown to developers | 295,508 | 70.06% of patches |
| Previewed | 65,069 | 22.02% of those shown |
| Applied | 28,554 | 43.88% of previews |
In Critique, 55,029 were previewed and 20,826 applied (37.84%). In Cider, 21,279 were previewed and 12,702 applied (59.69%). The two surfaces overlap, so the sum exceeds 28,554.
On the 781,874 attempts, changes averaged 10.57 files (median 6) and 16.49 failing tests (median 2, so a minority of changes pull the mean up), across 917 file extensions. A run averaged 21.37 tool calls and 424,149 tokens. Total usage was 643 billion tokens. Median time from notification to a patch was 9.85 minutes, and the 90th percentile was 25.83 minutes, still under the 27.48-minute human median.
Rejected suggestions include adding @Ignore, swallowing an assertion in try/catch, asserting greater-than 2 where the developer wanted equality to 3, and reverting the developer's edit. Some correct patches were retyped by hand anyway.
Nine developers from nine groups, each at Google at least six months and each having previewed or applied a fix, sat for 45-minute in-person interviews. Mock updates and mechanical refactors were the cases they called out as worth handing off. The root-cause summary saved a pass through the logs. Others said the patch was right but arrived after they had already fixed it, so they typed their own edit.
The paper calls this the largest industrial repair dataset reported so far, and it sits on the pre-submit path. The piece that can move elsewhere is the pipeline: on a failure notification, abstain with rules first, then generate in a sandbox and require a real test run, and show the suggestion only inside the review UI people already have open.
The 30-minute wall-clock cap is there to beat a 27.48-minute median edit. Temperature 0.1 is there so actions stay stable and reproducible. A 9.85-minute median depends on Google finishing build and test inside that budget. On a slower build system the same loop would hit the 30-minute cap first.
The apply rate after preview is 43.88%, and it is higher in Cider than in Critique. Only 22.02% of shown fixes are previewed at all. The 67.18% figure is an expert judgment on 195 cases, not the production apply rate. Passerine and Meta's Engineering Agent are discussed as post-submit systems. There is no head-to-head number on the same pre-submit failures.
The three reviewers did not own the production or test code, and the paper says they may have judged wrongly. The 195 failures were a random draw, not a guaranteed cross-section. Pre-execution filters keep only 30.45% of failing changes, using rules from a small early user group. Machine-authored edits, large diffs, flaky tests, and tests that were already failing never enter that 67.18%.
An apply click is not a correctness label. Developers keep editing after they apply a patch, and they sometimes retype a fix that was already right. The paper does not report whether tests stayed green after apply, or how much of the agent's diff survived to submit. Preview and apply counts are also unadjusted for how developers already feel about AI assistance.
The interviews are nine volunteers. People who ignore AI fixes were not in the room, which the paper flags as self-selection. The model is an internal Gemini 2.5 Pro, and small prompt changes can move the results. The data cannot be released, and there is no run on an external CI. Gemini also drafted some of the paper's tables and Colab plots.
The agent still suggests reverting a change or commenting out a failing test, which the interviews treat as a trust problem. A pre-display check for those patches is still being added. Critique could not show remaining time, so developers could not decide whether to wait. Attempts that produced no patch are blamed on the time and call caps. The paper expects newer models to recover a substantial share of them, and does not report that experiment's success rate.