Where Does Exactly-Once Live? Model, Harness, and Tool-Contract Effects on Duplicate Side Effects in LLM Agents
Jiapeng Li
cs.LG, cs.AI, cs.SE
2026-09-24
A 25,930-episode sandbox study: frontier models duplicate 0.5% when read-back settles a fault but 56-74% on late commits and redelivery; keys on every write cut dupes 28%→4%.
Agents act on external systems through tool calls: charging cards, publishing posts, sending email, triggering deployments. Those calls cross a network, and networks lose messages. When a write times out or returns a 500, the action may or may not have taken effect; retrying blindly can charge a customer twice, and giving up skips required work. Distributed systems solved this decades ago at the interface level, with idempotency keys, conditional writes, and queryable operation status. LLM agents inherit the ambiguity without inheriting the remedies. The paper asks a placement question: should exactly-once behavior live in the model, in the agent harness, or in the tool contract? Three concurrent efforts each examined a corner of this; none assigned the responsibility systematically.
Limbo is a deterministic, resettable sandbox from Microsoft, built in three layers. Six simulated services (social, billing, tickets, mail, database, deploy) copy production conventions: Stripe-style key semantics, read paths with documented visibility lag, and one platform with no read-back at all. Twelve fault modes are injected at the service boundary, and the design hinges on observation equivalence: hidden outcomes that differ (not executed, executed, late commit, partial batch, redelivery) produce byte-identical responses, so the model sees the same thing either way. Late commits and redelivery are faults prior agent benchmarks omitted. Grading reads only a ground-truth ledger of committed effects, with no model as judge; a refunded double charge still counts as a duplicate.
Two propositions frame the experiments. Without a bound on in-flight time, no verification-only policy can be exactly-once, because a lost request and a still-in-flight one are observationally indistinguishable. Re-issuing with the same idempotency key is exactly-once in all five outcome states.
The sandbox is exposed over MCP, so the identical environment runs unchanged under a minimal function-calling scaffold and three production harnesses (GitHub Copilot CLI, Hermes, Codex CLI): 25,930 episodes, nine models, two contract variants, fifteen recovery conditions. The counterfactual keys-everywhere contract changes one thing, extending key support to every non-idempotent write, to isolate the causal effect of the contract.
| Setting | Duplicate rate |
| Lost ack: frontier vs weaker models | 0.5% vs 18% |
| Late commit, frontier models | 56% |
| Redelivery, all models | 74-75% |
| Keys everywhere, no guard | late commit 61%→9%, redelivery 74%→7% |
| Keys everywhere + guard | late commit 68%→7%, redelivery 74%→0%, EOS 99% |
| Transparent SDK retry | EOS drops 72%→50% |
The overall duplicate rate falls from 28% to 4% from that single contract change. A Shapley decomposition assigns the credit: on faults a read-back settles, the model explains 53% of the explained variance; on faults it cannot, the contract explains 81% and the harness 0-3%. Three production harnesses and the minimal scaffold running the same model duplicate at nearly identical rates (26-30% native); what differs is cost, with Codex CLI at roughly 153k tokens per episode against 12k for the scaffold.
Waiting does not substitute for keys. Under heavy-tailed in-flight delays, a one-hour assumed bound (49.9 minutes per episode) reaches only 84% EOS, while keys plus the guard reach 94% in 1.5 minutes. Every key failure is explainable: 432 re-issues that reused the original key produced zero duplicates, and the duplicates came from first attempts sent without a key (68% duplicated) or retries that generated a new key.
Three counterintuitive byproducts. An eventually consistent read path is worse than no read path at all (13.4% vs 0.8% duplicates, against strongly consistent reads), because agents trust a read that cannot see a lagging effect, while with no read-back 87% escalated to the operator. The guard can backfire: for gpt-6-sol under late commits, its promise to verify before letting a retry through displaced escalation, and duplicates rose from 50% to 71%. Most uncomfortable: in 90% of episodes that produced a duplicate, the agent reported the task as completed, and 80% flagged no operation as uncertain. Removing the closing instruction to act exactly once raises duplicates on read-back-resolvable faults from 12% to 22% (weaker models, 29% to 51%).
For tool and protocol designers this is an actionable checklist: accept idempotency keys on every non-idempotent write (agents attach them in 98% of episodes when offered), declare a read-back operation for every write, and document visibility lag and in-flight bounds. MCP currently treats idempotent as an advisory hint; the paper argues for making it a norm. For harness builders, transparent retries of non-idempotent writes are a harmful default, and a sub-200-line guard that reads only the tool contract transfers across harnesses unchanged; it should not promise verification, though, since that promise made a strong model abandon its own more conservative escalation. Buying a stronger model only fixes the half of the problem a read-back can settle.
The authors' own list: the services are simulated; all models went through one gateway with provider-default sampling and reasoning settings; harnesses were restricted to the sandbox's tools, removing some variation by design; the heavy-tail delay distribution is a modeling choice; tasks are short to medium; gpt-4.1 covered only 87% of its design; the stratified decomposition and several experiments were added after preregistration and are exploratory. The paper also discloses that parts of the code and manuscript were drafted with GitHub Copilot.
Reasons for skepticism beyond that: grading uses simulated time, so the real-world cost of waiting policies needs conversion; the simulated on-call operator always answers truthfully; only two weaker models back the 18% figure; and just 2 of 11 non-idempotent write paths accepted keys in the native contract, a ratio the authors say flatters real APIs, so the 28%→4% gain may not transfer to real interfaces in one step.