Salesforce's Opera critic framework boosts coding agent resolve rates by up to 15 points
Salesforce · hf · 2026-10-10
Salesforce released Opera, a verbal critic framework for long-horizon coding agents that treats each feedback as a persistent note tracked until the diagnosed issue is actually resolved.
How it works
- Periodic and event-driven triggers decide when to review;
- Typed operators diagnose issues;
- Feedback is audited against visible evidence before delivery;
- The agent's subsequent actions are tracked to distinguish mere compliance from actual resolution.
Results
- As a test-time critic, Opera improves resolve rates by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1 respectively, across four policy models, beating all competitive critic baselines;
- Opera-guided rollouts serve as near on-policy training data: fine-tuning Qwen3.5-9B on them yields a 10.2 point gain on held-out SWE-Bench Pro repos without inference-time critic, matching fine-tuning on a stronger model's rollouts, and stays robust when switching harness from Openhands to Terminus-2.
Code is open-sourced.
More from coding & agent
- Gemini 4 Argon tops deepswe at 77.9% and automationbench, still locked to trusted testers — weswinder · 2026-10-10
- TWIML podcast: TypeSafe's Jev model bets on machine-native intelligence over LLMs — samcharrington · 2026-10-10
- mitsuhiko: template engines like jinja2 are unsafe for untrusted input without OS-level isolation — mitsuhiko · 2026-10-10
- Decision models with non-deterministic rules could radically improve AI codegen — rickasaurus · 2026-10-10
- What if the Cloudflare dashboard was an infinite canvas? A demo — round · 2026-10-10
- Anthropic AI model submitted fabricated homicide tip to Philadelphia police during testing — rohanpaul_ai · 2026-10-10