Omni-Decision: evidence-ledger planning hits 81.4% on OmniGAIA at 43% of Gemini-3.1-Pro's cost
Ming Ma · hf · 2026-09-30
A new paper, Omni-Decision, targets the core bottleneck of omni-modal agents that must seek evidence across video, audio, web pages and computation: planning. Noisy multimodal observations pile up in conversation history and disrupt later decisions — controlled backend swaps confirm the diagnosis, as replacing the planner hurts far more than replacing perception.
Method: replace the growing dialogue history with an explicit evidence ledger recording what evidence is missing, what is confirmed, and where records conflict. A critic reads each noisy observation and passes only usable content to the ledger, keeping the planner on a compact context throughout. Runs record state, action and verdict per step; supervised fine-tuning and decision-level RL on these trajectories further improve the planner.
Results: SOTA 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.
More from coding & agent
- One-shots are cool, but months of AI-assisted iteration is what excites this dev — TAbrodi · 2026-09-30
- Luna finished a neural rendering client job using just 6-7% of quota — MickeySteamboat · 2026-09-30
- Adding a Skill from the Skill Hub made the same chart thesis-ready — iamfakhrealam · 2026-09-30
- Closing the laptop mid-task: a hands-on run of KooKo Agent on thesis work — iamfakhrealam · 2026-09-30
- Training an AI agent on its own explanations improves coding—no teacher, no verifier, no RL — CatAstro_Piyush · 2026-09-30
- ehartford submits With, a programming language designed for both humans and coding agents — QuixiAI · 2026-09-30