Omni-Decision: evidence-ledger planning hits 81.4% on OmniGAIA at 43% of Gemini-3.1-Pro's cost

Ming Ma · hf · 2026-09-30

A new paper, Omni-Decision, targets the core bottleneck of omni-modal agents that must seek evidence across video, audio, web pages and computation: planning. Noisy multimodal observations pile up in conversation history and disrupt later decisions — controlled backend swaps confirm the diagnosis, as replacing the planner hurts far more than replacing perception.

Method: replace the growing dialogue history with an explicit evidence ledger recording what evidence is missing, what is confirmed, and where records conflict. A critic reads each noisy observation and passes only usable content to the ledger, keeping the planner on a compact context throughout. Runs record state, action and verdict per step; supervised fine-tuning and decision-level RL on these trajectories further improve the planner.

Results: SOTA 81.4% accuracy on OmniGAIA at roughly 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.

Original post →

More from coding & agent

coding & agent channel →