OmniSeek turns Omni-LLMs into evidence-seeking agents, +15.5 points on VideoHolmes
mohitban47 · x · 2026-10-10
Adobe researchers introduce OmniSeek, an agentic framework that turns an Omni-LLM into a multi-turn reasoning agent that actively acquires evidence instead of passively processing a full audio-visual sequence.
Key components:
- Native multi-turn tool use: the model decides when to look, listen, and which temporal windows to revisit, retrieving raw audio/video clips back into context;
- OmniTraj-170K: 170K synthesized multi-hop CoT trajectories with interleaved audio-visual evidence for cold-start SFT;
- Two-stage RL with an Audio-Visual Necessity reward that discourages single-modality shortcuts.
It outperforms the Qwen3-Omni-Instruct backbone on all 10 audio-visual benchmarks, including a +15.5-point gain on VideoHolmes.
More from Multimodal
- Midjourney starts testing its MCP with a limited creative community — midjourney · 2026-10-10
- Creator renders AI music video with local video model after 48 hours on two GPUs — tetsuoai · 2026-10-10
- Telling an LLM to "believe in yourself" helps it write 3D SDF models, but not enough — keenanisalive · 2026-10-10
- This 3D dragon is 27KB of LLM-generated GLSL, not a mesh or NeRF — keenanisalive · 2026-10-10
- Nikon Rescinds Microscopic Video Contest Win Over Generative AI Use — nordicinst · 2026-10-10
- AI fake videos have hit another level of realism, researcher warns — rohanpaul_ai · 2026-10-10