OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence
Haibo Wang · hf · 2026-10-02
OmniSeek is an agentic framework that converts an Omni-LLM from passive single-pass audio-visual processing into a multi-turn reasoning agent: it dynamically decides when to look or listen and over which temporal window, retrieving sparse but critical cross-modal evidence and appending raw segments back into context.
Key points:
- A data engine synthesizes OmniTraj-170K, multi-hop chain-of-thought trajectories with interleaved audio-visual evidence for cold-starting multi-turn tool use;
- Policy is further optimized via two-stage RL with verifiable rewards;
- An Audio-Visual Necessity objective rewards trajectories that genuinely depend on both modalities, discouraging single-modality shortcuts;
- Consistent gains across a wide range of audio-visual reasoning benchmarks.
More from coding & agent
- Open-source Harness fork moves coding agents out of the app into orca — dee_hw · 2026-10-02
- LAHacks build Residue uses acoustic analysis and AI agents to personalize your study environment — jonmarkgo · 2026-10-02
- awesome-jev indexes 700 production tools around TypeSafe AI's decision model Jev — Remarkable-Gur719 · 2026-10-02
- $10k of AI inference ports TS to C++ in days, a job for expert teams over years — kristoph · 2026-10-02
- Basis Theory launches revocable credentials letting AI agents act without touching your secrets — km · 2026-10-02
- smolvm v1.22 ships near-instant VM resume for undoing agent actions, 6.5k stars — LoganGrasby · 2026-10-02