Zhiyuan's WITA-Omni Tops DailyOmni, Beating Qwen and Gemini

智东西 · wechat · 2026-07-29

Zhiyuan Robotics released the full-modal large model WITA-Omni, which topped the authoritative embodied multimodal benchmark DailyOmni with a score of 85.21, surpassing general LLMs like Qwen, Gemini, and Doubao. The model demonstrated leading performance in core metrics such as audio-visual alignment and cross-modal reasoning.

WITA-Omni adopts a native Thinker–Talker–Actor architecture, breaking the delays and fragmentation of traditional serial processing. It elevates actions and expressions to first-class outputs alongside speech, achieving synchronized "thinking, speaking, and moving" on a unified timeline. The model was trained on tens of millions of hours of multimodal data and refined via high-quality post-training strategies.

In speech interaction evaluations, WITA-Omni also outperformed models like GPT-Realtime, demonstrating excellent full-duplex conversation capabilities. It naturally handles pauses, turn-taking, and interruptions, completing Zhiyuan's "Trinity of Intelligence" in embodied AI.

Original post →

More from Embodied

Embodied channel →