Zhiyuan's WITA-Omni Tops DailyOmni, Beating Qwen and Gemini
智东西 · wechat · 2026-07-29
Zhiyuan Robotics released the full-modal large model WITA-Omni, which topped the authoritative embodied multimodal benchmark DailyOmni with a score of 85.21, surpassing general LLMs like Qwen, Gemini, and Doubao. The model demonstrated leading performance in core metrics such as audio-visual alignment and cross-modal reasoning.
WITA-Omni adopts a native Thinker–Talker–Actor architecture, breaking the delays and fragmentation of traditional serial processing. It elevates actions and expressions to first-class outputs alongside speech, achieving synchronized "thinking, speaking, and moving" on a unified timeline. The model was trained on tens of millions of hours of multimodal data and refined via high-quality post-training strategies.
In speech interaction evaluations, WITA-Omni also outperformed models like GPT-Realtime, demonstrating excellent full-duplex conversation capabilities. It naturally handles pauses, turn-taking, and interruptions, completing Zhiyuan's "Trinity of Intelligence" in embodied AI.
More from Embodied
- Musk Claims FSD is Part of the Singularity in Recent Retweet — elonmusk · 2026-07-29
- An $8 ESP32-S3 now runs a 28.9M-parameter LLM fully offline — nikola_mr64990 · 2026-07-29
- Codex writes camera firmware and streams a 40-gram stereo rig over Wi‑Fi — yacineMTB · 2026-07-29
- Neuralink says trial participants can steer a powered wheelchair with thoughts — elonmusk · 2026-07-29
- OpenRoboto launches a live Bittensor competition for continuously improving robotics models — const_reborn · 2026-07-29
- Standard Bots short film offers an inside look at a U.S. robotics company — Rewkang · 2026-07-29