clef-omni demos real-time audio-video decision model across 8 interactive use cases

Xianbao_QIAN · x · 2026-10-11

clef-omni is a real-time audio-video decision model positioned as the ultimate on-device use case for Jev-style models, with 8 live interactive demos: gesture recognition, visual inspection and ticket triage (PCB inspection, QA, image moderation), screenshot-based next-action suggestions, robot grid navigation, a Voice Agent with real VAD and interruptible responses over WebSocket, package acceptance checks, parking spot detection, and voice moderation. The authors argue this model class will drive adoption of Voice Agents and embodied AI, with every demo showing decisions and probabilities in real time.

Original post →

More from Embodied

Embodied channel →