clef-omni demos real-time audio-video decision model across 8 interactive use cases
Xianbao_QIAN · x · 2026-10-11
clef-omni is a real-time audio-video decision model positioned as the ultimate on-device use case for Jev-style models, with 8 live interactive demos: gesture recognition, visual inspection and ticket triage (PCB inspection, QA, image moderation), screenshot-based next-action suggestions, robot grid navigation, a Voice Agent with real VAD and interruptible responses over WebSocket, package acceptance checks, parking spot detection, and voice moderation. The authors argue this model class will drive adoption of Voice Agents and embodied AI, with every demo showing decisions and probabilities in real time.
More from Embodied
- LILYGO BB keyboard device impresses in hands-on with Meshtastic mesh networking — glenbeer · 2026-10-12
- Paradromics BCI lets clinical participant communicate via imagined speech — PeterDiamandis · 2026-10-12
- Xpeng begins mass production of its Iron humanoid robot, Europe rollout in H2 2027 — LinusEkenstam · 2026-10-12
- Sim scaling laws praised in robot training: why pay millions for real-world data? — lukas_m_ziegler · 2026-10-11
- AHA warns wearable blood pressure devices are inaccurate and may pose patient safety risks — EricTopol · 2026-10-11
- Galbot humanoid robots staff Hong Kong 24-hour stores, grabbing drinks off shelves — CyberRobooo · 2026-10-11