inclusionAI releases Realtime-Venus: 9B full-duplex audio-visual omni model
jacek2023 · reddit · 2026-09-18
inclusionAI open-sourced two checkpoints of Realtime-Venus on Hugging Face:
- Realtime-Venus-Omni (9B), adapted from MiniCPM-o 4.5: continuously watches and listens, decides when to respond, and generates text and speech on a shared causal timeline. It supports proactive interaction, semantic interruption handling, and training-free long-video memory that archives informative visual moments and reassembles query-relevant context.
- Realtime-Venus-Audio: the same streaming backbone for audio understanding and conversation.
Highlights include native full-duplex conversation (distinguishing backchannels, interruptions, corrections), in-stream <delegate> requests so external tasks never block dialogue (requires the Realtime-Venus-Harness runtime), and native speech output via bundled Token2wav with a reference voice.
More from Models
- Matt Shumer asks if Jev could help with scalable oversight and alignment checks — mattshumer_ · 2026-09-20
- FrontierSWE v2 opens 24.1-point gap: Claude Fable 5.1 scores 56.29% vs GPT-5.6's 32.2% — geoffwolfe · 2026-09-20
- 22M local model beats JEV 93% vs 80% on Banking77 in 8ms on CPU — Prompt Engineering · 2026-09-20
- Jev loses to Gemini on 1,565-email classification benchmark, but dev still wants it in production — socialwithaayan · 2026-09-20
- Jev Detector scans ~10,000 words for AI slop in ~2 seconds, free with no sign-up — socialwithaayan · 2026-09-20
- Open-source 395M "System One" model Von runs on CPU in 25-300ms, beats JEV on all benchmarks — wFXx · 2026-09-20