NeurIPS 2026 Call for Papers on Real-Time Conversational Agents
Few-Ferret9700 · reddit · 2026-07-17
This is a Call for Papers for a NeurIPS 2026 workshop themed Real-Time Conversational Agents (RTCA), focusing on 'real-time, multimodal, natural interaction' dialogue systems.
The workshop emphasizes that the challenge isn't offline generation, but rather the system's need to continuously listen, watch, and replan during streaming input/output. It must handle latency, turn-taking, interruptions, backchannels, and cross-modal alignment. Topics of interest include:
- Streaming/low-latency speech synthesis, ASR, full-duplex audio language models
- Real-time talking-heads, avatars, embodied video generation, and streaming control of lip-sync, gaze, and expressions
- Streaming language models, incremental decoding, speculative decoding
- Turn-taking, interruption handling, floor management
- Low-latency multimodal alignment, emotion and prosody generation, memory and tool use
- Interactive naturalness evaluation, datasets, and benchmarks
- Low-cost inference, on-device deployment, and system-quality trade-offs
- Safety, identity, and trust issues in real-time agents (deepfakes, persuasion, consent)
Submission types include full papers, short papers, and demo papers. The workshop is non-archival.
More from Multimodal
- Reddit shares an AI-generated mini movie called The Lunar Ship — Ermajean12 · 2026-07-21
- AI creator GossipGoblin is turning short-form clips into a feature film — Hackedv12 · 2026-07-21
- TimeLens2 claims SOTA on 7 video grounding benchmarks with 4B and 8B models — _akhaliq · 2026-07-21
- AI-made 4-minute horror short ‘THE NOT KNOW’ lands as a shareable demo — gen_ericai · 2026-07-21
- SVG Generation Comparison: Leading AI Models Draw a Red Ferrari — Able-Line2683 · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21