Missing API for general real-time LLM agents: AsyncLLM preprint sparks interface debate

phill1992 · reddit · 2026-09-30

A Reddit discussion asks whether real-time LLM interaction can be generalized: models already handle mid-thought stimuli in specific cases (robotics agents, GPT-4o voice, video calls, mid-turn steering in Codex), but there's no general API for building real-time agents the way developers write custom MCP tools.

Points raised: vendor real-time APIs (Gemini Live, OpenAI Realtime) target voice/coding niches; compelling use cases include coding agents you steer mid-debug, deep-research agents that react to feedback mid-search, and voice-driven assistants. The closest general approach is the AsyncLLM preprint, where programmers write asyncio coroutines with shared-memory blocks for agent communication — but it still requires building a low-level inference pipeline. The open question: what would the right interface look like for a general async-agent API that OpenAI/Anthropic could expose for feeding live event streams?

Original post →

More from coding & agent

coding & agent channel →