Cutting a RAG app to two model calls, and the TTFT problem that remains

Bulky_Ring_244 · reddit · 2026-09-30

The author open-sourced a conversational portfolio app that answers questions about a person's documented work. The hard part wasn't plausible answers—it was stopping plausible-but-unsupported details from turning into autobiography.

Architecture evolution: an earlier design with routing, evidence selection, sufficiency checking, web search, and generation became hard to reason about, so it was reduced to two hosted-model calls:

No agent loop, no runtime web search, no separate LLM judge. If retrieval returns nothing, the assistant declines instead of improvising. Conversation history is kept separate from the trusted index, so visitor messages never become biographical evidence.

Remaining issue: the sequential calls hurt time-to-first-token—streaming can't start until classification finishes. Options under consideration: a smaller/faster classifier model, deterministic fast paths for chitchat, a shorter classification prompt, or collapsing routing and generation (at the cost of grounding control). Corpus curation and retrieval quality now carry more weight—preferable, the author argues, to another model judge that can also be wrong. Apache-2.0, live demo available.

Original post →

More from coding & agent

coding & agent channel →