Cutting a RAG app to two model calls, and the TTFT problem that remains
Bulky_Ring_244 · reddit · 2026-09-30
The author open-sourced a conversational portfolio app that answers questions about a person's documented work. The hard part wasn't plausible answers—it was stopping plausible-but-unsupported details from turning into autobiography.
Architecture evolution: an earlier design with routing, evidence selection, sufficiency checking, web search, and generation became hard to reason about, so it was reduced to two hosted-model calls:
- Classify the message and rewrite it as a standalone retrieval query when evidence is needed.
- Generate from up to five documents returned by a local Tantivy index.
No agent loop, no runtime web search, no separate LLM judge. If retrieval returns nothing, the assistant declines instead of improvising. Conversation history is kept separate from the trusted index, so visitor messages never become biographical evidence.
Remaining issue: the sequential calls hurt time-to-first-token—streaming can't start until classification finishes. Options under consideration: a smaller/faster classifier model, deterministic fast paths for chitchat, a shorter classification prompt, or collapsing routing and generation (at the cost of grounding control). Corpus curation and retrieval quality now carry more weight—preferable, the author argues, to another model judge that can also be wrong. Apache-2.0, live demo available.
More from coding & agent
- This developer's perfect AI coding stack costs $420/month across ChatGPT, Claude and AmpCode — iannuttall · 2026-09-30
- A tricky UI test for agents: scrolling up to load images in long chat apps — gethackteam · 2026-09-30
- Cloudflare Birthday Week: 6x faster containers, AI Gateway auto-routing, Workers error monitoring — ritakozlov · 2026-09-30
- Tianqi Chen's team open-sources a book on compiler-driven agentic GPU kernel optimization for MLSys — sh_reya · 2026-09-30
- computesdk cuts cold-start latency 6x: median 4.05s to 648ms via co-located containers and VM restore — dinasaur_404 · 2026-09-30
- Golem Builds a Log Incident Pipeline With Jev Classifier and Durable Agents — blaizedsouza · 2026-09-30