How a ChatGPT-like system actually works: request-to-stream architecture explained
jawadhamza · reddit · 2026-10-06
- A breakdown of the full production flow of a ChatGPT-like system: user → API gateway → auth/rate limiting → AI orchestrator → context & memory → LLM → safety/processing → streaming response.
- The hard part isn't calling the LLM API but engineering at millions of concurrent requests: request distribution, conversation history storage, context budget, token-by-token streaming, provider failover, rate limits and queues, latency and cost control, and horizontal scaling.
- Accompanied by episode 2 of the author's "System Design in the AI Era" video series with a full visual walkthrough.
More from Infra
- Modal Runtime keynote ships Endpoint Candidates, Clusters and VM Sandboxes — graceisford · 2026-10-06
- Dev shelving Cloudflare K2 support: no event push to Workers, API not in spec — samgoodwin89 · 2026-10-06
- PyTorch Consolidates Media Stack: TorchCodec Becomes Single Home for Decode/Encode — PyTorch · 2026-10-06
- AI-generated key-value stores beat general-purpose DBs via specialized architecture — CShorten30 · 2026-10-06
- Only 30-40k robots installed in the US last year — the case for self-replicating factories — ihorbeaver · 2026-10-06
- Reflection says Beam hit ~19% MFU with 92% goodput during its 4-week RL run — alexpolozov · 2026-10-06