FailureAtlas: most severe LLM gateway failures return HTTP 200 and silently corrupt your app
its_vayishu · x · 2026-09-28
metriqual published a paper, FailureAtlas, taxonomizing failure modes in multi-provider LLM serving infrastructure (gateways/reverse proxies). Key finding: the most operationally severe failures are silent — they return HTTP 200, pass every standard health check, and corrupt conversation state, requiring semantic-level observability to detect.
Highlights:
- Two-axis taxonomy: failures classified by origin layer (Network/Transport, Streaming/Protocol, State/Session, Model Behavior, Governance/Cost) and detectability (Loud vs. Silent).
- Five verified entries sourced from public bug reports and first-hand stress testing, each with root-cause analysis; three include standalone reproduction scripts.
- Two first-hand silent failures: a concurrency race causing history loss, and a streaming index collision corrupting tool-call payloads.
Implication: HTTP status codes and health checks alone are insufficient for production LLM features; semantic monitoring is required.
More from Infra
- US AI datacenter spending as share of GDP now exceeds historic railroad and highway buildouts combined — rvp · 2026-09-28
- Used RTX 3090 prices creep toward $1,500 on eBay amid GPU shortage — sleight42 · 2026-09-28
- gufo inference doubles prefill speed vs llama.cpp forks for Qwen 3.8 Flash Next on Strix Halo — fallingdowndizzyvr · 2026-09-28
- Scraping p50 stabilized at 2s: keep your app and databases colocated — DanielLockyer · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- Why credit quality matters most in compute: clusters are opaque, people are trackable — AccBalanced · 2026-09-28