What actually breaks when LLM features meet real users, from production experience
Early-Sir-932 · reddit · 2026-09-30
A developer shares hard-won lessons on how LLM features fail in production — rarely the model itself, but the layers around it: no eval set (changes are vibes without a scored set), prompt-level guardrails that models talk their way out of (deterministic code enforcement is a must), runaway costs from non-idempotent retries and over-routing to big models, garbage inputs (use docling/llamaparse for documents), and silent failures you only learn about from user tickets without per-step tracing.
More from coding & agent
- Figma restricts remote MCP server to whitelisted clients, mitsuhiko slams it for missing the point of open protocols — mitsuhiko · 2026-09-30
- Redditor builds a standalone Windows game with ~150 Claude prompts — Ill-Range-4954 · 2026-09-30
- RemCTL 2.0 turns Apple Reminders into a native, open-source ChatGPT extension — rudrank · 2026-09-30
- SKATE: an open-source workshop memory OS built with Claude Code, plus a 3D-printed mic — lilweedbitch69 · 2026-09-30
- Claude Code offers $250 free cloud credits via /claim-credit command — daniel_mac8 · 2026-09-30
- AWS Shows How to Build a Multi-Agent Music Pipeline on Bedrock AgentCore Runtime Instances — AWS ML Blog · 2026-09-30