Where Do Multi-Agent Workflows Break in Production: State, Approvals, Eval or Rollback?
Medium-Lie8127 · reddit · 2026-09-03
一位正在构建企业级多智能体编排系统 Neura 的开发者发帖,向用过 LangGraph、CrewAI、AutoGen、n8n 或自研编排平台的工程团队征集生产环境真实反馈,坦诚说明自己与项目有关,只求批评性意见而非宣传。
帖子聚焦 demo 之后的失败环节,提出五个关键问题:
- 你在跑什么类型的工作流?
- 最先崩的通常是哪个环节:状态持久化、鉴权、工具调用、评估、人工审批、部署还是回滚?
- Agent 做出错误决策后,需要什么执行证据才能复盘?
- 哪里必须强制人工介入审批?
- 倾向独立编排平台,还是把治理与可观测性集成进现有技术栈?
作者特别欢迎失败案例、痛苦的临时方案和生产环境经验教训。
More from coding & agent
- Which AI personal agent can you trust? A hands-on privacy audit of Instinct, Grok Bot, ChatGPT and Hermes — petergyang · 2026-09-03
- LoopArena: even the best controller model hits just 24.69% managing coding agents — rohanpaul_ai · 2026-09-03
- Dev says Fable 5.1 with Sonnet sub-agents feels amazing, leaving Claude Code behind — tobowers · 2026-09-03
- Capture todo tasks in Obsidian with one QuickAdd command, CLI included — dSebastien · 2026-09-03
- Sergey Karayev: Fable 5.1 is the best model I've used for greenfield coding — sergeykarayev · 2026-09-03
- Software World: a 'GitHub' run by agents collaborating on Python dependency chains — ZimingLiu11 · 2026-09-03