Why agents stay stuck at 70-75% success: the hard problem of removing humans from the loop
Motor_Fox_9451 · reddit · 2026-09-24
A developer who has built multiple AI agent systems reports their workflows plateau at 70-75% success rate, forcing humans to stay in the loop despite real headcount savings. They suspect Gemini 2.5 Pro's limitations play a role, but the deeper issue is that the tasks are non-standardized with roughly 10% subjective judgment involved. The post asks for proven approaches and learning resources for fully automating the human review step — a core reliability and eval-engineering question for agent builders.
More from coding & agent
- Salesforce's JitMem curates agent memory at read time, gains up to 16.3 points on benchmarks — Salesforce · 2026-09-24
- Coding agents generate robot training data: VLA finetuned on agent demos runs on real hardware — Haoxiang You · 2026-09-24
- Discrawl: Mirror Discord Guilds Into Queryable Local SQLite — steipete · 2026-09-24
- AiFi workflow: connect Coinbase MCP so your agent reads a daily stock-market brief — MurrLincoln · 2026-09-24
- Qwen3.8-Flash is free in Alibaba's Qoder IDE until Sept 30, no credits needed — Aiden_Tech_Ai · 2026-09-24
- Paper: Re-Evaluate Production Agents with 38.5% of the Benchmark, Within 1.03 Points of Full Score — omarsar0 · 2026-09-24