Why agents stay stuck at 70-75% success: the hard problem of removing humans from the loop

Motor_Fox_9451 · reddit · 2026-09-24

A developer who has built multiple AI agent systems reports their workflows plateau at 70-75% success rate, forcing humans to stay in the loop despite real headcount savings. They suspect Gemini 2.5 Pro's limitations play a role, but the deeper issue is that the tasks are non-standardized with roughly 10% subjective judgment involved. The post asks for proven approaches and learning resources for fully automating the human review step — a core reliability and eval-engineering question for agent builders.

Original post →

More from coding & agent

coding & agent channel →