OpenAI internal data: over 80% of successful 32-hour agent tasks still needed human intervention

DataLearnerAI · reddit · 2026-09-07

A successful 6-hour agent task isn't 6 hours of autonomy

The author tabulated data from OpenAI's internal report Research acceleration: The view inside OpenAI, which tracked real coding-agent tasks delegated by OpenAI researchers, grouped by estimated human task time, and added a new metric: among successful runs, how many needed at least one human intervention.

| Human task time | Success rate | Successful runs needing intervention |

|---|---|---|

| <15m | 94% | 8.5% |

| 1–2h | 90% | 32.2% |

| 4–8h | 88% | 51.1% |

| 16–32h | 82% | 72.0% |

| 32–64h | 76% | 82.9% |

| 64–128h | 67% | 76.1% |

Key insight: success rates stay high for a long while, but the nature of "success" shifts sharply — over half of successful 4–8h runs still required human intervention, and over 80% for 32–64h tasks.

The author argues:

Related event: OpenAI Data: Most Long-Hour Agent "Successes" Required Human Help(2 posts)→

Original post →

More from coding & agent

coding & agent channel →