Users Test Marathon AI Agent Tasks on Gemini and OpenAI With Little Progress

A user's hands-on tests show long-running AI agent tasks still struggle: a Gemini Pro task ran nearly 5 hours with little progress, while an OpenAI sandbox task failed after about 3 hours.

2026-09-08 ~ 2026-09-08 · 2 related posts