Best AI Agent Finishes Year-Long Simulation With Just 27.3% of Human Earnings

rohanpaul_ai · x · 2026-10-02

Quoting a long-horizon agent stress test: given a year of interconnected decisions, delayed feedback, and consequences of past actions, leading models collapse relative to humans. Eight frontier models were tested; the best setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant — far from dependable long-horizon execution.

Original post →

More from Models

Models channel →