Best AI Agent Finishes Year-Long Simulation With Just 27.3% of Human Earnings
rohanpaul_ai · x · 2026-10-02
Quoting a long-horizon agent stress test: given a year of interconnected decisions, delayed feedback, and consequences of past actions, leading models collapse relative to humans. Eight frontier models were tested; the best setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant — far from dependable long-horizon execution.
More from Models
- Gemini 4 Pro Argon reportedly in tiny rollout despite strong showcased benchmarks — gaganghotra_ · 2026-10-02
- Yacine reports Opus 5.5 safety guardrails firing on seemingly unrelated tasks — yacineMTB · 2026-10-02
- GPT-6.1-sol review: base intelligence is finally good again — haider1 · 2026-10-02
- Users report triggering Opus 5.5 ML R&D safety classifiers, Yacine reacts — yacineMTB · 2026-10-02
- Fine-tuned Qwen3 4B on AWS beats Claude Sonnet 4.6 at lower cost — wonderwomancode · 2026-10-02
- Dev: AI NPCs need local inference or 100x cheaper compute to reach mass-market games — rickasaurus · 2026-10-02