AgentWorld Benchmark: Best LLM Only Hits 52% on Long-Horizon Multi-Agent Collaboration

Raphael Shu · hf · 2026-09-28

AgentWorld is an open-source benchmark for evaluating long-horizon multi-agent collaboration among LLMs, and current models fare poorly.

Original post →

More from coding & agent

coding & agent channel →