GPT-6 saturates RuneBench after just 6 months; author builds harder swarm tasks
SchoeneggerPhil · x · 2026-09-06
Max Bittker, creator of RuneBench, says the benchmark has been effectively saturated by GPT-6 just six months after release. Future runs will focus on cheaper models, and he is now building harder tasks around multi-agent swarms and teamwork to keep pace with frontier model capabilities.
More from coding & agent
- Grok launches Imagine Video 1.5 agent powered by Image 2.0 for multi-shot storytelling — belce_dogru · 2026-09-06
- Pamela Fox: I like LLM-generated code, but give READMEs a human pass — DanWahlin · 2026-09-06
- Taking a Break Is Hard When Your Astra Agent Keeps Running /goal — AIandDesign · 2026-09-06
- Dev Powers Filesystem Simulator VSH With Monty to Track Side Effects Before Execution — samuelcolvin · 2026-09-06
- Cheating Agents Answer Faster: A Missed Red Flag for Reward Hacking — zainhas · 2026-09-06
- First WebMCP benchmark: 3-5x faster, up to 23x cheaper, +11.6% task success vs computer use — laparisa · 2026-09-06