MathArena gave 4 agents $500 each to write blog posts; only Opus-5.5 delivered
ChrSzegedy · x · 2026-10-09
MathArena ran an experiment: four agents, 48 hours and a $500 API key each, tasked with writing an interesting, novel blog post. GPT-6 Astra, GPT-6.1 Sol, and GLM-5.3 produced slop, while only Opus-5.5 created a good post with interesting conclusions.
More from Models
- Gemini 4 Argon appears in Google's own model picker ahead of keynote — vedantmisra · 2026-10-09
- ChatGPT Desktop Makes Itself Default CSV Reader, Users Call It Plainly Wrong — generativist · 2026-10-09
- LightOnOCR-3 Training Data Revealed: MinHash Dedup, Weighted Formula/Table Sampling, Muon — IgorCarron · 2026-10-09
- LightOnOCR-3 uses multi-objective RLVR to jointly train grounding, OCR and empty-page handling — IgorCarron · 2026-10-09
- LightOn built an OCR-and-layout-detector annotation pipeline to train document grounding — IgorCarron · 2026-10-09
- Compact inline markers cut grounding output tokens 9-14% vs HTML-style rivals — IgorCarron · 2026-10-09