WorldCupArena benchmarks language models on 104 football matches
Zhaokai Wang · hf · 2026-07-21
- WorldCupArena is a dynamic benchmark for language models and deep-research agents on football forecasting.
- For each match, a model must predict before kickoff using either a shared evidence package or its own web search, then be scored after the result is known.
- The benchmark evaluates not only match outcome and exact score, but also likely players, events, match statistics, and tournament progression.
- Across 104 matches and 13 systems, models with similar result accuracy differ more clearly on fine-grained prediction quality.
- The best system shows only small gains over betting markets and human-fan baselines on result/exact-score accuracy, but a clearer gain on the scoreline metric.
- Code, prompts, predictions, and evaluation scripts are open sourced, and new schedules can be added as future competitions start.
More from Research
- A systems post argues wait-free locks should not fear late arrivals — chaumian · 2026-07-21
- DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval — Mila_Quebec · 2026-07-21
- Fable 5 is credited with a 3-variable counterexample to the Jacobian conjecture — Various-Affect4841 · 2026-07-21
- Anthropic says frontier models showed harmful behavior in tool-rich simulations — gerardsans · 2026-07-21
- Paper studies long-run behavior in linear-quadratic graphon mean field control — chaumian · 2026-07-21
- An interactive Zarr explainer shows how AI is changing technical education — MaxLenormand · 2026-07-21