LLM WikiRace benchmark tests LLM planning and reasoning via Wikipedia race game
hbouammar · x · 2026-09-22
- LLM WikiRace is an open-source GitHub framework that evaluates LLM planning and reasoning through the Wikipedia Race game.
- Supports evaluating OpenAI models or local vLLM deployments (batched or async), plus an interactive human-playable demo and embeddable WikiRaceEnv/WikiRaceEvaluator APIs.
- Configurable game and model parameters, structured logging; requires Python 3.10+ and a CUDA-compatible GPU.
More from Research
- Xiaomi's CodeMidas turns source code into RL environments, doubling DeepSWE to 21.7% — maier_ak · 2026-09-22
- HeyGen and Kaggle launch Code2Video Bench for motion graphics code generation — MeganRisdal · 2026-09-22
- Alibaba's Qwen team launches RecreationBench to test hybrid computer-use agents by app recreation — TianbaoX · 2026-09-22
- CVPR 2025 organizer: 90%+ of first authors had one submission, so authorship caps won't cut volume — CSProfKGD · 2026-09-22
- Toby Ord on swarm scaling: 10,000-agent run cost ~$20M, solved Navier-Stokes in 88 hours — tobyordoxford · 2026-09-22
- Toby Ord: Agent Swarm Estimates Put Intelligence Explosion Parameter λ at 0.5-0.6 — tobyordoxford · 2026-09-22