GPT-5.6 Tops FutureSim Forecasting Agent Leaderboard
maksym_andr · x · 2026-08-06
Following its Best Paper award at the ICML Forecasting Workshop, the FutureSim benchmark updated its test set to replay world events from April-June 2026. GPT-5.6-Sol now leads the board, executing over 15,000 tool calls in a single 24+ hour run. While starting with similar accuracy to Fable-5, GPT-5.6 demonstrated strong test-time adaptation.
More from Models
- Anthropic's Guardrails Trigger False Positives, Frustrating Devs in Coding Workflows — Cozyboy02 · 2026-08-07
- Without a Single Dominating Model, AI Routers Can Outperform Any Individual LLM — Muennighoff · 2026-08-07
- Google Demos Fully Offline Gemma Translator Powered by Raspberry Pi 5 — GlennCameronjr · 2026-08-07
- DeepSeek Apologizes for API Price Hike Controversy, Offers Refunds — teortaxesTex · 2026-08-07
- Leaked Benchmark Scores for Gemini 3.5 Pro — Miserable-Archer-631 · 2026-08-07
- Meta AI Competes in Five STEM Olympiads, Achieves Perfect Physics Scores and Math Gold — AIatMeta · 2026-08-06