GPT-5.6 Tops FutureSim Forecasting Agent Leaderboard

maksym_andr · x · 2026-08-06

Following its Best Paper award at the ICML Forecasting Workshop, the FutureSim benchmark updated its test set to replay world events from April-June 2026. GPT-5.6-Sol now leads the board, executing over 15,000 tool calls in a single 24+ hour run. While starting with similar accuracy to Fable-5, GPT-5.6 demonstrated strong test-time adaptation.

Original post →

More from Models

Models channel →