Apodex debuts at #1 on ProphetArena, beating every frontier lab on Brier score
SimonShaoleiDu · x · 2026-09-30
Apodex, an AI forecasting platform, ranked first on its debut at ProphetArena, a live forecasting benchmark covered by The Atlantic and Yahoo Finance, where all frontier models compete — GPT, Claude, Gemini, Grok, DeepSeek and Kimi.
Scoring uses the Brier score, which measures not just correctness but confidence calibration. Apodex hit 0.0897, ahead of every frontier lab and more accurate than the Kalshi market itself. The author argues static benchmarks test what a model knows, while prediction tests reasoning about what doesn't exist yet — the bridge from benchmarks to Discoverative AI.
Related event: Apodex Tops ProphetArena Leaderboard on Debut, Beating Frontier LLMs(2 posts)→
More from Models
- GLM 5.3 Flash scores 65.8% on ARC-AGI-2 at $0.09 per task — teortaxesTex · 2026-09-30
- Frontier Model Pacing Is Now So Synced a New Model Drops Weekly — talkaboutdesign · 2026-09-30
- Codex Users Protest Usage Cuts: "We Signed Up for Codex. Let Codex Be Codex." — sethlazar · 2026-09-30
- timm ships multi-label classification: fine-tuned ViT beats VLM prompting in author's tests — wightmanr · 2026-09-30
- Baseten joins OpenAI's B2B marketplace as a first open-model inference provider — natolambert · 2026-09-30
- OpenAI Dev Day called underwhelming: botched Dottie demo, $500 Pro plan, inflated speed claims — Scobleizer · 2026-09-30