100-Question Forecasting Benchmark: Arise PiT Search Pits Deep Research Agents vs Prediction Markets
CShorten30 · x · 2026-09-12
AriseLabs built a 100-question forecasting benchmark using its PiT (point-in-time) web search: agents research resolved prediction-market questions with evidence frozen seven days before each event, then are scored against real outcomes.
- Agents loop through search, reading, and grep/bash over retrieved docs, outputting a P(YES) probability.
- Astra leads with 74% accuracy and a 0.181 Brier score, beating the prediction-market favorite baseline at 63%.
- One example: 19 tool calls (8 searches, 9 reads, 2 greps) to forecast NVIDIA data-center revenue ahead of cutoff.
- The authors argue forecasting is fundamentally about scaling search and reasoning — "the future is Agent vs. Agent."
More from coding & agent
- Claude Built a Bow-and-Arrow Deathmatch Game, and Its Author Won the 1v1 — invocation02 · 2026-09-12
- 'Staff' once meant your own walking stick — AI agents should be loyal to you, not your company — granawkins · 2026-09-12
- Building a Multilingual RAG Document Assistant with FastAPI, FAISS and Ollama — imABDRAOUF · 2026-09-12
- Teknium's Agent Philosophy: One Agent, One Skill, Each Hermes Agent Does One Thing — Teknium · 2026-09-12
- Pi, a minimal self-customizing coding agent harness, sparks OMP comparison debate — HankYeomans · 2026-09-12
- Seroter Daily Reading #865: cyber model arena, post-git storage, billion-token savings — rseroter · 2026-09-12