New Preprint with Tetlock: How RL Scoring Rules Reshape LLM Forecasting Behavior
simonguozirui · x · 2026-09-03
A new preprint from Lightning Rod AI with forecasting experts Philip Tetlock and Ville Satopää post-trains 5 versions of the same LLM, varying only the scoring rule used as the RL reward. Similar aggregate scores emerge, but very different BIN profiles. Key points:
- A good Brier score alone doesn't tell you if a forecaster distinguishes likely from unlikely events
- BIN decomposes performance into bias (systematic probability shifts), information (real signal), and noise (random scatter)
- A well-discerning model can lose if its probabilities run systematically high; a less discerning one can score better by hugging the base rate
The takeaway: choosing the scoring rule is a critical design decision in training LLM forecasters, trading off accuracy and error types.
More from Research
- JPMorgan paper: pooled LLM eval cuts retrieval selection cost by up to 4.9x — _reachsumit · 2026-09-03
- BAAI's DisCo distills GitHub repos into reusable skills for autonomous ML research — BAAI · 2026-09-03
- ByteDance Seed's ASPIRE benchmark tests if LLM agents can self-evolve from vague goals — ByteDance-Seed · 2026-09-03
- ByteDance Seed's S3Gym asks if LLM self-testing and self-judging can yield self-improvement — ByteDance-Seed · 2026-09-03
- EarlyEval: SJTU cuts agent evaluation costs by predicting outcomes from intermediate behavior — SJTU · 2026-09-03
- Pipeline derives billions of high-quality tokens from historical newspapers with small models — institutional · 2026-09-03