LLM Forecasting Study: Activations More Reliable Than Text
burny_tech · x · 2026-07-11
Goodfire released a paper on LLM forecasting titled "What LLM Forecasters Know but Don’t Say".
The paper points out that while LLM forecasters often appear confident, their calibration is unstable, and their chain-of-thought may not reveal why predictions change. The research found that internal model activations reflect the true state better than the output text.
By using small probes to read these activations, the study aims to:
- Improve confidence calibration
- Detect "silent" shifts in evidence
- Partially recover prediction outcomes before reasoning begins
The authors argue this offers a cheaper and more faithful approach to auditing LLM predictions.
More from Research
- Structural ensembles beat single predictions in TCR:pMHC generalization study — quaidmorris · 2026-07-22
- RSS launches under OMSF to push structural biology data modeling at scale — MoAlQuraishi · 2026-07-22
- enFoldX tops 8 neoantigen scans and an unseen-peptide benchmark — quaidmorris · 2026-07-22
- enFoldX reaches AUC 0.82 on human VDJdb and transfers to mouse at 0.76 — quaidmorris · 2026-07-22
- enFoldX gains accuracy as AF3 ensemble disagreement rises for non-binders — quaidmorris · 2026-07-22
- A 3D ray plot shows how hard this Jacobian counterexample is to read — moultano · 2026-07-22