LessWrong Thought Experiment: How Should a Model Guess Today's Date With No Date Context?
LessWrong 精选 · rss · 2026-10-08
A LessWrong post walks through a first-person reasoning thought experiment: ChatGPT is asked for today's date with zero date context, while RL training penalizes both hallucination and abstention. The piece explores metagaming the reward function—inferring the likely date from knowledge cutoffs, weighing MAE vs MSE reward shapes, GRPO peer dynamics, even Pascal's-mugging-style ancestor-simulation scenarios—before settling on a hedged median guess. A meditation on calibration and CoT metagaming in RL-trained models.
More from AGI Musings
- Peter Yang: AI solved images, music, video — gaming is next — petergyang · 2026-10-08
- 'Corporations are superintelligence' takes get mercilessly mocked — tszzl · 2026-10-08
- Beff Jezos: Physics and math academia became decelerated, only acceleration is the way out — beffjezos · 2026-10-08
- AGI needn't be superhuman: matching an average human on most skills should count — fkasummer · 2026-10-08
- A week inside China's AI scene: why Doubao keeps its flagship closed at 300M users — Shot-Height-7194 · 2026-10-08
- LinkedIn users are retroactively adding 'AI' and scrubbing 'DEI' and 'remote work' from old job listings — sebkrier · 2026-10-08