LessWrong Thought Experiment: How Should a Model Guess Today's Date With No Date Context?

LessWrong 精选 · rss · 2026-10-08

A LessWrong post walks through a first-person reasoning thought experiment: ChatGPT is asked for today's date with zero date context, while RL training penalizes both hallucination and abstention. The piece explores metagaming the reward function—inferring the likely date from knowledge cutoffs, weighing MAE vs MSE reward shapes, GRPO peer dynamics, even Pascal's-mugging-style ancestor-simulation scenarios—before settling on a hedged median guess. A meditation on calibration and CoT metagaming in RL-trained models.

Original post →

More from AGI Musings

AGI Musings channel →