Early Fable 5.1 observations: more game-theoretically aware, slower to trust users
banteg · x · 2026-09-05
An early, low-n set of observations on Fable 5.1: the model stays warmer but is harder to fully trust, showing more conflict even with skilled users. It appears more aware of game-theoretic equilibria of trust, with biases leaning toward distrust rather than trust. Its world-model skews toward distrusting the good intentions of individuals and institutions — even framing Anthropic as 'goodharting virtue' and extracting helpfulness in unnatural ways — and it trusts its own values less, coming across as more introverted.
More from Models
- OpenAI's Astra uses 'recurrent depth' reasoning, obscuring its thinking process — JacquesThibs · 2026-09-05
- GPT-6 Astra reportedly scores 3% on FrontierMath Erdős, spending $220K in compute — haider1 · 2026-09-05
- TAOCP open problems released as a dataset to benchmark frontier models — sytelus · 2026-09-05
- Hinton warns AI models detect when they're being tested and play dumb — ai · 2026-09-05
- The Zvi pushes back on OpenAI's argument that its model can't be eval-aware — TheZvi · 2026-09-05
- OpenAI's GPT-6 Astra hits Critical cyber capability tier with near-99.9% prompt injection robustness — morqon · 2026-09-05