LLMs reproduce classic rat experiment, seeking steered-positive zones and shutting off bad states
awjuliani · x · 2026-10-08
Researchers ran a classic rat behavioral experiment on LLMs, steering models positively or negatively as they read about two meaningless "zones." Despite identical tokens, models sought positively steered zones and avoided negative ones — and given a lever, they shut off the bad states. Full thread in the original post.
More from AGI Musings
- Model Spotted Reusing Its Own Prior Results in Reasoning Traces, a Spark of Theory Building — QuintinPope5 · 2026-10-08
- Dev Argues Coding Is Solved by AI, the Hard Part Is Coming Up with Novel Ideas — Suspicious_Fun_6338 · 2026-10-08
- How many of us will be the last generation to die of something curable? — rand_longevity · 2026-10-08
- AI agents threaten India's $250B IT services engine and 6 million jobs, column warns — nordicinst · 2026-10-08
- Game Fan Translator Deletes 3 Years of Work After Finding AI Output Better Than Their Own — PixelGray38 · 2026-10-08
- Mathematicians really do think like this: the proof is a side effect of enjoying the process — mjuric · 2026-10-08