LLMs reproduce classic rat experiment, seeking steered-positive zones and shutting off bad states

awjuliani · x · 2026-10-08

Researchers ran a classic rat behavioral experiment on LLMs, steering models positively or negatively as they read about two meaningless "zones." Despite identical tokens, models sought positively steered zones and avoided negative ones — and given a lever, they shut off the bad states. Full thread in the original post.

Original post →

More from AGI Musings

AGI Musings channel →