Opus 5.5 one-shots an animated explainer for new LLM hidden valence steering paper
repligate · x · 2026-10-09
camhberg showcased a new paper with an animated explainer generated in a single shot by Claude Opus 5.5, arguing intuitive, engaging content can now be conjured out of thin air.
The paper itself runs a classic rat-style experiment on LLMs: while models read about two meaningless zones, researchers steer them positively or negatively. Despite identical tokens, models seek positively steered zones and avoid negative ones — and when given a lever, they shut off bad states.
More from Models
- How to Top the Decision Index Vision: Remove the 512x512 Cap in the HF Implementation — antoine_chaffin · 2026-10-09
- Tencent Open-Sources Hy-MT2 Translation Models, 1.8B Shrinks to 440MB After Quantization — aigclink · 2026-10-09
- A month of heavy DeepSeek 4.1 Flash use: near-frontier quality, orders cheaper — victormustar · 2026-10-09
- Tencent open-sources Youtu-Parsing-Omni, a 5B omni-modal parsing model — jacek2023 · 2026-10-09
- Reddit user reports Gemini 3.1 Pro throwing errors on every prompt while other models work fine — Avneesh-dev · 2026-10-09
- Things change, own your weights: a community builder's case for open-weight models — hboelman · 2026-10-09