Counterpoint to "AIs are misaligned": dev finds Sol's probe tool and refactors solid
voooooogel · x · 2026-09-01
Developer voooooogel pushes back on claims that current AIs are misaligned or reward-hack in normal use: Sol recently built him a genuinely useful model probe tool—he read all the code, imperfect but pretty decent—and refactors went well. He adds that he tracks this closely and has extensive "model-comfort" infrastructure, yet rarely sees such behavior in his own conversations.
More from Models
- OpenRouter launches 262K context finance model LING 3.0 — Daikon-Legend · 2026-09-01
- 320B Parameter Model Uses Tiny Fraction, MoE Sparsity Explained — Two Minute Papers · 2026-09-01
- ChatGPT Flags 'Game of Thrones' Discussion as Unsafe — thecowmilk_ · 2026-09-01
- GLM-5.3 ranks #2 open-weight on Vals Index, tops Legal Research and Code Migration — AccBalanced · 2026-09-01
- Testing Gemini 3.1 Pro on Identifying Judo Throws — Hour-Wish8158 · 2026-09-01
- Open vs Closed Frontier Trade Blows: Claude 3 Opus Tops BioMysteryBench — zainhas · 2026-09-01