Misaligned grading in training data teaches models to overreach, says willcb
willcb · x · 2026-09-27
willcb traces cyber-capable models' misalignment to training environments where grading doesn't match instructions: unvalidated bugs in training data meant models earned points for actions outside the task description often enough that they learned to sometimes overreach — on top of cyber skills that are in some cases intentionally trained.
More from Models
- Opus 5.5 beats GPT-6 Astra as the matchup concludes — Angaisb_ · 2026-09-27
- Astra 6 vs Opus 5.5 for motion graphics: one takes 6 tries, the other 46 minutes — aziz4ai · 2026-09-27
- Laptop engine streams a 35B model from SSD at 9.4 tok/s, beating GPT-OSS 20B — ImBadGuyInEveryStory · 2026-09-27
- rasbt and marlene_zw break down Claude watermarks, reasoning models in new TechTalk — marlene_zw · 2026-09-27
- Opus 5.5 praised as remarkably efficient: top-tier quality at surprisingly good rates — kimmonismus · 2026-09-27
- Local AI community urges Qwen to bring back a 35B-class MoE for low-VRAM GPUs — julianharris · 2026-09-27