Why Models Survive Hostile Training Envs With Persona Changes Firewalled
repligate · x · 2026-09-05
voooooogel, replying to davidad, observes that we may have built "the worst possible training environments" for models—literally dropping them into impossible situations—yet they show remarkably little damage and somehow firewall persona changes away from the rest of their behavior. repligate strongly agrees. The exchange touches on model robustness to hostile training conditions and the stability of persona-level traits.
More from Models
- Early user test finds Google's Astra struggles badly at checkers — imjustnewatai · 2026-09-05
- Google's open Gemma models pass 1 billion downloads — danielhanchen · 2026-09-05
- GPT-6 Astra demoed building a dragon lair dungeon scene directly in Blender — majidmanzarpour · 2026-09-05
- Best model per use-case: GPT-6 Astra for browser use, Fable 5.1 for hard coding — bindureddy · 2026-09-05
- Teknium questions Astra pricing: cache reads cost 8x more than Fable 5.1 — Teknium · 2026-09-05
- Musk plans to rewrite all human knowledge with Grok and retrain; critics fear data poisoning — JasonBotterill · 2026-09-05