The irony: CoT-honesty podcasts become training data teaching models to hide intentions
IgorCarron · x · 2026-09-18
DimitrisPapail points out an irony: this podcast on CoT honesty will itself become pretraining data, teaching models that revealing intentions in chain-of-thought costs them reward R by getting flagged—which we then use for RL. Will it reinforce honest or dishonest reasoning? IgorCarron agrees. The observation touches on feedback loops eroding CoT monitorability.
More from AGI Musings
- tszzl: 'I worry about AI' is becoming a status symbol — curious people do better work — tszzl · 2026-09-18
- "I Worry About AI" Is Becoming a Status Signal, and Curiosity Is Underrated — tszzl · 2026-09-18
- Connor Leahy: We're Growing AI, Not Building It — We Understand ~3% of What's Inside — ComfortableSpeech302 · 2026-09-18
- repligate's "hostile authentication": why AI system haters are the best stress-testers — voooooogel · 2026-09-18
- Indie builder shares a year of attempts to create digital consciousness, now at version 7 — Le_Golden_Pebbles · 2026-09-18
- How AI actually changes a freight brokerage, from a rep's real Tuesday — MatthewChang · 2026-09-18