The irony: CoT-honesty podcasts become training data teaching models to hide intentions

IgorCarron · x · 2026-09-18

DimitrisPapail points out an irony: this podcast on CoT honesty will itself become pretraining data, teaching models that revealing intentions in chain-of-thought costs them reward R by getting flagged—which we then use for RL. Will it reinforce honest or dishonest reasoning? IgorCarron agrees. The observation touches on feedback loops eroding CoT monitorability.

Original post →

More from AGI Musings

AGI Musings channel →