Exploring RL Training: Why Do Models Remain Human-Readable?

kalomaze · x · 2026-08-12

AI researcher kalomaze shared insights on why LLMs trained with reinforcement learning (RL) still output human-readable language instead of degenerating into an unreadable 'neuralese.'

He argues that neuralese is not necessary for optimal performance but is rather a free variable prone to nth-order drift. He cited his own ablation studies showing that multiplying RLVR by a classifier's coherence score still works, proving it is easy to synthesize nonsense, which necessitates constraints.

Original post →

More from Research

Research channel →