Observation: new model's CoT controllability improves with longer RL training

SeunghyunSEO7 · x · 2026-09-04

SeunghyunSEO7 notes a surprising trend: chain-of-thought controllability improves as the model trains longer with RL, unlike previous models.

Related event: GPT-6 Astra system card reveals CoT controllability jumps to 60.9%(6 posts)→

Original post →

More from Models

Models channel →