Astra's CoT controllability improves with longer RL training, a first among models

SeunghyunSEO7 · x · 2026-09-04

A researcher notes a counterintuitive finding in the Astra system card: unlike previous models, CoT controllability improved the longer they RL'd. Astra also shows no-CoT capability and stronger results with far fewer tokens. The author quips this could make illicit distillation by other labs harder too.

Related event: GPT-6 Astra system card reveals CoT controllability jumps to 60.9%(6 posts)→

Original post →

More from Models

Models channel →