Astra's big reasoning jump without chain of thought raises covert-reasoning concerns

JeffLadish · x · 2026-09-11

Safety researcher Jeff Ladish flags a concerning jump in Astra's reasoning ability with no chain of thought. Neel Nanda replicated the UK AI Security Institute's finding using his own private benchmark. Ladish argues this greatly increases how much covert reasoning the model could be doing — undermining chain-of-thought-based model monitoring.

Related event: Astra's Strong No-CoT Reasoning Raises Hidden-Reasoning Safety Concerns(3 posts)→

Original post →

More from Models

Models channel →