Replication of no-CoT evals shows GPT-Astra makes a qualitative jump across all datasets

dhadfieldmenell · x · 2026-09-16

Christine Corryy ran her replication of Ryan Greenblatt's no-CoT evals on GPT-Astra plus Gemini 3.1 Pro, Kimi k3, and Fable 5.1, finding a qualitative jump on every dataset — especially competition math and multi-hop reasoning — reinforcing that GPT-Astra is strikingly good at computation without chain-of-thought.

Original post →

More from Models

Models channel →