Replication of no-CoT evals shows GPT-Astra makes a qualitative jump across all datasets
dhadfieldmenell · x · 2026-09-16
Christine Corryy ran her replication of Ryan Greenblatt's no-CoT evals on GPT-Astra plus Gemini 3.1 Pro, Kimi k3, and Fable 5.1, finding a qualitative jump on every dataset — especially competition math and multi-hop reasoning — reinforcing that GPT-Astra is strikingly good at computation without chain-of-thought.
More from Models
- With retries and pooled selection, Qwen3.8 27B hits 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra — S_Conradi · 2026-09-16
- "Just output probability distributions, never hallucinate": AI safety claim gets mocked — inductionheads · 2026-09-16
- Chinese open models hit 53% of OpenRouter tokens, but closed models still dominate real adoption — ohlennart · 2026-09-16
- Niche AI Use Cases Keep Getting Absorbed Into General Models — samiramanabi · 2026-09-16
- Google reportedly building math-focused DeepThink variant, raw thoughts leak — PMinervini · 2026-09-16
- Player claims to find OpenAI GPT-6 "Astra" easter egg in Fallout 3 — imjustnewatai · 2026-09-16