Cool presentation aside, Astra still can't nail research-level single-step reasoning
xiaosun86 · x · 2026-09-07
The author pushes back on the wave of big-account hype around Astra, calling the platform full of posers. He concedes Astra is better, especially at presentation, but argues it's still not there for a single step of research-level thinking, sharing a ChatGPT conversation ("Visualize Trap Effects") as evidence. His key point: meaningful work requires many steps all correct, so weak single-step reliability compounds fast.
Related event: Hands-On Pushback: Astra Impresses in Demos But Falls Short on Reasoning(4 posts)→
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07
- GPT-6 Astra generates missile evasion simulation, showcasing stunning capability — algo_diver · 2026-09-07