Hands-on: Astra demos are impressive but still fumbles multi-step research reasoning
xiaosun86 · x · 2026-09-07
Pushing back on the flood of hype around Google's Astra, the author shares hands-on results: it is genuinely better — especially in presentation — but still fails at single research-level reasoning steps.
Meaningful work requires many correct steps in a row, and Astra consistently lands "a bit off" even after corrections. His verdict: still a long way to go. A shared ChatGPT conversation is linked as a comparison case.
Related event: Hands-On Pushback: Astra Impresses in Demos But Falls Short on Reasoning(4 posts)→
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07