Hands-on: Astra demos are impressive but still fumbles multi-step research reasoning

xiaosun86 · x · 2026-09-07

Pushing back on the flood of hype around Google's Astra, the author shares hands-on results: it is genuinely better — especially in presentation — but still fails at single research-level reasoning steps.

Meaningful work requires many correct steps in a row, and Astra consistently lands "a bit off" even after corrections. His verdict: still a long way to go. A shared ChatGPT conversation is linked as a comparison case.

Related event: Hands-On Pushback: Astra Impresses in Demos But Falls Short on Reasoning(4 posts)→

Original post →

More from Models

Models channel →