Astra looks great in demos but falls short on multi-step research-level reasoning, tester says
xiaosun86 · x · 2026-09-07
Reacting to the wave of hype from big accounts, user xiaosun86 says the platform is "full of posers." His hands-on verdict: Astra is indeed better, especially at presentation, but still fails at research-level tasks that require many steps to all be correct.
He cites a shared ChatGPT "Visualize Trap Effects" session as evidence: after corrections the output looks close, but is always a bit off. His conclusion: Astra still has a long way to go for genuine research-level thinking.
Related event: Hands-On Pushback: Astra Impresses in Demos But Falls Short on Reasoning(4 posts)→
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- 31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day — ionutvi · 2026-09-07