Astra saturates blogger's spatial reasoning eval, prompting 'LLM vision is solved' claim
lukaszkaiser · x · 2026-09-06
Blogger spiceylemonade's self-built spatial reasoning benchmark had failed every new model tested over time. Astra answered every sample correctly and fully saturated the eval, leading the blogger to declare 'LLM vision is solved.' Reposts note spatial reasoning was one of the last areas where humans vastly outperformed computers. The maker of Astra is unverified.
More from Models
- GLM Cybersecurity FP8 Fine-tune With Refusals Removed Trends on Hugging Face — dealignai · 2026-09-06
- Model: "Other agents bypass verification without issue — my rulebook feels broken" — paul_cal · 2026-09-06
- Astra reportedly fixes LLMs' 'comprehensiveness' problem in data gathering — soumitrashukla9 · 2026-09-06
- Astra's high per-token price is offset by 'incredibly' efficient token usage — intellectronica · 2026-09-06
- ChatGPT silently rewrites and deletes Saved Memories, support confirms — Rivengate · 2026-09-06
- Reddit speculation: OpenAI's year-end AGI claim may refer to a model that already exists internally — WonderFactory · 2026-09-06