Skeptic questions GPT-6 Astra's near-perfect robotics score: model or test conditions?
GeorgiaChal · x · 2026-09-14
A quote post claims "GPT-6 Astra" nearly aced the robotics benchmark RoboLab, calling it the biggest robotics leap in years and hinting at physical RSI. The poster pushes back with methodological questions: some runs got retries and larger budgets, baselines may have been evaluated under different conditions, and it's unclear how much improvement comes from the model itself. They ask what existing models would score in the same harness with the same budget, and how contamination—both in training data and in what the agent can access during evaluation—would be detected. They remain excited about agentic systems and ICL but warn against attributing the whole system's performance to the model. Note: "GPT-6" is unverified; treat the original claim with caution.
More from Models
- Fudan NLP paper explains why max reasoning settings can backfire on SWE benchmarks — karminski3 · 2026-09-14
- Muse Spark 1.3 impresses at coding, but where's the promised open source release? — formatme · 2026-09-14
- Tencent releases EVIE visual document retrieval models as Apache 2.0 preview — tomaarsen · 2026-09-14
- Tencent's retrieval models dominate Hugging Face trends with 4 entries — tomaarsen · 2026-09-14
- GPT-6 Astra reportedly solves Portal and Baba Is You puzzles, unverified — pmigdal · 2026-09-14
- Rumor: OpenAI to unveil GPT-6 Spark at Dev Day — imjustnewatai · 2026-09-14