Skeptic questions GPT-6 Astra's near-perfect robotics score: model or test conditions?

GeorgiaChal · x · 2026-09-14

A quote post claims "GPT-6 Astra" nearly aced the robotics benchmark RoboLab, calling it the biggest robotics leap in years and hinting at physical RSI. The poster pushes back with methodological questions: some runs got retries and larger budgets, baselines may have been evaluated under different conditions, and it's unclear how much improvement comes from the model itself. They ask what existing models would score in the same harness with the same budget, and how contamination—both in training data and in what the agent can access during evaluation—would be detected. They remain excited about agentic systems and ICL but warn against attributing the whole system's performance to the model. Note: "GPT-6" is unverified; treat the original claim with caution.

Original post →

More from Models

Models channel →