GPT-6 Astra hits 95.9% on BenchCAD, suggesting 3D intelligence's 'ceiling' was never fundamental

hanjie_chen · x · 2026-09-08

BenchCAD author Haozhe Zhang reports that GPT-6 Astra scored 95.9% on BenchCAD, his team's benchmark for executable 3D CAD reasoning — a new SOTA. Beyond the leaderboard, he argues this dismantles the strongest argument against MLLMs: that language-dominated pretraining and scarce aligned multimodal data impose a hard ceiling on 3D intelligence. With the right synthetic data, multimodal post-training, executable feedback and agentic tool use, once-unreachable capabilities can be systematically engineered. 3D, long seen as AI's hardest frontier, is rapidly becoming solvable.

Original post →

More from Models

Models channel →