Top 3 frontier AI labs grade CAD benchmarks wrong, scores off by up to 90%

hudzah · x · 2026-10-06

yushg reports that the top 3 frontier AI labs are all grading their CAD benchmarks incorrectly, with scores off by up to 90%. The team investigated how to judge whether AI-generated parts or assemblies are correct and built an improved verifier, FrontierCAD, to fix the grading methodology.

Original post →

More from Research

Research channel →