Top 3 frontier AI labs grade CAD benchmarks wrong, scores off by up to 90%
hudzah · x · 2026-10-06
yushg reports that the top 3 frontier AI labs are all grading their CAD benchmarks incorrectly, with scores off by up to 90%. The team investigated how to judge whether AI-generated parts or assemblies are correct and built an improved verifier, FrontierCAD, to fix the grading methodology.
More from Research
- PerturBot Breaks Shortcut Priors in Vision-Language-Action Models With Perturbative Training — Mingyu Liu · 2026-10-06
- Georgia Tech's TextReg Fixes Prompt Distributional Overfitting, Gains Up to +11.8% OOD — GeorgiaTech · 2026-10-06
- SourceLearn Builds Source-Specific Agent Competence, Wins 13 of 15 Benchmarks — GeorgiaTech · 2026-10-06
- Representation-Space MMD Post-Training Boosts Diffusion LMs, More Parallel Decoding at 16B — yresearch · 2026-10-06
- Attention Relay makes embedding models instruction-aware without training via LLM attention weights — _reachsumit · 2026-10-06
- Programmatic Search Agents boost task success by up to 7.56 points over query-based agents — _reachsumit · 2026-10-06