AutoCAD-Bench puts GPT 5.6 sol at 46% on precise CAD tasks via computer use
DevvMandal · x · 2026-07-24
AutoCAD-Bench measures whether models can do precise CAD work through computer use
DevvMandal says they released AutoCAD-Bench, a benchmark for testing whether AI models can complete precise AutoCAD tasks using computer use alone.
- The benchmark focuses on hands-on CAD操作 rather than simple text answers.
- In the reported results, GPT 5.6 sol leads with 46%, and is said to solve many basic and intermediate tasks in one shot.
- Other models in the chart include Terra (14%), Fable 5 (10%), Luna (8%), while Opus 4.8, Kimi K2.5, and Qwen3.7 score 0% in this setup.
The post points readers to a full report for details.
Related event: AutoCAD-Bench Released: GPT-5.6 Sol Leads in CAD Tasks(4 posts)→
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- giffmana: the env being used in training is part of the point — giffmana · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11