SimuVerity Benchmark: Best Agent Scores Only 42.86 on Engineering-Grade Simulink Generation
Ruiqi Zhang · hf · 2026-10-05
SimuVerity is a new benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains, designed to test whether generated models actually satisfy engineering requirements rather than merely compiling or resembling references.
- Tasks ground specs in executable-system profiles plus four families of native simulation scenarios
- A hierarchical evaluator checks artifact delivery, executability and engineering qualification, then scores qualified models on six dimensions: accuracy, output quality, mechanistic fidelity, control/causal integrity, robustness, and dynamic response
- Six agent systems evaluated; the best scores only 42.86 overall, showing structural similarity is a poor proxy for engineering performance
- Some high-scoring models still exhibit severe visual-layout disorder
More from coding & agent
- Photon raises $4.5M seed for open-source messaging layer connecting AI agents to iMessage, WhatsApp, Telegram — SucceededMind · 2026-10-05
- Human-AI collaboration pushes 11-square packing lower bound to 3.875, closing 95.89% of the gap — ctjlewis · 2026-10-05
- Follow-up on the same 11-square packing breakthrough: 300-line verifiable proof via 1,121 weighted dots — ctjlewis · 2026-10-05
- Mobilerun scores perfect 116/116 on Google's AndroidWorld, beating Artemis — SimplyAnnisa · 2026-10-05
- Gemini Robotics-ER 2 plans dual Franka FR3 manipulation, open-sourced with Isaac Sim — Stefania_druga · 2026-10-05
- Dimillian: as models get smarter, intent-based single-thread prompting beats agent orchestration — Dimillian · 2026-10-05