SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo, Chaoting Wang, Guoyu Mou, Keyu Lai, Hanchao Lv, Jiaxu Wang, Yibo Zheng, Aijun Yang, Xiaohua Wang
cs.SE, cs.AI, cs.LG
2026-10-02
SimuVerity scores six agents on 101 executable Simulink tasks in ten domains; Opus 4.8 leads at 42.86, against a 96.01 reference-system mean.
Simulink is the graphical environment used in industry for dynamic models, control design, and simulation. LLM agents can already turn a natural-language spec into a runnable .slx, simulate it, and revise. Existing Simulink benchmarks mostly check whether the file arrives, whether it compiles and runs, and whether the diagram resembles a reference model.
A clean run only shows that the toolchain did not error. Treating one reference diagram as the standard misses valid models with a different structure, and it passes models that look right and behave wrong. Requirement-level tests exist for automotive control, but they stay inside software logic and a few control functions. Power electronics, fluids, batteries, and aerospace sit outside that scope.
SimuVerity, from Xi'an Jiaotong University, is 101 text-to-executable model tasks, 9 to 11 in each of ten engineering domains.
Each task starts from a reference system that runs in native MATLAB/Simulink. Experts execute it and write an executable-system profile: boundary, interfaces, mechanisms, control relations, operating range, and dynamics. The public spec and the tests come from that profile. The prompt states the system, interfaces, operating range, mechanisms, and dynamics. It does not prescribe blocks. A tiltrotor VTOL task, for example, has to hold hover, transition in wind, and waypoint flight. Another diagram is acceptable if the behavior qualifies.
Experts turned the requirements into 595 native scenarios in four families: operating envelope 205, interaction and fault 189, temporal process 79, and mechanism and causal checks 122. Scoring is gated. The candidate must deliver an .slx, then finish native load, diagram update, compile, and simulation. Gate G then checks functional roles, signal paths, and whether the model emits valid evaluation evidence. A failed gate scores 0. Models that pass receive six scores from 0 to 100: A objective accuracy, Q output quality, M mechanistic fidelity, C control and causal integrity, R operating-domain robustness, and D dynamic response and recovery. C applies to 35 tasks, M to 93, D to 98, and A, Q, and R to all 101.
The task score is the gate multiplier times a weighted geometric mean of the applicable dimensions, so one strong dimension cannot hide a critical miss. Some tasks also aggregate by the worst condition. Scenarios, thresholds, and scoring scripts are hidden from the agent, and reference-model directories are not visible. Each system runs each task three times from a clean session. Reported figures are means.
Six agent systems saw the same 101 tasks. The reference systems themselves average 96.01.
| System | Overall | Delivered | Executable | G pass |
| Opus 4.8 + Claude Code | 42.86 | 94.72% | 86.47% | 76.90% |
| GPT-5.5 + Codex | 41.72 | 98.68% | 94.72% | 75.58% |
| DeepSeek-V4-Pro + Claude Code | 28.40 | 97.03% | 88.12% | 61.39% |
| Qwen3.8-Max + Claude Code | 25.01 | 89.77% | 80.86% | 53.80% |
| GLM-5.3-Flash + Claude Code | 4.98 | 96.04% | 48.18% | 13.20% |
| Qwen3.8-27B-FP8 + Claude Code | 1.60 | 30.36% | 13.53% | 6.27% |
| Reference | 96.01 | 100% | 100% | 100% |
Opus 4.8 with Claude Code scores 42.86, with a 95% task-level interval of [37.03, 48.73]. GPT-5.5 with Codex scores 41.72, interval [35.29, 48.31]. The paired test gives p=0.86, so the two are not separable. DeepSeek-V4-Pro, Qwen3.8-Max, GLM-5.3-Flash, and Qwen3.8-27B-FP8 score 28.40, 25.01, 4.98, and 1.60. Adjacent tiers differ at p at or below 0.012. Within a tier they do not.
Delivery, native executability, and G-pass are 94.72%, 86.47%, and 76.90% for Opus, and 98.68%, 94.72%, and 75.58% for GPT-5.5. Pooled across every task-system run, the three rates are 84.43%, 68.65%, and 47.85%. A file that runs can still fail as an engineering implementation.
Averaged across the six systems, end-to-end A/Q/M/C/R/D are 29.03, 38.46, 31.99, 29.21, 27.69, and 27.82. Output quality is highest. Robustness and dynamic recovery are lowest. Opus reaches a Q of 63.24 end to end and 82.21 after G, while post-G R and D stay at 60.73 and 58.65.
Opus also splits by where it fails. Power electronics scores 25.99 with a G-pass rate of 44.44%, so the bottleneck is qualification. Batteries pass G 86.67% of the time and then average only 38.29, so the bottleneck is performance after qualification.
Structural similarity is a poor proxy. Of 95 models from Opus run 1, 39 are reference-aligned, 43 partly similar, and 13 structurally distinct. The latter two groups are 58.95%. End-to-end means are 52.13, 44.05, and 42.37. After G they are 59.80, 52.61, and 61.20. Twelve of 24 top-quartile models are not reference-aligned. Seven reference-aligned models score zero, and two of those still pass G. Two structurally distinct models, wheel-loader coordination and an electric-aircraft powertrain, score 85.10 and 80.07.
A high engineering score does not mean a readable diagram. Experts marked 8 of 31 Opus models scoring at least 70 as severely disordered. Moving blocks and rerouting existing lines, without changing the implementation, removed all block overlap and cut line-crossing density by 22.7% to 85.1%.
On ten cross-domain tasks Opus already solved, full MATLAB MCP scores 78.46. Disabling simulation feedback drops the score to 7.88, a gap of 70.58. Delivery, executability, and G-pass fall from 100% to 53.33%, 33.33%, and 13.33%, and all ten tasks get worse. Removing MCP and leaving batch scripts scores 18.23, with eight of ten tasks worse.
Two experts blindly rated 30 models from Opus run 1. Weighted Cohen's kappa is 0.90. Spearman correlation with the mean expert rating is 0.94, against 0.87 between the experts. On the 18 models with a nonzero score, the correlation is still 0.88.
The leading system delivers 94.72% of artifacts and still scores 42.86, against 96.01 for the references. A benchmark that stops once the model runs will overstate engineering ability.
Distance to the reference is a bad reward. After G, structurally distinct models do not score lower. Simulation feedback is harder to drop: on tasks the agent already solved, the score falls to single digits without it. Layout is separate. The engineering score does not care whether the lines are tangled.
Against older checks of completeness, executability, and similarity to a reference, this bar is stricter. Across ten domains, current agents do not yet have stable engineering-grade Simulink modeling.
The paper says evaluation depends on MATLAB/Simulink, native simulation is slow, and the pipeline does not yet scale.
The main study leaves other gaps. Dimension C covers only 35 of 101 tasks, so it cannot carry a cross-domain claim. Structural labels, the layout review, and the 30-model expert check all use one Opus run. The ablation uses ten tasks the agent already handled, so it does not describe failures on the full set. Gate G and the dimension weights are hand-set per task and tied to the reference profile, with no second expert panel. Five systems run in Claude Code and GPT-5.5 runs in Codex, so the comparison is model plus harness.