MindTrial: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 hits 96/98

Correct_Tomato1871 · reddit · 2026-10-04

The MindTrial maintainer tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as predecessors (39 text + 59 visual, Python/scientific libs available, 10-call limit per task, xhigh effort, no skipped tasks):

| Model | Passed | Hard errors | Request time |

|---|---|---|---|

| Sonnet 5 | 72/98 | 6 | 5:30:44 |

| Sonnet 5.5 | 94/98 | 1 | 1:23:07 |

| Opus 5 | 88/98 | 4 | 3:40:24 |

| Opus 5.5 | 96/98 | 0 | 0:58:35 |

Sonnet improves most: 22 additional passes with 74.9% less request time and 67.4% fewer output tokens including reasoning; 24 newly passed tasks (18 prior wrong answers, 6 prior errors), only 2 regressions; on the second visual collection it goes from 12/26 to 25/26. Opus gains 8 passes and cuts request time 73.4% with zero hard errors; its two failures are square counting and circle-piece matching.

Tool traces are mixed: Sonnet logs 41 undefined-variable errors across 24 tasks (though 23 tasks with failed calls still pass); Opus has none and only 2 nonzero exits across 155 Python calls. Both use Python far less than predecessors (Sonnet -55%, Opus -51%), suggesting headroom in Sonnet's tool-execution reliability. Versus Astra high, Opus leads by one task (96 vs 95) but Astra is faster once local Python wall time is counted.

Original post →

More from Models

Models channel →