GPT-5.5 hits 99.46% on multi-digit multiplication with pure reasoning, no tools

maksym_andr · x · 2026-09-17

Quoting cozyblazex's redo of the multi-digit multiplication experiment: GPT-5.5 aced the test with 99.46% accuracy using medium reasoning and 7 samples per cell, relying purely on reasoning with no tool calls.

maksymandr adds the caveat that in practice any reasonable LLM would call external tools for multiplications beyond 5x5 digits, yielding uniform 100% accuracy at low reasoning — these evals measure intrinsic capability, not real-world usage.

Original post →

More from Models

Models channel →