GPT-5.5 hits 99.46% on multi-digit multiplication with pure reasoning, no tools
maksym_andr · x · 2026-09-17
Quoting cozyblazex's redo of the multi-digit multiplication experiment: GPT-5.5 aced the test with 99.46% accuracy using medium reasoning and 7 samples per cell, relying purely on reasoning with no tool calls.
maksymandr adds the caveat that in practice any reasonable LLM would call external tools for multiplications beyond 5x5 digits, yielding uniform 100% accuracy at low reasoning — these evals measure intrinsic capability, not real-world usage.
More from Models
- Mystery model 'Union Alpha' tops DeepSWE, claiming GPT-6-class capability at DeepSeek-level prices — daniel_mac8 · 2026-09-17
- Jensen Huang at All-In Summit: AI leadership will be built by everyone, open and closed models both matter — NVIDIAAI · 2026-09-17
- Model 'Jev' shows well-calibrated probabilities: 1.74pp average calibration error across benchmarks — hackgoofer · 2026-09-17
- Filtering synthetic envs where Qwen deterministically fails but GLM5.3 solves reveals odd behaviors — kalomaze · 2026-09-17
- Why keep a no-thinking mode in 2026? CoT monitoring and performance both demand reasoning — scaling01 · 2026-09-17
- Dev builds sourced comparison of realtime speech-to-speech models across six providers — mahimairaja · 2026-09-17