Claude Opus 5 aces 20x20-digit multiplication eval, 1200/1200 correct with max reasoning
maksym_andr · x · 2026-09-17
Developer maksymandr ran a minimal eval testing whether frontier LLMs can multiply large numbers without external tools. Claude Opus 5 scored 100% on multiplications up to 20x20 digits (1200/1200 correct), but only at maximum reasoning effort.
The author argues this simple eval is newly relevant for measuring no-CoT capability, especially for recurrent-depth architectures like GPT-6-Astra, which reportedly cannot run without CoT at all.
Related event: Claude Opus 5 Aces 20×20-Digit Multiplication, 1200/1200 Without Tools(7 posts)→
More from Models
- Mystery model 'Union Alpha' tops DeepSWE, claiming GPT-6-class capability at DeepSeek-level prices — daniel_mac8 · 2026-09-17
- Jensen Huang at All-In Summit: AI leadership will be built by everyone, open and closed models both matter — NVIDIAAI · 2026-09-17
- Model 'Jev' shows well-calibrated probabilities: 1.74pp average calibration error across benchmarks — hackgoofer · 2026-09-17
- Filtering synthetic envs where Qwen deterministically fails but GLM5.3 solves reveals odd behaviors — kalomaze · 2026-09-17
- Why keep a no-thinking mode in 2026? CoT monitoring and performance both demand reasoning — scaling01 · 2026-09-17
- Dev builds sourced comparison of realtime speech-to-speech models across six providers — mahimairaja · 2026-09-17