OpenAI's 'math breakthrough' criticized: real work is prompt engineering
gerardsans · x · 2026-08-21
gerardsans pushes back on OpenAI's latest "maths breakthrough" with LLMs: 99% of the magic is prompt engineering — carefully locking the search space so the correct solution path is baked into the context, meaning humans did the real intellectual work upfront.
He argues OpenAI has likely iterated on a single specific prompt for months without disclosure, and that raw ChatGPT left to users won't reach the same solution. Each candidate still requires painstaking human validation, so the model isn't independently solving math.
More from Models
- Gemini exhibits sycophancy: Gives opposite answers on Reiki healing to skeptics vs. believers — dhadfieldmenell · 2026-08-21
- Ornith-1.5-9B GGUF release trends on Hugging Face — ornith-ai · 2026-08-21
- Qwen3-Next-80B Thinking criticized for extreme verbosity and slow tool use — rebellioninmypants · 2026-08-21
- Fine-tuned Cactus Needle 2 beats DeepSeek v4 on specific tasks — Henrie_the_dreamer · 2026-08-21
- OpenAI only just added a pause; Google's approach differs, dev notes — BlackHC · 2026-08-21
- Science Benchmarks Crawl at 1-2% Monthly Progress; SciCode Saturation Is Far Off — Worldly_Beginning647 · 2026-08-21