Paper says a claimed “100% wrong” MATH set missed many valid answers
burny_tech · x · 2026-08-04
- The paper argues that a prior claim about a "100% wrong" MATH dataset was flawed because the dataset construction missed many valid equivalent answers.
- The authors show that MATH problems can have multiple correct answers and multiple equivalent expressions, so labeling an answer wrong is not always correct.
- The attached figure is from "Spurious Rewards: Rethinking Training Signals in RLVR", which studies how weak or even spurious training signals affect MATH-500 accuracy across models such as Qwen2.5-Math-7B, Qwen2.5-7B, Llama3.1-8B-Instruct, and OLMo2-7B.
- The main takeaway: some reward signals can boost certain Qwen models, but the same signals do not transfer uniformly to other models with different reasoning priors.
More from Research
- A neurosymbolic model is pitched as a way to build ultra-reliable numerical solvers — GaryMarcus · 2026-08-04
- Claude Code Under the Hood: System Prompts Exceed 75%, Size Surges — ClaudeCodeLog · 2026-08-04
- Self-distillation from production traces could make models improve with use — ypatil125 · 2026-08-04
- A new reading list links refactoring economics, agent skills, and Google’s Agent Skills — rseroter · 2026-08-04
- A well-executed PhD full of negative results should still be defendable — prajdabre · 2026-08-04
- Pre-registered study finds a universal floor for hallucination detection, but no universal detector — k01234n · 2026-08-04