Only right answers still teach chemistry: popular chemistry benchmark has shortcuts
tak3sh8 · x · 2026-10-06
tak3sh8 reports an odd finding while SFT-ing a small model on a chemistry MCQ dataset: training on only the correct answers — with no explanations — still taught the model a fair amount of chemistry.
Digging in, he found the explanation: the benchmark contains exploitable shortcuts, so models can score well without genuine chemical understanding. He cautions others to be careful with this dataset and concludes the chemistry benchmark is "terrible as is."
More from Research
- A weak paper may earn more citations — everyone cites it to beat it — LucaAmb · 2026-10-06
- tszzl: Mechanistic interpretability is the bare minimum to make AI alignment an engineering discipline — tszzl · 2026-10-06
- Causal Decision-Making Preprint Presented at Simons Institute Trustworthy AI Workshop — murat_kocaoglu_ · 2026-10-06
- Claude produces O(n^1.9992) 3SUM algorithm with Lean proof, vetted by top experts — thegautamkamath · 2026-10-06
- ReSteer open-sourced: fixing VLA policies that ignore mid-execution instruction switches — siddkaramcheti · 2026-10-06
- Researcher presenting Latent Policy States in Reasoning Models at COLM this week — hunarbatra · 2026-10-06