Follow-up: the chemistry benchmark is "terrible as is", shortcuts everywhere
tak3sh8 · x · 2026-10-06
In a follow-up to his earlier finding, tak3sh8 notes that other chemistry examples in the dataset have shortcuts as well, and concludes that the chemistry benchmark is "terrible as is" — models can score well via shortcuts without real chemical understanding.
More from Research
- OCBench offers human-like scripted policies for scalable robot BC/RL research — kevin_zakka · 2026-10-06
- Apple research team opens 2027 PhD internships in video models, 3D/4D reconstruction — HildeKuehne · 2026-10-06
- MemAdapter uses counterfactual reasoning to curb memory-induced sycophancy in LLM agents — Ruqing Ning · 2026-10-06
- Peking University's Code2Games gets coding agents to build playable UE5 game worlds — PekingUniversity · 2026-10-06
- QuantCode: domain pretraining + SFT lifts Qwen trading-code pass from 27.8% to 58.2% — Alexey Chernysh · 2026-10-06
- Subsampling and extrapolation keep the Mandelbrot area estimate unbiased near the boundary — geoffreyirving · 2026-10-06