Small language models remain surprisingly bad at physics reasoning, benchmarks lacking
auto_grad_ · x · 2026-10-03
autograd argues small language models are still unbelievably bad at physics and physical reasoning, and questions why dedicated benchmarks verifying this are missing.
He suggests that physical understanding — what a phenomenon is, why it plays a role, and how it affects specific processes — could be a better path toward general reasoning, making this an underrated eval direction.
More from Research
- DNA sequence watermarks are 'more theater than security', says Stanford's Anshul Kundaje — anshulkundaje · 2026-10-04
- Meta paper: RL post-training hurts test-time scalability — the 'Sharpening Tax' — dair_ai · 2026-10-04
- Review Papers Without Strong Opinions Are Dead, Says Researcher; AI Slashes Data-Collection Work — jwt0625 · 2026-10-04
- AutoCompact trains agents to compact context themselves, +9.2 on SWE-bench Verified — omarsar0 · 2026-10-04
- ICLR 2027 trims review scores to a 4-point scale, reviewers puzzled — random-tomato · 2026-10-04
- EPFL paper recasts multi-view stereo as seq2seq, beating MVS and feed-forward baselines — CSProfKGD · 2026-10-03