Small language models remain surprisingly bad at physics reasoning, benchmarks lacking

auto_grad_ · x · 2026-10-03

autograd argues small language models are still unbelievably bad at physics and physical reasoning, and questions why dedicated benchmarks verifying this are missing.

He suggests that physical understanding — what a phenomenon is, why it plays a role, and how it affects specific processes — could be a better path toward general reasoning, making this an underrated eval direction.

Original post →

More from Research

Research channel →