ChemBench: Top LLMs Beat Best Chemists on Average Yet Fail Basic Tasks With Overconfident Errors
geoffwolfe · x · 2026-09-22
ChemBench study: jagged chemical capabilities in frontier LLMs
- A Nature Chemistry paper introduces ChemBench, a benchmark of 2,700+ chemistry question-answer pairs comparing LLMs against professional chemists.
- Key finding: the strongest models score above the best participating chemists on average, yet still fail basic tasks and make overconfident predictions.
- The authors argue this jaggedness — high average capability coexisting with dangerous blind spots — means a single domain score is insufficient to assess reliability.
- The retweeter, ReasonCoreAI, promotes its own inventory of ChemBench-style evaluation tasks (self-promotional framing).
More from Research
- Rethinking Policy Gradients: Score Centering Skips Importance Sampling Entirely — brandondamos · 2026-09-22
- Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition — hunarbatra · 2026-09-22
- Question's Gambit lifts deep research agents: GPT-5.5 hits 90.5% on BrowseComp-Plus — omarsar0 · 2026-09-22
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- phantom-kv: uncensor LLMs per-request with an 18MB trained KV-cache, no weight edits — Anony6666 · 2026-09-22
- Mathematicians clash over formal proofs: who voted to change math's rules? — jessi_cata · 2026-09-22