Lanyon's benchmark claims 20–250x gains over frontier models on PDEs
burny_tech · x · 2026-07-23
- This is the fuller benchmarking post for Lanyon's neurosymbolic architecture on simple linear PDE problems.
- Across multiple trials, Lanyon is said to beat frontier models such as Claude Fable 5, Claude Opus 4.8, GPT-5.5, GPT-5.6 Sol, and Kimi K3 by roughly 20–250× in speed and 50–250× in token efficiency.
- The post highlights repeated failure modes in frontier models: wrong numerical scheme order, failure to limit discontinuous solutions, and other mathematical/algorithmic mistakes even under detailed prompting.
- It also points to Lanyon's symbolic theorem-proving stack as a differentiator for verifying numerical solvers.
Related event: Lanyon's Neurosymbolic Solver Outpaces Frontier Models by 250x in PDE Tests(2 posts)→
More from Research
- Patch Policy beats a fine-tuned 7B VLA by 18% with 0.7% of the parameters — ylecun · 2026-07-23
- A year-built personal agent was finally beaten by a one-day-old competitor — Antony_Richards · 2026-07-23
- Inkling scores 836 Elo on AA-Briefcase, trailing top open-weight models — ArtificialAnlys · 2026-07-23
- Robotics paper says VLA and world models are not enough for grounded supervision — hbouammar · 2026-07-23
- Google Research: Towards a Quantum Computer That Learns From Its Errors — donutloop · 2026-07-23
- AI could compress decades of biomedical research into days, says Derya Unutmaz — DeryaTR_ · 2026-07-23