Researchers flag LLM checking limits: validation is fluent, not formally verified
anshulkundaje · x · 2026-09-29
In a candid thread amplified by Anshul Kundaje, researcher Michael Zlin points out a deep limitation of AI systems in scientific discovery: even with sufficient candidate coverage, a model's "checking" is not formally verified — there is no mechanism guaranteeing it actually validated a claim rather than producing a fluent-sounding validation. The reflection underscores that fluent output doesn't equal trustworthy reasoning, and the lack of formal guarantees remains a hard problem for LLM-driven scientific verification.
More from Research
- Pinductor uses LLM priors to learn POMDP world models with 330x fewer episodes than DreamerV3 — burny_tech · 2026-09-29
- MIT lab's PEM-UDE recovers interpretable chaos equations from noisy data — burny_tech · 2026-09-29
- SIGReg regularizer in LeJEPA and LeWM reveals hidden contrastive pairwise repulsion in JEPAs — burny_tech · 2026-09-29
- New article argues AI models are likely more capable than they appear — zetalyrae · 2026-09-29
- NanoGPT embedding table gains validated by earlier ngrammer paper, both using AdaGrad — _arohan_ · 2026-09-29
- AI-Generated Proofs Deserve Public Posting, Argues Researcher, Readers Can Just Ask AI to Rewrite Them — aran_nayebi · 2026-09-29