Researchers flag LLM checking limits: validation is fluent, not formally verified

anshulkundaje · x · 2026-09-29

In a candid thread amplified by Anshul Kundaje, researcher Michael Zlin points out a deep limitation of AI systems in scientific discovery: even with sufficient candidate coverage, a model's "checking" is not formally verified — there is no mechanism guaranteeing it actually validated a claim rather than producing a fluent-sounding validation. The reflection underscores that fluent output doesn't equal trustworthy reasoning, and the lack of formal guarantees remains a hard problem for LLM-driven scientific verification.

Original post →

More from Research

Research channel →