giffmana and NandoDF debate: LLMs still fail at experiment design, with no fix in sight

giffmana · x · 2026-09-07

NandoDF argues LLMs (citing Fable-type models) still suck at running scientific experiments and generalizing out-of-distribution or beyond verifiable domains. giffmana agrees, saying he probes experiment design and sensible analysis/next-step reasoning every other week and it's consistently weak — though he can't tell whether that's a "for now" or "for many years" problem.

Original post →

More from AGI Musings

AGI Musings channel →