giffmana and NandoDF debate: LLMs still fail at experiment design, with no fix in sight
giffmana · x · 2026-09-07
NandoDF argues LLMs (citing Fable-type models) still suck at running scientific experiments and generalizing out-of-distribution or beyond verifiable domains. giffmana agrees, saying he probes experiment design and sensible analysis/next-step reasoning every other week and it's consistently weak — though he can't tell whether that's a "for now" or "for many years" problem.
More from AGI Musings
- Michael Nielsen: writing 50 paragraphs isn't 50x harder — scaling is nonlinear — michael_nielsen · 2026-09-07
- OpenAI Chief Scientist Jakub Shares Candid Thinking on Recursive Self-Improvement and ASI Alignment — BorisMPower · 2026-09-07
- Dev argues CS unemployment spike is macro-driven and SWE is just hit first by AI — menhguin · 2026-09-07
- Debunking the 'CS Unemployment Is Up Because of AI' Narrative in Three Points — menhguin · 2026-09-07
- Even If CS Degrees Lost Absolute Value, Their Relative Value Is Rising — menhguin · 2026-09-07
- Hugging Face CEO says publicly disclosing an agent cyberattack reinforced his push for 100x more AI transparency — kuchaev · 2026-09-07