Yarin Gal: open-ended research makes LLMs collapse problems into known work
yaringal · x · 2026-09-23
Oxford professor Yarin Gal tested Fable as a "student" researching an open-ended hypothesis and found himself arguing with it when it misread a paper as refuting his hypothesis. His takeaway: LLMs handle well-defined optimization objectives fine, but on open-ended problems they collapse the task onto known work instead of actually answering it — a useful data point on the limits of AI-scientist products.
Related event: Oxford's Yarin Gal: LLMs are poor at real science without scaffolding(3 posts)→
More from Research
- Multi-Agent Efficiency Isn't Just Shorter Speakers — It's Listeners Who Know When to Act — lileics · 2026-09-23
- HANDRAISER Generalizes to Unseen GPT-4o Speakers and Multi-Interrupter Settings — lileics · 2026-09-23
- HANDRAISER's Per-Task Numbers: 48.9% Cost Cut on Debate, 32.2% on Average — lileics · 2026-09-23
- CMU × Meta's HANDRAISER Cuts Multi-Agent Communication Cost 32.2% by Learning to Interrupt — lileics · 2026-09-23
- One Attention Head Matters Across Five ICL Task Families: Mech Interp Paper Accepted at COLM 2026 — JacobSteinhardt · 2026-09-23
- Google Fellow John Platt: ERA AI scientist grew out of an attempt to automate Kaggle — Latent Space · 2026-09-23