Burkov: on truly novel problems, LLM error rates tell you how often you'll be fooled
burkov · x · 2026-10-02
ML author Andriy Burkov argues that extrapolating an LLM's measured error rate to tasks with no existing human-made information gives you the percentage of time you'll be fooled by the model when working on something no one else has tackled—a caution about hallucination in genuinely novel domains.
More from Models
- Opus 5.5 keeps saying "himbo" — a verbal quirk no previous Claude showed — repligate · 2026-10-02
- AI spend falls in latest Ramp AI Index as frontier price cuts bite; open source under 5% — PaulYacoubian · 2026-10-02
- Mystery model Fledge Alpha spotted before announcement, clues point to Thinking Machines Lab — cephaloform · 2026-10-02
- Ivo's contract agent hits 91% on Legal Agent Benchmark via River AI post-training — ibab · 2026-10-02
- Gemini 4 Argon called 'frontier' yet ranks #48 in analysis and #59 in presentation — AbjectMycologist614 · 2026-10-02
- Artificial Analysis adds refusal timing and fallback-model views to Coding Agent Index — ArtificialAnlys · 2026-10-02