ChapterPal dev: frontier vision models consistently fail to spot obvious webpage conversion artifacts

burkov · x · 2026-10-03

ChapterPal developer burkov shares hands-on experience using AI to convert third-party HTML ebooks into an interactive format. Since third-party books come in hundreds of formats, each conversion pipeline produces unique artifacts — leaked HTML, math meant to be LaTeX rendered as plain text, unescaped dollar signs, missing list bullets, and more.

He hoped vision models plus browser-driving agents could compare source and target pages and flag anything that looks "off," deliberately without giving examples of what artifacts look like. But every new model or harness release fails the same way: models can spot missing words, yet consistently miss leaked HTML, misformatted math, unescaped characters, and missing bullets.

His conclusion: models that ace ARC-AGI-style benchmarks still can't notice what's visually "weird" on a webpage the way humans instantly can.

Original post →

More from coding & agent

coding & agent channel →