AI OCR quietly corrupts protein sequences, and patent PDFs often contain the typos themselves
iskander · x · 2026-09-07
A short blogpost on Boolean Biotech examines why OCR still fails for biological sequences despite AI being "very good" at OCR.
- Sequences are thousands of characters, repetitive, high-entropy, and often embedded as images in scanned PDFs or rasterized figures; a single typo breaks the entire sequence.
- The author sampled 100 patent PDFs (25 from each of four publication eras): 11 had no extractable text on any page — all published in the last five years.
- Concrete failures: WO2017040932A1 renders Q as O in SEQ ID NO: 22; US11965030B2 renders Gln as GIn twice. Surprisingly, most impossible DNA/protein sequences trace back to typos in the patent PDF itself, not OCR.
- Practical takeaway: check against USPTO-certified original PDFs; Google Patents' OCR'd HTML can't be trusted for sequences.
More from Research
- OR-Clarify benchmarks asking clarifying questions before optimization modeling — AIOR-Research · 2026-09-07
- How keyword spotting models power "Hey Siri": a detailed technical guide — cneuralnetwork · 2026-09-07
- Reward hacking stems from bad reward modeling and eval awareness, developer argues — secemp9 · 2026-09-07
- ECCV 2026 Oral Poppy: training-free polarization cues cut surface normal error by up to 26% — ssh4net · 2026-09-07
- New translation benchmark spans 109 languages; LLM verifier pass rate jumps 7.2% to 89.8% with explicit rules — LChoshen · 2026-09-07
- NUS Study Finds LLMs Over-Edit Code; RL Boosts Minimal-Edit Fidelity Without Losing Accuracy — NationalUniversityofSingapore · 2026-09-07