15.9M-parameter Kraken OCR model beats VLMs hundreds of times larger on historical text
vanstriendaniel · x · 2026-09-03
Kraken PP-OCRv6, a specialist OCR model with just 15.9M parameters, ranks #4 on reading CER across a 2,165-page historical eval dataset — and #1 when long-s, ligatures, and case are preserved — outperforming general-purpose VLMs hundreds of times its size.
- Kraken is an open-source automatic text recognition system built for historical documents and non-Latin scripts
- Fully trainable layout analysis, reading order, and character recognition; RTL/BiDi/vertical script support
- Ships with CLI and Python API, ALTO/PageXML/abbyyXML/hOCR output, and a public model repository (HTRMoPo)
- The poster notes the model performed so well he had to adjust his plotting logic to display it
A strong data point for small specialist models in low-resource and long-tail digitization tasks.
Related event: Tiny 15.9M-parameter Kraken OCR model beats far larger VLMs(2 posts)→
More from Models
- China's token plans priced far above overseas LLM subscriptions, per finance mag piece — yihui_indie · 2026-09-03
- Hands-On Ranking: Fable 5.1 Beats Claude Opus on RL Environment Creation — HarveenChadha · 2026-09-03
- Philosopher says Fable 5.1 found Arrhenius's favored population ethics impossibility theorem is false — willmacaskill · 2026-09-03
- Reddit users report mass OpenAI account bans hitting paying customers — Existing-Slide7395 · 2026-09-03
- Gemini Speech-to-Text Shines in Real-World Test—Dev Open-Sources Windows Tool AivoRelay — lvvy · 2026-09-03
- Meta's new Muse model suddenly outputs Chinese characters, sparking distillation rumors — zsakib_ · 2026-09-03