CER no longer fits full-page OCR: layout and style decisions all count as errors
staghado · x · 2026-09-03
staghado argues CER made sense for line/word-level eval with CRNN+CTC models — fixed vocab, no formatting, no reading order, no layout, so there was no convention surface. Full-page OCR now forces models to make dozens of stylistic decisions where the ground truth picked one particular way ten years ago, and CER treats all of them as reading errors. He concedes CER is still fine for line/word-level benchmarks.
Related event: 0.8B Qwen Fine-Tuned on Medieval Manuscripts Sparks CER Debate(4 posts)→
More from Models
- Astra solve-rate barely improves at max compute, undercutting the 'too smart to throttle' RL theory — zainhas · 2026-09-05
- Sam Altman teases next-gen OpenAI models: 'much, much, much more capable' and 'sobering for everybody' — ChrisGPT · 2026-09-05
- Astra burns tokens at high effort for no gains, finds dev testing Terminal Bench 4.0 — zainhas · 2026-09-05
- GPT 6 Astra lets you change reasoning effort mid-conversation without breaking the cache — intellectronica · 2026-09-05
- "Don't use past tense for models": users mourn Claude Opus 3 — repligate · 2026-09-05
- Astra hits 74% on DeepSWE with 30k tokens, half the cost steps of GPT-5.6 Sol — haider1 · 2026-09-05