Baidu’s Unlimited OCR uses R-SWA to parse dozens of pages with a 3B model

vista8 · x · 2026-07-22

A Chinese post highlights Baidu’s open-source Unlimited OCR project and says Yann LeCun has noticed it.

The project proposes Reference Sliding Window Attention (R-SWA) instead of the usual page-by-page parse-and-stitch pipeline. The idea is to keep KV cache at a constant size during decoding by borrowing how humans read and copy long documents, allowing a 3B-parameter model to continuously parse dozens of pages in one pass.

The post also claims the project quickly became a major hit after open-sourcing:

It is framed as infrastructure for large-model training-data cleaning.

Related event: Baidu's Open-Source Unlimited OCR Draws Attention for Long Document Parsing(2 posts)→

Original post →

More from Models

Models channel →