Baidu’s Unlimited OCR uses R-SWA to parse dozens of pages with a 3B model
vista8 · x · 2026-07-22
A Chinese post highlights Baidu’s open-source Unlimited OCR project and says Yann LeCun has noticed it.
The project proposes Reference Sliding Window Attention (R-SWA) instead of the usual page-by-page parse-and-stitch pipeline. The idea is to keep KV cache at a constant size during decoding by borrowing how humans read and copy long documents, allowing a 3B-parameter model to continuously parse dozens of pages in one pass.
The post also claims the project quickly became a major hit after open-sourcing:
- GitHub Trending and Hugging Face download charts went to No. 1 on launch day
- GitHub stars have passed 16.5K
- Hugging Face downloads are over 2.24M
- it has returned to the trending charts and is now ranked second
It is framed as infrastructure for large-model training-data cleaning.
Related event: Baidu's Open-Source Unlimited OCR Draws Attention for Long Document Parsing(2 posts)→
More from Models
- Hugging Face reportedly used open-weight GLM 5.2 after proprietary models failed — rasbt · 2026-07-22
- Google Exec Seeks Feedback on Gemini 3.6 Flash & 3.5 Flash-Lite Performance — patloeber · 2026-07-22
- Gemini 3.6 Flash is 2x faster and 18% cheaper, but independent tests say it is not smarter — etherd0t · 2026-07-22
- Critic says OpenAI incident coverage confuses bad reward functions with autonomy — ambaonadventure · 2026-07-22
- Google Launches Gemini 3.5 Flash Cyber Model for Security Teams — pushmeet · 2026-07-22
- Rumor says GPT-5.6 Sol could hit 750 tok/s after Cerebras upgrades — haider1 · 2026-07-22