Building a million-page OCR pipeline with a 500GB RAM used server plus LLM extraction

oilmutt · x · 2026-10-01

An engineer shares his low-cost pipeline for digitizing millions of scanned pages: a used eBay server with 500GB+ RAM (plus hand-me-down desktops and external drives) runs traditional OCR for indexing, then LLMs extract real data to build datasets — cutting compute while boosting accuracy.

On handwritten scans, free OCR struggles badly; by his tests Claude uses 10,000 tokens per page and Grok under 3,000, with both far more accurate than free OCR. He ranks Claude first and Grok second on images.

Related event: Engineer Builds Million-Page OCR Index with Cheap Secondhand Server(2 posts)→

Original post →

More from coding & agent

coding & agent channel →