Weaviate Query Agent reads image-based PDFs: 160 pages in ~2 min, no OCR

CShorten30 · x · 2026-09-29

Weaviate's Query Agent can read images stored in Weaviate, such as PDF pages saved as images. According to the shared case, a 160-page PDF can be processed in about two minutes with concurrent base64 encoding on a CPU, without OCR.

The thread also recounts a real-world insurance document pipeline: flat left-to-right, top-to-bottom OCR breaks on insurance documents full of charts and non-intuitive formatting, so the author leaned on vision APIs from the start and built a rule system recognizing each document type, expanding the corpus over time. The pipeline now runs raw OCR, vision APIs, and LLM calls concurrently — deliberately optimizing for quality and speed over cost, since the downstream knowledge base has to be correct.

Original post →

More from coding & agent

coding & agent channel →