Weaviate Query Agent reads image-based PDFs: 160 pages in ~2 min, no OCR
CShorten30 · x · 2026-09-29
Weaviate's Query Agent can read images stored in Weaviate, such as PDF pages saved as images. According to the shared case, a 160-page PDF can be processed in about two minutes with concurrent base64 encoding on a CPU, without OCR.
The thread also recounts a real-world insurance document pipeline: flat left-to-right, top-to-bottom OCR breaks on insurance documents full of charts and non-intuitive formatting, so the author leaned on vision APIs from the start and built a rule system recognizing each document type, expanding the corpus over time. The pipeline now runs raw OCR, vision APIs, and LLM calls concurrently — deliberately optimizing for quality and speed over cost, since the downstream knowledge base has to be correct.
More from coding & agent
- Garry Tan Backs AI Browsers Like AsideAI: Agent Password Management Is the Key Layer — garrytan · 2026-09-29
- Meeting Prep Agent Separates App State in SQLite from Long-Term Memory — Datrika_Nandhini · 2026-09-29
- Chestnut unveils 18-DOF Aero Hand plus matching exoskeleton for zero-gap humanoid data capture — chris_j_paxton · 2026-09-29
- RubyLLM hits 1.1M downloads in a month, 13M total — kieranklaassen · 2026-09-29
- Incident Memory Agent Turns Postmortems into Verified Triage Context via Hindsight — nagakeerthan · 2026-09-29
- InstaCloud launches agent-native serverless cloud that lets Claude Code and Cursor provision infra via one command — testingcatalog · 2026-09-29