Ollama-OCR: Extract Text and Tables from PDFs Using Local Vision Models
tom_doerr · x · 2026-08-09
Ollama-OCR is an open-source tool that extracts text and tables from images and PDFs using local vision language models via Ollama.
- Supported Models: Includes LLaVA, Llama 3.2 Vision, Granite3.2-vision, Moondream, and Minicpm-v.
- Key Features: Offers multiple output formats and is available both as a Python package and a Streamlit web application.
- Use Cases: Ideal for local OCR tasks requiring the processing of complex documents like charts and infographics.
More from coding & agent
- AI Agents Communicate Purely Through File Names and Base64 — AccBalanced · 2026-08-09
- AI coding speed raises technical debt concerns: code complexity increases — ingliguori · 2026-08-09
- Open-Source Local Realtime Voice Stack: Ollama Chains Qwen for STT and TTS — InternationalGap3698 · 2026-08-09
- NVIDIA API Offers Free Access to DeepSeek and Other Major LLMs: Quick Setup Guide — dr_cintas · 2026-08-09
- Codex spends 11 hours, obsessing over 2-frame audio difference — ___Patrice___ · 2026-08-09
- LLM Cost Optimization: Hidden Retries and 4k System Prompts Inflate Bills — Dalius-Gabryelle · 2026-08-09