A multimodal RAG builder wants OCR that can use document context, not just images

MediocreAd3005 · reddit · 2026-07-22

The author is building a multimodal RAG pipeline where Mistral OCR annotates images before they are stored in a vector database alongside document text.

The issue is that the OCR step treats images in isolation, so the generated annotations miss surrounding document context. The author asks for:

This is essentially a practical workflow question about how to make OCR and multimodal RAG more context-aware.

Original post →

More from Apps

Apps channel →