Seeking open-source methods to extract tables from Indian bank PDFs

OmPatel110 · reddit · 2026-08-29

A developer building a credit underwriting agent is struggling to extract structured transaction data from Indian bank statement PDFs that have wildly varying layouts. Standard tools (pdfplumber, OCR, regex, LLMs) fail on unstructured text and scanned documents. The user is seeking advice on open-source/on-premise architectures for reliable, bank-agnostic data extraction.

Original post →

More from coding & agent

coding & agent channel →