Seeking open-source methods to extract tables from Indian bank PDFs
OmPatel110 · reddit · 2026-08-29
A developer building a credit underwriting agent is struggling to extract structured transaction data from Indian bank statement PDFs that have wildly varying layouts. Standard tools (pdfplumber, OCR, regex, LLMs) fail on unstructured text and scanned documents. The user is seeking advice on open-source/on-premise architectures for reliable, bank-agnostic data extraction.
More from coding & agent
- MCP Usage Explodes: Vercel Tool Calls Up 564% in 3 Months — jasonkneen · 2026-08-29
- From Chatbots to Doing Work: X17z Demonstrates Terminal Operators — Scobleizer · 2026-08-29
- Agents on Omarchy create tools for themselves, including a task workbench — BLUECOW009 · 2026-08-29
- 'Abundant Constraints Beat Abundant Implementation': An Essay on Directing AI Capability — aishashok14 · 2026-08-29
- fbtee 4.0 released, fully rewritten in Rust with Oxc — cnakazawa · 2026-08-29
- Together AI: cascading GLM-5.3 Flash to GLM-5.3 cuts cost 57% while boosting DeepSWE to 80.9% — togethercompute · 2026-08-29