Where Should the LLM Sit in a Scraping Pipeline? Practitioner Maps Costs, Reliability Tradeoffs
minexa_ai · reddit · 2026-10-02
A practitioner maps how scraping fits into AI workflows (RAG feeds, lead lists, market research), arguing the LLM works better on top of extraction than as the extractor.
Key points:
- Classic pipeline (render → parse DOM → CSS/XPath extraction → clean) breaks on anti-bot layers and layout changes
- Swapping extraction for an LLM removes selectors but hits context limits, grows cost with page size, and can hallucinate missing fields
- Best practices: exponential-backoff retries, keeping raw HTML snapshots, deterministic validation (regex/enums), schema-mismatch alerting, cron freshness
- The post promotes the author's Minexa.ai, a Chrome-extension-trained deterministic DOM scraper with fileurls passthrough
More from coding & agent
- Aviation's ASD-STE100 controlled language as an anti-AI-slop prompt hack, and where it fails — Paimaamu · 2026-10-02
- Pi Durable as statecharts: an interactive demo of crash-safe LLM agent harnesses — sloppenheimer · 2026-10-02
- Microsoft open-sources NVX, an ultra-light OpenVMM-based micro-VM sandbox for agentic workloads — unixterminal · 2026-10-02
- exe.dev's 'Run Fewer Agents': why task management isn't the fix for agent sprawl — charles_irl · 2026-10-02
- Building an agentic ML team: multi-agent pipeline with 40% token savings — kmeanskaran · 2026-10-02
- Claude Code creator: I don't prompt anymore, I write loops — a PM starter — aakashgupta · 2026-10-02