Our RAG bot answered "accept all cookies" as competitor pricing: scraping tool shootout
Typical-Code-7006 · reddit · 2026-09-28
A developer shares a debugging story from an internal competitor-analysis RAG bot: the first version chunked raw HTML, so half the index was cookie banners and nav menus while the actual pricing tables (JS-rendered) were never fetched — the bot confidently reported "accept all cookies" as a competitor's price.
He then spent a week comparing three scraping tools:
- Jina Reader: prefix the URL and get Markdown back, working in ten minutes — great for prototyping; but JS-heavy pages come back thin, and you still find URLs and write extraction yourself.
- Context.dev: extract mode takes a domain plus a JSON schema and finds relevant pages on its own — it even pulled an SSO answer from a security page he hadn't found; costs more per call since it crawls several pages, and docs are thin.
- Firecrawl: the most polished — clean Markdown, JS pages render fine, plugs into LangChain with almost no work, plus an untested extract endpoint; the catch is credits burn faster than expected once extracting structured fields, making monthly cost fuzzy.
He asks whether anyone has run these in production for a few months and how they held up — more useful than homepage benchmarks.
More from coding & agent
- Open-sourced Claude skill finds sales leads for $0.0002 per signal, replacing $167/mo Clay plans — SimplyAnnisa · 2026-09-28
- DeepMind's Veo/Gemini Omni lead rebuilds his site with Antigravity—'it just works' — dumierhan · 2026-09-28
- Always-on agents from Anthropic and OpenAI will be big for engineering—but hard to monetize beyond it — cnakazawa · 2026-09-28
- Open-source Three.js game agent skills: 80-run tuning, 5x cheaper QA screenshots — majidmanzarpour · 2026-09-28
- three.js hits 19M weekly npm downloads, and most may be ChatGPT and Claude — cloneofsimo · 2026-09-28
- Former doomer dev: capable models now let him build '10x faster, 10x better' — bennash · 2026-09-28