7 open-source repos for scraping millions of web pages, from Scrapling to ScrapeGraphAI
JafarNajafov · x · 2026-09-22
A curated rundown of 7 open-source repos for large-scale web data extraction:
- Scrapling: adaptive scraper that relocates elements after layout changes, with JS rendering, proxy rotation, concurrent crawling, clean Markdown output, and an MCP server for AI agents
- Botasaurus: all-in-one Python framework covering browser automation, caching, parallel processing, proxies, and anti-detection; can package scrapers as desktop apps
- Crawlee: serious JS/TS crawling framework switching between fast HTTP requests and full browsers, with retries, queues, sessions, proxies, screenshots, and storage
- ScrapeGraphAI: describe what you want and the AI-driven pipeline scrapes it (text truncated)
Useful as a tooling shortlist for feeding data to AI agents.
More from coding & agent
- User Vibecodes Tampermonkey Scripts to Patch Long-Ignored UX Bugs in University Systems — craig_macdonald · 2026-09-22
- Speaking UI layouts into existence with Jev, shadcn and a local transcription model — _AustinCalvert_ · 2026-09-22
- ai-copywriter hits 1k stars: an agent skill that writes converting copy with zero AI tells — tom_doerr · 2026-09-22
- ContextBridge: an open-source pool that routes AI tasks across local machines, APIs and shared hosting — IamAngusU · 2026-09-22
- Microsoft's 44-page playbook: 100+ internal AI projects show licenses alone fail — alex_verem · 2026-09-22
- signal-mcp: open-source local-first MCP server exposing Signal to AI clients — googlarz · 2026-09-22