3 GitHub tools that let your AI agent scrape data from almost any website
Roger_M_Taylor · x · 2026-09-18
A roundup of three open-source tools you can hand to an AI agent for data collection across X, YouTube, Reddit and arbitrary sites:
- Agent-Reach: bundles reading/search for X, YouTube, Reddit, GitHub, Bilibili, XiaoHongShu into one CLI with zero API fees; 82k+ stars on GitHub.
- Patchright Enhanced: a Playwright-based browser approach that listens to site requests and extracts data via script, useful when no proper API exists.
- Scrapling: a general-purpose web scraping tool.
Example use: have an agent collect posts from 100 X accounts over the past month and surface the most popular ones.
More from coding & agent
- Distributed-Slides: an MCP server that compiles agent-written talks into offline presentations — arthurcolle · 2026-09-18
- Vercel skills CLI adds Notion-hosted agent skills, no Git repo required — ivanhzhao · 2026-09-18
- Vercel teams have long written GTM and on-call skills in Notion — ivanhzhao · 2026-09-18
- eve adds automatic model selection: agents pick models by task difficulty before inference — cramforce · 2026-09-18
- DHH on Lex Fridman: Nov 2024 split coding into two universes, agentic coding is transformative — zakkohane · 2026-09-18
- Vercel CLI now deploys static artifacts in under one second, build step skipped — cramforce · 2026-09-18