How to Reliably Scrape Academic Paper Data
Aromatic-Ad-6711 · reddit · 2026-07-19
After getting their IP blocked by Google Scholar for excessive requests while trying to collect papers for a university research institute, the author asked how to build a more stable and compliant scraping system.
Listed approaches include:
- Proper rate limiting and caching
- Exponential backoff and resuming from breakpoints
- Manual CAPTCHA handling
- Switching to alternative sources like OpenAlex, Crossref, ORCID, and Semantic Scholar
The core question is how to build a long-running academic literature collection agent without violating platform rules or triggering further bans.
More from coding & agent
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- ty adds first-class Pydantic support, including strict and lax field handling — charliermarsh · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22