LLM extraction silently dropped listings at chunk boundaries; 2K overlap and two-step dedup fixed it
InsideDebt6345 · reddit · 2026-10-06
While building a lead-generation agent with the OpenAI Agents SDK and ZenRows, the author hit a silent data-loss bug: a 543K-character Clutch directory page was extracted in 40K-character chunks, and the prompt's instruction to skip partial listings caused records split across chunk boundaries to be skipped by both chunks.
The fix and lessons:
- Add a 2K-character overlap between chunks so boundary listings appear complete at least once
- Overlap creates duplicates, so dedup needs two checks: normalize URLs (www/trailing slash variants) before comparing, and fall back to company name when the website field is empty—otherwise all empty-website records collapse into one
- Unwrapping the directory's tracking URLs cut 86K characters and stopped the model from treating tracking links as company sites
The pipeline extracted 78 companies from one page with none missing, though counts still vary run to run.
More from coding & agent
- Matt Pocock launches The AI Coding Dictionary to standardize AI coding terms like harness, spec and cache tokens — mattpocockuk · 2026-10-06
- These agents ate 780GB of disk space in 3 weeks — kevinkern · 2026-10-06
- Google's AIM paper: research agents improve faster by mapping and auditing ideas, beating baselines up to 3.1x sooner — rohanpaul_ai · 2026-10-06
- banteg shares how agents help him coordinate a 150-PR cross-project effort — banteg · 2026-10-06
- Most Lawyers Use AI Legal Tools Only at Basic One-Shot Prompt Level, Says Attorney — jkubicki · 2026-10-06
- GPT-2 × Codex builds 'alien eye' tracker web app that admits when it can't see — MikePFrank · 2026-10-06