How do you benchmark recall for research agents without a gold-standard crawler?
Spirited-Cheek8436 · reddit · 2026-09-13
The author runs a challenge where agents research companies, finding people, locations, jobs and financials with sourced facts. The hard part is measuring recall: comparing against your own crawler only bounds results by what it already found. His proposed approach: pool facts from all submissions plus his own crawlers, verify and dedupe into a shared truth set, and score each agent against it — crediting facts only one agent found. The known hole is correlated misses; he considers manually researching a random sample of companies to estimate the gap but is unsure that suffices, and asks for similar evaluation experience.
More from coding & agent
- "AI won't kill me—but I'll be buried under an ever-growing review pile" — tokenbender · 2026-09-13
- ProTip: lock down code sections in agents.md to stop agents from breaking them — cyrus_zei · 2026-09-13
- Open-source thesys-core highlights exact paragraphs behind AI answers in 100+ page PDFs — Flat-Phone-1596 · 2026-09-13
- Dev uses idle Opus 5 credits to build split-screen multiplayer into his game — gandamu_ml · 2026-09-13
- freecad-mcp: 2.2k-star MCP server lets Claude Desktop drive FreeCAD for CAD and FEM — tom_doerr · 2026-09-13
- Dev fine-tunes Qwen 3.8 on 81,837 book annotations, writing quality up 86% in blind tests — Scobleizer · 2026-09-13