Sales agent hallucinated for a month — the culprit was cookie banners and stale cache

MantisReka · reddit · 2026-10-02

The author built a research agent for their sales team that scraped target companies' sites and news to qualify leads. For about a month it kept producing wrong conclusions — bad pricing, products discontinued a year earlier, even claiming a company that laid off 50% of staff in July was "hiring aggressively."

They swapped models twice, tweaked the system prompt dozens of times, and added a reflection loop (then a second one). Notes got longer and more confident but stayed wrong; at one point self-checking cost more than human research. The breakthrough came from logging raw tool responses: the model was behaving normally all along — it was reasoning over cookie banners, stale cached pages and empty React routes as if they were real content, fabricating plausible answers when given empty strings.

The fix was simple: web search migrated to contextdev (which only caches changed pages, fixing the stale careers-page issue), documents handled similarly, plus sanity checks on all tool outputs before they reach the model — empty responses, pages older than a week, or login screens are flagged to the agent. This took 3 hours and beat a month of prompt engineering. The author asks the community: should such checks live in tool filtering, or should the model handle junk data itself?

Original post →

More from coding & agent

coding & agent channel →