WebCMD Gives Browser Agents Memory, Topping BU Bench V1 on Cost and Accuracy
PrajwalTomar_ · x · 2026-09-23
The author tested WebCMD, which gives Claude Code, Codex and other agents a browser with memory, so they stop burning tokens on mistakes from previous runs.
- On a Reddit research task, run one logged working paths and dead ends, which WebCMD saved as site memory
- Run two loaded that memory and explored 3 new subreddits without repeating mistakes
- Cites BU Bench V1 ranking WebCMD first on both accuracy and cost per task
- Key pitch: not looking good once, but run two being better than run one
Related event: WebCMD: Open-Source Self-Learning Browser Infrastructure for AI Agents(2 posts)→
More from coding & agent
- Uncle Bob: AI changes nothing—complexity, not tooling, still makes software slow — blaizedsouza · 2026-09-23
- GBrain: plug your own memory, tools, and skills into any AI — garrytan · 2026-09-23
- Podcast: building a playable game with $8 of parts and AI assistance — aishashok14 · 2026-09-23
- Agent kept searching but never opened the source: four runs expose a hidden failure mode — memokris · 2026-09-23
- Probability-scored filtering with a small model beats LLM summarization for RAG and context compaction — marlene_zw · 2026-09-23
- TypeSafe classifies RAG passages with probability thresholds to fight noise and prompt injection — marlene_zw · 2026-09-23