Web Draw: MCP server reads pages as text instead of screenshots, Amazon page ~750 tokens
ahstanin · reddit · 2026-09-03
A browser extension plus MCP server that renders the visible DOM as text with a handle on every control, so agents act by handle instead of guessing coordinates. An Amazon search page reads 750 tokens; a full eBay checkout 550. Five tools; browseract takes batched steps and returns the updated view, so form-filling is one call. The author's hardest problem was honest failure reporting: rejected submits used to look identical to successes, so the agent kept acting on a stuck screen—now refusals return the page's message and abandon remaining steps. Install via npx @olib-ai/web-draw-mcp; limits: canvas apps and div controls without ARIA roles are invisible. Free, no account, no telemetry, localhost only.
Related event: Web Draw: Open-Source MCP Reads DOM Text Instead of Screenshots(2 posts)→
More from coding & agent
- GLM-5.3-Flash priced at 0.06x: Factory confirms all-day droid usage without rate limits — matanSF · 2026-09-03
- Color grading with Imagine agent: one-shot clips, Grok grid previews and previz workflow — Kyrannio · 2026-09-03
- Agent stacks silently burn budgets: a 6-step checklist to catch runaway loops before the bill hits — Rough-Green-7067 · 2026-09-03
- Claude Code 2.1.259 adds managed MCP servers and headless permission mode — ClaudeCodeLog · 2026-09-03
- Claude Code 2.1.259 full changelog: managed MCP servers, headless permission denials — ClaudeCodeLog · 2026-09-03
- Claude Code 2.1.259 ships 37 changes: org-wide managedMcpServers, headless permission mode — ClaudeCodeLog · 2026-09-03