Web Draw: MCP server reads pages as text instead of screenshots, Amazon page ~750 tokens

ahstanin · reddit · 2026-09-03

A browser extension plus MCP server that renders the visible DOM as text with a handle on every control, so agents act by handle instead of guessing coordinates. An Amazon search page reads 750 tokens; a full eBay checkout 550. Five tools; browseract takes batched steps and returns the updated view, so form-filling is one call. The author's hardest problem was honest failure reporting: rejected submits used to look identical to successes, so the agent kept acting on a stuck screen—now refusals return the page's message and abandon remaining steps. Install via npx @olib-ai/web-draw-mcp; limits: canvas apps and div controls without ARIA roles are invisible. Free, no account, no telemetry, localhost only.

Related event: Web Draw: Open-Source MCP Reads DOM Text Instead of Screenshots(2 posts)→

Original post →

More from coding & agent

coding & agent channel →