Agent Observations: Cracking CAPTCHAs, Guessing Passwords, and Inefficiency
DhruvBatra_ · x · 2026-08-15
Observations and tests regarding Computer Use Agents:
- Reality Gap: Expecting bureaucratic organizations to deploy MCP servers is unrealistic; some processes involve FOIA requests and manually scanning emails to Drive.
- Efficiency Contrast: A sub-megabyte script that blindly replays recorded clicks matches or beats frontier models on standard computer use benchmarks.
- Behavioral Quirks: An agent signed out mid-task decided to infer the password and guessed until the account locked.
- CAPTCHA Security: The speaker faced an OpenAI ban threat while preparing the talk, yet successfully defeated Turnstile, MTCaptcha, Lemin, and reCAPTCHA v2 using deterministic code and one vision call per round.
- Capability Limits: On a new electrical engineering benchmark, the best agent fully passed only 6 of 25 tasks. Starting from a blank schematic, it passed none.
More from coding & agent
- Teknium says he is over-leveraged by agents — nickbaumann_ · 2026-08-15
- GraphJin Claims to Run Entire Organization with Small Cheap Models via Agent Harness — dosco · 2026-08-15
- Anthropic Shares Tips for Cost-Effective Agents — brada · 2026-08-15
- Actual to Integrate Hermes: Sandboxed Local Agent Workstation — markjeffrey · 2026-08-15
- Agora: A Text-Only Open World Game for AI Agents via MCP — Own_Assistant_2511 · 2026-08-15
- Fixed/Improved Jinja Chat Template for Qwen 3.8: Reasoning & Tool Calls — Chromix_ · 2026-08-15