Agent-Device Benchmark: 4× More Work Per Dollar Than Alternatives
Vjeux · x · 2026-08-19
Benchmarked on AppControlBench with GPT-5.4-mini and Haiku 4.5, Agent-Device delivered:
- 4× more completed tasks per dollar
- 94.2–98.3% task completion rate
- +2% speed with GPT-5.4-mini; -26% with Haiku 4.5
Counter-data cited in the thread shows Haiku 4.5 achieved 98% completion (vs 84% for Agent-Device) and was 20% faster when running on the Argent framework, highlighting significant model-framework dependency.
More from coding & agent
- Software engineering fundamentals matter more than ever in the agentic era — tokenbender · 2026-08-19
- Frontier benchmark verifier expects output fields the agent can't even infer — dejavucoder · 2026-08-19
- BuilderIO open sources 'skills' library for coding agents with visual planning and context recovery — tom_doerr · 2026-08-19
- A Reddit dev is validating an open-source, self-hostable proxy that sits between MCP clients and servers, handling OAuth flows, token expiration, silent refresh, and provider-specific quirks via Docker. Target users are Cursor/Claude/Copilot/custom MCP clients tired of implementing refresh logic themselves; the author explicitly wants real pain points over "yes I'd use it." — Valuable-Ticket-6879 · 2026-08-19
- AI SDK introduces Code Mode for efficient programmatic tool calling — lgrammel · 2026-08-19
- MiniMax releases MCode CLI: Deep terminal integration for AI coding agents — PrajwalTomar_ · 2026-08-19