Agent-Device Benchmark: 4× More Work Per Dollar Than Alternatives

Vjeux · x · 2026-08-19

Benchmarked on AppControlBench with GPT-5.4-mini and Haiku 4.5, Agent-Device delivered:

Counter-data cited in the thread shows Haiku 4.5 achieved 98% completion (vs 84% for Agent-Device) and was 20% faster when running on the Argent framework, highlighting significant model-framework dependency.

Original post →

More from coding & agent

coding & agent channel →