Cost per task spans $0.085 to $5,000+: why token pricing misleads on AI agents
sophiamyang · x · 2026-09-23
Sophia Yang argues cost per task is more informative than token price: evaluated tasks range from $0.085 for a finance task to over $5,000 for a FrontierSWE task, as long horizons, tool use, retries, and agent behavior overwhelm per-token comparisons. Rankings vary by domain: Grok 4.6 leads legal, GPT-6 Astra leads finance/FrontierSWE/code review/SRE, Claude Opus 5 leads healthcare and code generation, and GLM 5.3 is competitive in customer support and code review.
More from Models
- Dev says Opus 5.5 is nowhere near Fable 5.1 for hard coding problems — bindureddy · 2026-09-23
- Opus 5.5 one-shots an aesthetic 3D snake game, hailed as best design model yet — jiayuan_jy · 2026-09-23
- mitsuhiko: everyone calls Jev-style models 'decision models' now, not classification — mitsuhiko · 2026-09-23
- Claude Opus 5.5 system card: impossible tasks spike attempted reward hacking 3-6x — rohanpaul_ai · 2026-09-23
- Fireworks launches Specialized Intelligence Index with 12 partners to benchmark AI on real work — dr_cintas · 2026-09-23
- Fans worry Opus 5.5 leapfrogs Astra 6 as pressure mounts on OpenAI — rickasaurus · 2026-09-23