$40 Benchmark Shows Qwen Matches Claude Opus at a Third of the Price

An AI engineer spent $40 running ExtractBench on 36 documents, finding that Claude Opus 5 scored around 0.94 F1 while Qwen3.8 matched it at roughly a third of the price.

2026-08-25 ~ 2026-08-25 · 2 related posts