$40 Benchmark Shows Qwen Matches Claude Opus at a Third of the Price
An AI engineer spent $40 running ExtractBench on 36 documents, finding that Claude Opus 5 scored around 0.94 F1 while Qwen3.8 matched it at roughly a third of the price.
2026-08-25 ~ 2026-08-25 · 2 related posts
- $40 experiment: Opus 5 hits ~0.94 F1 on ExtractBench, Qwen3.8 matches at 1/3 the price — Ok-Challenge-7810 · 2026-08-25
- Experiment: Qwen small model matches Claude at a fraction of the cost — Ok-Challenge-7810 · 2026-08-25