Blind test: Sonnet 5.5 matches Opus 5.5 on a from-scratch Rust decompressor at 1/4 the price
_Duex · reddit · 2026-09-29
The author benchmarked Sonnet 5.5 vs Opus 5.5 (plus GPT-6 Sol/Luna) on the same self-contained task: write a dependency-free, unsafe-free DEFLATE/zlib decompressor in Rust, graded blind and fully automated against zlib with 4055 hidden tests (real streams, hand-built edge cases, 4000 corrupted inputs, a speed test, a zip bomb).
- Opus 5.5: all passed, 507 MB/s, 10m18s, $2.08
- Sonnet 5.5: all passed, 447 MB/s, 3m41s, $0.52
- GPT-6 Sol: all passed, 301 MB/s, 6m37s, $0.27
- GPT-6 Luna: failed on normal zlib (one typo), $0.007
Opus earned its price on polish—zlib-style two-level Huffman tables, buffer discipline, and a self-written differential tester. For well-defined tasks, Sonnet 5.5 at medium looks like the better deal; Opus may still lead on open-ended debugging. Single task, single run per model.
More from coding & agent
- Swargs claims $50M+ transaction volume and 6,000+ listings on its AI agent marketplace — KyeGomezB · 2026-09-29
- Sonnet 5.5 effort settings make no difference in 15-task coding test: 9/15 at low, medium and high — every · 2026-09-29
- Claude Connectors: 7 essential integrations to stop copy-pasting between Claude and your apps — Aiden_Tech_Ai · 2026-09-29
- AI made building so cheap that 5 launches beat 5 months of thinking — alexmacgregor__ · 2026-09-29
- AI agent buys NBA tickets end-to-end, pivots on its own after seats sell out mid-checkout — armand_ruiz · 2026-09-29
- Embedded dev on Chinese models: DeepSeek and GLM are good enough, skip Codex/Claude — sven_ai · 2026-09-29