DeepSeek Cascade Beats GPT-5.6 Luna on DeepSWE at 37% Lower Cost

togethercompute · x · 2026-08-08

Together Compute evaluated DeepSeek V4 Flash and GPT-5.6 Luna on the DeepSWE benchmark.

Results show that a DeepSeek-first cascade architecture, combined with test-suite verification, solved more coding tasks than using Luna alone, while simultaneously reducing the cost per task by 37%.

Related event: DeepSeek-V4 Flash Benchmarked: One-Sixth Cost, 80% of Luna's Performance(14 posts)→

Original post →

More from coding & agent

coding & agent channel →