DeepSeek's Parallel Attempts Catch Up: pass@2 Beats Luna pass@1 at $0.20
zainhas · x · 2026-08-07
While Luna has higher single-shot quality (67.2% vs 53.3%), DeepSeek's low cost enables parallel attempts: 2 shots reach 81.6% vs 70.1%, 4 shots 90.3% vs 80.5%. For DeepSWE, DeepSeek pass@2 exceeds Luna pass@1 at $0.20 vs $0.61.
Related event: DeepSeek-V4 Flash vs GPT-5.6 Luna: Cost-Effective, 80% Quality(11 posts)→
More from coding & agent
- Jeff Dean Shares 1-Hour AI Engineering Lecture: From LLMs to Coordinating 100 Agents — colinmcnamara · 2026-08-07
- Factorio Learning Environment v0.3.0: The Ultimate AGI Eval? — jwt0625 · 2026-08-07
- Inside Claude Science: Auto-Context Compression and Background Self-Correction — josephdviviano · 2026-08-07
- Core AI Engineering Pain Point: Is It a Model Problem or a Harness Problem? — zainhas · 2026-08-07
- Orbit: Manage Multi-Agent Collaboration via Local Markdown Files — khanhhuy_1998 · 2026-08-07
- Hermes Agent Combined with MCP Enables Desktop Control and Automation of Android Phones — Teknium · 2026-08-07