Same model, different agent harnesses: SWE-bench gap from 61% to 75%
Benchmarks using the same DeepSeek V4.1 Flash model across five coding-agent harnesses showed SWE-bench Lite success rates ranging from 61% to 75%, with accompanying research finding that more compute doesn't equal better accuracy and harness costs vary widely.
2026-09-30 ~ 2026-10-01 · 2 related posts
- Swing Coding-Agent Harnesses and Success Jumps 61% to 75% on SWE-bench Lite — shensi · 2026-09-30
- New study: more compute doesn't mean better accuracy across agent harnesses — xiye_nlp · 2026-10-01