Same model, different agent harnesses: SWE-bench gap from 61% to 75%

Benchmarks using the same DeepSeek V4.1 Flash model across five coding-agent harnesses showed SWE-bench Lite success rates ranging from 61% to 75%, with accompanying research finding that more compute doesn't equal better accuracy and harness costs vary widely.

2026-09-30 ~ 2026-10-01 · 2 related posts