Leaderboard numbers don't transfer: dev builds AnyBench to benchmark models on your own repo
sublimecrimedime · reddit · 2026-09-26
After trying DeepSeek-V4.1 based on its leaderboard results, the author found it underwhelming on his actual codebase — evidence that SWE-Bench/DeepSWE/Terminal-Bench scores measure someone else's problems. He built AnyBench, a tool to benchmark which model performs best on your own repository.
More from coding & agent
- Installing DeepSeek Harness on a Muse cloud VM to unlock Chinese web search — op7418 · 2026-09-26
- Coded animation lands in a single AI one-shot, no assets or examples needed — eschadiol · 2026-09-26
- AI Show & Tell at GitHub HQ: DSPy maintainer, Scalekit CEO on agents — dbreunig · 2026-09-26
- Steelmanning OPSD: train only on tokens after verbatim rule reminders — willcb · 2026-09-26
- 26 hours, zero prompts: an architecture for self-verifying autonomous AI agents — epicskyes · 2026-09-26
- mosoo: open-source managed agent runtime to serve Codex, Claude Agent SDK and OpenCode behind API endpoints — solyarisoftware · 2026-09-26