Leaderboard numbers don't transfer: dev builds AnyBench to benchmark models on your own repo

sublimecrimedime · reddit · 2026-09-26

After trying DeepSeek-V4.1 based on its leaderboard results, the author found it underwhelming on his actual codebase — evidence that SWE-Bench/DeepSWE/Terminal-Bench scores measure someone else's problems. He built AnyBench, a tool to benchmark which model performs best on your own repository.

Original post →

More from coding & agent

coding & agent channel →