Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80
PawelHuryn · x · 2026-09-10
- Bug Hunt Bench hides 105 real bugs in two production repos and lets frontier models (GPT-6, Claude, Grok, Gemini, DeepSeek, Kimi, GLM) hunt them in their own agentic CLIs, with blind grading against a withheld key.
- DeepSeek-V4.1-Flash (high): 24/105 at max effort ($1.80, 42.6 min); 19/105 at high effort ($0.31, 26.1 min).
- Comparison: Gemini 3.8 Flash (high) 20/105 at $9.78; GLM-5.3 (default) 19/105 at $19.73. deepseek-v4-pro results coming next.
- FAQ notes these are genuine tough bugs frontier models initially struggled with; unplanted bugs don't count (OpenAI models report many real but irrelevant issues); each model is judged by a different model family with verified calibration. Live leaderboard on GitHub.
Related event: Bug Hunt Bench Tests 105 Real Bugs, DeepSeek-V4.1-Flash Stands Out on Value(2 posts)→
More from Models
- DeepSeek unveils V4.1-Flash, smallest model in new family with native vision — NVIDIAAI · 2026-09-10
- Official confirmation: opted-out prompts and replies never used for training in any capacity — BlackHC · 2026-09-10
- Same Bug Benchmark: GPT-6 Astra Medium Fixes 34/105, Low Scores 27 — PawelHuryn · 2026-09-10
- ChatGPT can't stop second-guessing you, and users blame its safety training — Due-Conference-5134 · 2026-09-10
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10
- Huge share of post-2022 web data is AI content mislabeled as human-written — menhguin · 2026-09-10