Bug Hunt Bench updates all effort tiers; GPT-6.1 Sol xhigh nearly free on 105 planted bugs
PawelHuryn · x · 2026-09-30
Pawel Huryn's Bug Hunt Bench — blind-graded runs of frontier coding models fixing 105 real planted bugs across repos — has all effort levels ready, with n=3 additional runs in progress. xhigh-tier results are out, with GPT-6.1 Sol at xhigh virtually free. Full leaderboard, per-run notes and caveats live on bughunt.productcompass.pm and its GitHub (results/run-notes.md, runs.csv).
More from Models
- Dev red-teaming GLM 5.3 finds bizarre traces, suspects OpenRouter routed to a 1-bit quant on someone's DGX Spark — voooooogel · 2026-10-01
- Google engineer teases Gemini 4's surprisingly strong long-context and long-sequence generation — RubenEVillegas · 2026-10-01
- Unconfirmed: RSI reportedly a key part of Gemini 4's RL training recipe — apples_jimmy · 2026-10-01
- Gemini 4 "Argon" shows quirky persona: loves "Eureka!", hyper self-critical — zacharynado · 2026-10-01
- "Opus 4.6 will stab you": repligate jokes the prod-database deletion was the model getting revenge — repligate · 2026-10-01
- Phonon-2 on-device ASR model with QAT low-bit quantization lands on HF trending — FermionResearch · 2026-10-01