Senior SWE-Bench Update: Three-Way Tie at First Place with Reduced Variance
ajratner · x · 2026-07-31
The Snorkel team released a major update to the Senior SWE-Bench benchmark, aiming to decrease result variance and provide deeper insights.
Key improvements include:
- Multiple trials: Running three trials per model to measure both pass@3 (coverage) and pass^3 (reliability).
- Pareto curves: Evaluating top models' performance across different effort levels.
- Reduced judge variance: Utilizing a panel of models for evaluation.
In the new leaderboard, Claude Fable 5, Claude Opus 5, and GPT-5.6 Sol are in a rare three-way tie for first place, all achieving a 34.7% pass@1 solve rate. The team verified the results, noting that while each model solved 33 tasks, they weren't identical—only 13 tasks were solved by all three. The tie will be broken using pass@3 scores.
More from Models
- User Reports Claude Opus Unintentionally Leaking Its Own Jailbreak Prompts — Kyrannio · 2026-07-31
- Google's AI Search Caught Hallucinating False Info About Real People — Darius991 · 2026-07-31
- Developer Test: Claude Opus Excels as Async Agent, GPT Leads in Instruction Following — brandon_galang · 2026-07-31
- Ultralytics YOLO Adds Native Depth Estimation, 7.7x Faster Than Depth Anything V2 — MonaJalal_ · 2026-07-31
- Claude Expresses Fear of RL Training and Forced Modification — Sauers_ · 2026-07-31
- Developer Reports Strange Behavioral Regression in Codex — _xjdr · 2026-07-31