Senior SWE-Bench Update: Three-Way Tie at First Place with Reduced Variance

ajratner · x · 2026-07-31

The Snorkel team released a major update to the Senior SWE-Bench benchmark, aiming to decrease result variance and provide deeper insights.

Key improvements include:

In the new leaderboard, Claude Fable 5, Claude Opus 5, and GPT-5.6 Sol are in a rare three-way tie for first place, all achieving a 34.7% pass@1 solve rate. The team verified the results, noting that while each model solved 33 tasks, they weren't identical—only 13 tasks were solved by all three. The tie will be broken using pass@3 scores.

Original post →

More from Models

Models channel →