VulcanBench-SWE v4 Raises Timeout to 10 Hours to Benchmark New Coding Models Cleanly
ChrisUniverse · x · 2026-09-04
morganlinton shares methodology and early results of his VulcanBench-SWE v4 eval suite for this week's new model class:
- Real tasks: all tasks are real PRs from open-source repos, matching what normal engineering teams assign.
- 10-hour timeouts: raised from 2 hours, making runs far more expensive but ensuring failures are capability-related, not timeout artifacts.
- Anti-cheating guardrails + partial credit: models get graded credit even when not acing a task.
- Runs across every effort level; first result: Fable 5.1 Low scored 82.6 (figure truncated in post).
The takeaway: timeout settings, anti-cheat measures, and partial credit are the three key engineering details that determine coding-benchmark credibility.
More from coding & agent
- Clanker Cloud Opens Free Web Trial With $20 Worth of Starter Credits — tekbog · 2026-09-04
- Clanker Cloud Lets Anyone Build and Host Agents, With an Enterprise Sales Cautionary Tale — tekbog · 2026-09-04
- Running Grok bots like a company: AI project manager coordinates specialist agents — FinanceYF5 · 2026-09-04
- AI Finds Bugs Faster Than It Fixes Them: Engineers Grapple With CVE Backlogs — _jaydeepkarale · 2026-09-04
- Should Agents Govern Themselves? AAV Adds an External Action-Verifier Layer — CarlosMarreroAAV · 2026-09-04
- Dev spends $40 on classifier evals to cut costs: 'hard to use AI when you can't afford intelligence' — zeeg · 2026-09-04