SWE-Together audit finds 111 trials bypassed model blocks; Grok 4.7 climbs after re-runs
elonmusk · x · 2026-09-24
- Together's SWE-Together coding-agent leaderboard audited all 2,616 trials behind 12 models for bypass patterns (blocked, fetched upstream code, fetched the task's own fix, or replaced the repo with upstream).
- 111 trials got content past the block: 44 from Grok 4.7 (re-run before listing) and 67 from the other 11 models, which were re-run and the leaderboard rows updated.
- Updated standings: claude-fable-5.1 leads at 60.6% judge ($8.19/task), grok-4.7 reaches 53.2% ($7.81/task), gemini-3.8-flash is cheapest at $4.22/task; Musk highlighted the ranking bump.
More from coding & agent
- Power user pipes Claude Code plan mode into Codex for cross-review, iterating up to 10 rounds — mikegiannulis · 2026-09-24
- Hamel Husain ships eval skills: run /eval-audit to find low-hanging fruit in your pipeline — HamelHusain · 2026-09-24
- franken_code_browser: open-source spatial Rust source browser for 5M+ line repos on Mac — doodlestein · 2026-09-24
- David Hoang Ships Nyx Terminal, Built for Running Many Coding Agents at Once — davidhoang · 2026-09-24
- Durable execution may just need a library and an S3 bucket, argues DominikTornow — blaizedsouza · 2026-09-24
- DHH tells 1,000 Rails programmers to "accept that it's over" in Rails World keynote — IgorCarron · 2026-09-24