ROME hits 57.4% on SWE-bench Verified with only 3B activated parameters
thisguyknowsai · x · 2026-10-06
Benchmark results for the ROME model:
- 57.4% on SWE-bench Verified
- 24.7% on Terminal-Bench 2.0
- Only 3B activated parameters (30B total, sparse activation)
The author claims this 30B model matches 480B+ models on real coding tasks, challenging assumptions about agent scaling.
More from coding & agent
- t3 code spawns agent types to review PRs with Alchemy's 20s deploys — samgoodwin89 · 2026-10-06
- Hugging Face adds an RL Environments filter to the Hub, supporting 4 frameworks — lmoroney · 2026-10-06
- Karpathy's arrival makes Anthropic's coders better, says Kuprel — Kuprel · 2026-10-06
- 'DOCX Will Be Replaced by Markdown Within 5 Years' Sparks Pushback from Formats Veteran — bytebot · 2026-10-06
- A browser tab refresh bug kept a server at 100% CPU for 4 days—load scaled with tabs squared — Ok_Negotiation_2587 · 2026-10-06
- Loop-engineering: 8 unattended agent loop patterns to run your repo while you sleep — JafarNajafov · 2026-10-06