Papers with Code adds parameter-size filtering, exposing sub-6B SWE-Bench Verified SOTA
NielsRogge · x · 2026-09-09
NielsRogge shares that Papers with Code now lets users filter benchmark entries by parameter size. A notable example: SWE-Bench Verified, the 500-sample human-validated subset testing real-world software issue solving, can now be filtered to show SOTA only among models with ≤6B parameters — handy for tracking small-model progress. The board aggregates evals from a wide range of papers spanning Claude, DeepSeek, Gemini, Kimi, MiniMax and more.
More from coding & agent
- Running a $10,000 AI model at home: Fireworks AI engineer's agentic workflow — David Ondrej · 2026-09-09
- FOSS Mac MCP server adds browser tab isolation, memory and delegated agents — bulutarkan · 2026-09-09
- Nous Research pitches Hermes Agent as the only way to truly own your AI stack — Teknium · 2026-09-09
- One Person + Codex + Apify + Apollo: The Exact Stack for AI-Powered Outbound That Isn't Spam — aryanXmahajan · 2026-09-09
- GitHub Details 4 Changes That Cut Copilot's AI Cost Without Hurting Task Quality — marlene_zw · 2026-09-09
- Mac MCP 2.0 open-sources 81 tools giving AI agents local macOS control and background browser automation — bulutarkan · 2026-09-09