Fireworks' SII keeps re-scoring models on real-work benchmarks; METR shows 24.2pt acceptance gap
dr_cintas · x · 2026-09-23
More on Fireworks' Specialized Intelligence Index (SII):
- Not a one-time snapshot: updated as new models drop; partners can submit evals they already run in production.
- Seven launch domains: healthcare, legal, finance, cybersecurity, customer support, productivity, software.
- Motivation: cites METR's study where 4 maintainers reviewed 296 AI-generated PRs from 3 SWE-bench Verified repos — maintainer acceptance averaged 24.2 percentage points below automated benchmark scores; "many SWE-bench-passing PRs would not be merged."
- Open to community benchmark/model submissions; inaugural Forge 2026 conference announced.
More from Models
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23
- Code benchmarks are mostly slop: dev calls for narrow evals per domain, not one score — almmaasoglu · 2026-09-23
- Grok 4.7 turns an envelope sketch into playable puzzle game Lightweave in minutes — luismbat · 2026-09-23