Fireworks launches Specialized Intelligence Index with 12 partners to benchmark AI on real work
dr_cintas · x · 2026-09-23
Fireworks launched the Specialized Intelligence Index (SII), aiming to measure AI on real work rather than static benchmarks: built with teams running evals in production, launching with 12 partners across healthcare, legal, cybersecurity, finance, customer support, productivity, and software.
Early highlight: on DepthFirst's dfbench v1 (defensive security: vulnerability detection, validation, differential analysis), dfs-large1 — a GLM 5.2 base post-trained with Fireworks via RL — set a new Pareto frontier, with gains from RL reward shaping, an effort penalty, and a soft finding-budget penalty.
More from Models
- Opus 5.5 tops AI index at 58, undercuts GPT-6 Astra by 60% but burns 4x more tokens — johnseach · 2026-09-23
- Altman clarifies: OpenAI's new release is voice, not video, 'sorry to disappoint' — sama · 2026-09-23
- Grok 4.7, Opus 5.5, and GPT-6 Sol/Luna all shipped in one insane September week — altryne · 2026-09-23
- Leaked GPT-6-Sol testing claims half the tokens and 1/5 the time of GPT-5.6-Sol — pvncher · 2026-09-23
- OpenAI blog hints new model is similar size; caching and inference gains double intelligence per dollar — eliebakouch · 2026-09-23
- Dev calls 3D render evals 'mid': rerunning the same prompt beats any model gap — BLUECOW009 · 2026-09-23