Decision Index 0.2.1: SGD benchmark pulled over bug, scoring fixes applied
multimodalart · x · 2026-09-26
Responding to a question about Decision Index 0.2.1 changes, multimodalart explains: category proportions vs. usefulness were rebalanced per feedback, several benchmark scoring bugs were fixed, and a bug in SGD requires re-running all models so it was removed ahead of v0.3. Other bugs (e.g., applying F1 to ACOS) were re-scored without re-runs.
Related event: Decision Index 0.2.1 Fixes SGD Benchmarks, Revamps Category Weights(2 posts)→
More from Models
- OpenAI's rumored persistent agent "o" may tie to old "rebranding to O" report — Dullydude · 2026-09-26
- ValsAI benchmarks reasoning levels on Proof Bench: Opus 5.5 overpriced, Astra exceeds needs — JenniferHli · 2026-09-26
- Data leak reveals Anthropic's 'Mythos' model, a 'step change' beyond Opus — Miles_Brundage · 2026-09-26
- User review: Sol is a real step up in writing and reasoning quality — dreamwieber · 2026-09-26
- 70K Responses Over 28 Days: Measuring How Volatile ChatGPT and AI Overviews Really Are — gaganghotra_ · 2026-09-26
- Developer says GPT-6-Luna is a regression: switching back to 5.6-Luna fixed everything — ElectronicAd4565 · 2026-09-26