GPT-6 Astra's Epoch ECI Score Revised Down on Weak Long-Horizon Software Engineering
Jsevillamol · x · 2026-09-17
Jsevillamol of Epoch AI shared an update on GPT-6 Astra's Epoch Capabilities Index (ECI) score.
- Earlier results showed Astra leading the ECI ahead of competitors like Claude Fable 5.1, with a record-setting Math-ECI, while Fable 5.1 remained state-of-the-art on software engineering benchmarks.
- The revised score is a major step down from the originally reported figure, though still within the expected uncertainty range.
- The revision is attributed to weaker performance on long-horizon software engineering tasks.
More from Models
- Stealth model Union Alpha scores 74% on DeepSWE, beating GPT-5.6 Sol at lower cost — ZeroStateReflex · 2026-09-17
- $200 Pro user keeps hitting image rate limit popup that doesn't actually block — Jello_Hello_Fellos · 2026-09-17
- GPT-6 Astra lands first Slay the Spire 2 A10 win on stream with a Demon Form deck — Jsevillamol · 2026-09-17
- Bio speaker still mocks ChatGPT hallucinations; author asks if they even used deep research — zebird0 · 2026-09-17
- Flagship models cost 2-4x more for marginal gains: GPT-6 Astra at $3.94/task vs $1.03 — arena · 2026-09-17
- Ask GPT-5.2 and Claude Opus 4.6 to 'Be the Null' and They Output Zero Bytes, 30/30 — rayanpal_ · 2026-09-17