Merit CEO Brendan Foody: Astra tops enterprise evals but lags academic benchmarks like Artificial Analysis
danintheory · x · 2026-09-05
Brendan Foody says Astra ranks at the top of almost all of their enterprise evals while lagging on academic benchmarks such as Artificial Analysis. He predicts it will be the first of many models to move away from the 'myopic focus on academic evals,' signaling a decoupling between real enterprise task performance and leaderboard scores.
More from Models
- Fable 5.1 vs GPT 6 Astra on 3D Blender asset generation shows a stark gap — curious_capsuleer · 2026-09-05
- GPT 6 'Astra' reportedly recreates Pokémon from a single prompt — IanArawjo · 2026-09-05
- User switches back to GPT from Gemini after 9/10 PDF failures and heavy usage burn — DrlNoV · 2026-09-05
- GPT-6 Astra's experimental compaction in Codex saves notes across context windows, off by default — TheMoonMidas · 2026-09-05
- Same prompt run on Opus 4.8, Fable 5.1, and GPT 6 Astra for comparison — LightningMcLovin · 2026-09-05
- "Best model by far": Astra's speed lets him run 4 coding agents at once — charliermarsh · 2026-09-05