Four labs shipped flagship models in one week: benchmarks and pricing compared
craigmullins · x · 2026-09-16
Memeburn compares four flagship models released in the same week of September 2026:
- GPT-6 Astra (OpenAI): leads math reasoning at 97.6% on FrontierMath and scores 100% on ExploitBench.
- Claude Fable 5.1 (Anthropic): best agentic coding at 55.8% on Terminal-Bench, strong at scientific research, cache costs 75% lower.
- Gemini 3.8 Flash (Google): cheapest at $0.75/$3.75 per million tokens but 13.3s first-token latency.
- Muse Spark 1.3 (Meta): top long-context coding at 75.4% on DeepSWE, though its best tier remains gated behind safety testing.
Premium tiers all cost $10/$50 per million tokens except Gemini, which undercuts everyone.
More from Models
- Google releases Gemini 3.8 Live and Live Extended Thinking, live in AI Studio — _philschmid · 2026-09-16
- Microsoft reportedly to limit future AI models, prompting Clippy-meme jokes — matt_slotnick · 2026-09-16
- Gemini 3.8 Live now tryable live in Google AI Studio — _philschmid · 2026-09-16
- Gemini 3.8 Live launches: #1 voice model at $0.005/min with async background thinking — _philschmid · 2026-09-16
- Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking with 97-language real-time voice — GoogleAI · 2026-09-16
- Google rolls out Gemini 3.8 Live audio models to consumers, developers and enterprises — GoogleAI · 2026-09-16