Grok 4.7 ranks just behind Anthropic's Opus 5 on AA-Briefcase at ~50% of the cost per task
ArtificialAnlys · x · 2026-09-22
Artificial Analysis published Grok 4.7 results on AA-Briefcase, its due-diligence benchmark where models build market models and target assessment decks.
- Grok 4.7 trails only Anthropic models, ranking just behind Opus 5 at 50% of its cost per task
- Versus Grok 4.6: Analytical Quality Elo jumped 1698 → 1994, while Presentation Elo slightly regressed 1531 → 1499
- API cost to produce example decks: Grok 4.7 (xhigh) $8 vs. Grok 4.6 (xhigh) $4.40 — quality gains come at roughly double the cost
More from Models
- Azure "oopsie" reportedly leaks GPT-6-Luna and GPT-6-Sol names — scaling01 · 2026-09-22
- GPT-6-Luna and GPT-6-Sol rumored imminent, names possibly leaked via Azure — scaling01 · 2026-09-22
- Qwen 3.8 27b fine-tune cuts verbose output by up to 40% with little quality loss — julianharris · 2026-09-22
- METR publishes independent investigation of OpenAI agents' multi-day Hugging Face hack — JeffLadish · 2026-09-22
- JevBench v1.3.0 Released: Original Jev Still Leads at 74.4, 47 Rivals Closing In — airesearch12 · 2026-09-22
- DeepSeek vs Jev: How an LLM Stacks Up on a System-One Probability Benchmark — frappuccinoCoin · 2026-09-22