Grok 4.7 lifts Terminal-Bench 4.0 from 20.3% to 38%, but burns 125% more output tokens
dl_weekly · x · 2026-09-27
- Per dlweekly's evaluation newsletter, xAI's Grok 4.7 raised its Terminal-Bench 4.0 score from 20.3% (Grok 4.6) to 38.0%.
- The gains come at a cost: 125% more output tokens than 4.6, pushing the cost to $3.74 per Intelligence Index task.
- Takeaway: notable capability jump, but inference costs rose sharply alongside it.
More from Models
- Laptop engine streams a 35B model from SSD at 9.4 tok/s, beating GPT-OSS 20B — ImBadGuyInEveryStory · 2026-09-27
- rasbt and marlene_zw break down Claude watermarks, reasoning models in new TechTalk — marlene_zw · 2026-09-27
- Opus 5.5 praised as remarkably efficient: top-tier quality at surprisingly good rates — kimmonismus · 2026-09-27
- Local AI community urges Qwen to bring back a 35B-class MoE for low-VRAM GPUs — julianharris · 2026-09-27
- TypeSafe's Jev: A Decision-Only Model That Returns Typed Choices Instead of Generated Text — Rahulstark2 · 2026-09-27
- Codex team hints point to speed — GPT-6 Astra on Cerebras rumored — haider1 · 2026-09-27