Internal eval: Grok-4.7 ranks #3 on 22 knowledge-work tasks at under $5 total cost
himanshustwts · x · 2026-09-22
The author evaluated Grok-4.7 on an internal sample of 22 knowledge-work tasks spanning engineering, bizops, finance and healthcare. It ranked #3 on the table at a total cost under $5, showing strong multi-document policy reasoning and precedence handling, performing well on large structured audits across CSV, JSON, Markdown and spreadsheets, and consistently producing the requested deliverable files. The RT quips that DeepSeek V4.1 Flash still leads on this eval, crediting Colossus 2 and Cursor data.
Related event: Grok-4.7 Ranks Third Across 22 Knowledge Tasks, Costs Under $5(2 posts)→
More from Models
- Xiaomi's MiMo-V2.6 lands: 1.02T-param MoE takes top open model spot on Artificial Analysis — _AndrewZhao · 2026-09-22
- Merge Gateway lists Grok 4.7 from xAI, citing coding gains on CursorBench and DeepSWE — shensi · 2026-09-22
- SemiAnalysis says open source is dying, yet 20+ open models shipped in the past month — _lewtun · 2026-09-22
- Grok 4.7 posts 59% recall on defensive cyber bench at half the cost of rivals — andreamichi · 2026-09-22
- Grok 4.7 jumps from #9 to #3 on BuildingBench with 0.783, 66% cheaper than Fable 5.1 — ZhitingHu · 2026-09-22
- Jev explained: why the AI community's new favorite isn't a traditional LLM — multiply_matrix · 2026-09-22