Internal eval: Grok-4.7 ranks #3 on 22 knowledge-work tasks at under $5 total cost

himanshustwts · x · 2026-09-22

The author evaluated Grok-4.7 on an internal sample of 22 knowledge-work tasks spanning engineering, bizops, finance and healthcare. It ranked #3 on the table at a total cost under $5, showing strong multi-document policy reasoning and precedence handling, performing well on large structured audits across CSV, JSON, Markdown and spreadsheets, and consistently producing the requested deliverable files. The RT quips that DeepSeek V4.1 Flash still leads on this eval, crediting Colossus 2 and Cursor data.

Related event: Grok-4.7 Ranks Third Across 22 Knowledge Tasks, Costs Under $5(2 posts)→

Original post →

More from Models

Models channel →