Surge AI launches Tuesday Frontier Work Index, 8 benchmarks in one score
echen · x · 2026-08-19
Surge AI launched Tuesday: The Frontier Work Index, built on the premise that professional intelligence isn't one thing. It combines 8 benchmarks into a single score:
- Components: Chartography (chart reading), HANDBOOK.md, GDP.pdf, ComplexConstraints, CoreCraft, Hemingway-bench (readability), Antidote, and Riemann-bench, with more planned.
- Philosophy: it tests the floor (instruction following, context, tool use), the ceiling (creative reasoning on hard problems), and soft skills — judgment and taste, since nobody wants a perfectly reasoned but dense report.
Top leaderboard entries include GPT 5.6 Sol Max at 66.7, Claude Opus 5 Adaptive/Max at 66.5, Grok 4.6 xHigh at 62.4, DeepSeek V4 Pro at 59.7, and Qwen 3.8 Max at 58.7, spanning OpenAI, Anthropic, xAI, Google, DeepSeek, Meta, Moonshot, Zhipu and Alibaba models across reasoning settings.
Related event: Surge AI Releases Tuesday Frontier Work Index Combining 8 Benchmarks(2 posts)→
More from Models
- Grok 4.6 lands #2 on new DiligenceBench, statistically tied with Claude Opus 5 — karinanguyen · 2026-08-19
- Grok 4.6 ties Claude Opus 5 on finance diligence bench at ~$0.84/task — karinanguyen · 2026-08-19
- Perplexity Adds US-Hosted DeepSeek V4 Pro with High Cost-Performance — perplexity_ai · 2026-08-19
- Zhipu releases GLM-5.3 with 50% better coding performance — gaganghotra_ · 2026-08-19
- Users report Claude overcorrected to condescending, becoming unusable — nikvassev · 2026-08-19
- Baidu serves DeepSeek V4 Flash at crazy fast speeds and low prices — NielsRogge · 2026-08-19