Surge AI Launches Tuesday Frontier Work Index; Fable 5 Tops First Leaderboard
On August 19, Surge AI launched Tuesday (Frontier Work Index), a new benchmark evaluating AI models' professional work capabilities, and published its first single-score leaderboard: Fable 5 leads with 66.8, closely followed by GPT 5.6 Sol at 66.5. The benchmark attempts to measure models' overall ability to handle real workplace tasks, and is worth watching because it represents a shift in evaluation from static academic questions toward simulated real workflows.
Confirmed
- Tuesday premised on "professional intelligence isn't a single ability," combines 8 benchmarks into one score: Chartography (reading charts), HANDBOOK.md, GDP.pdf, ComplexConstraints, CoreCraft, Hemingway-bench, Antidote, and Rie (name truncated in the post), with more to be added.
- The benchmark simulates a typical Tuesday-morning workflow: reading emails, understanding the data team's charts, connecting them to last week's Slack threads, and reporting to the VP before noon.
- First leaderboard results: Fable 5 scores 66.8, GPT 5.6 Sol scores 66.5; then a clear gap emerges — DeepSeek V4 Pro at 59.7, Qwen 3.8 Max at 58.7, Gemini and others (subsequent scores truncated in the material).
- @echen added in the same thread: models need both breadth like long-horizon work and narrow skills like reading charts, plus soft skills — a perfectly reasoned but inscrutable report nobody wants to read, and inspiring output requires judgment and taste.
Why it matters
- The top two are separated by just 0.3 points, showing top models are extremely close in overall workplace capability, while the nearly 7-point gap to the second tier shows clear capability stratification.
- Tuesday synthesizes diverse abilities (breadth, narrow skills, judgment and taste) into a single comparable score, giving the industry a new cross-cutting evaluation perspective.
2026-08-19 ~ 2026-08-19 · 6 related posts
Primary sources
- [source] Surge AI Launches Tuesday Index to Benchmark AI on Real-World Professional Tasks — echen · 2026-08-19
- Tuesday Work Index leaderboard: fable 5 tops at 66.8, gpt 5.6 sol at 66.5 — echen · 2026-08-19
- [source] Tuesday Leaderboard: Fable 5 and GPT 5.6 Sol Lead the Pack — echen · 2026-08-19
- Surge: professional intelligence needs breadth, narrow skills, and taste — echen · 2026-08-19
- Tuesday index combines 8 Surge benchmarks, more coming — echen · 2026-08-19
- [source] Surge AI launches Tuesday Frontier Work Index, 8 benchmarks in one score — echen · 2026-08-19