A 97-task benchmark built from funded AI startups: the best models fully complete about 30%

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang

cs.AI

2026-08-18

StartupBench derives 97 end-to-end tasks from 20+ market-validated AI startup products; nine frontier models top out at a 73.67 average score, and none completes more than a third of tasks under a 90-point pass line.

What problem this solves

Agent benchmarks share a soft spot: tasks are invented by researchers to probe capabilities, not drawn from work users actually pay AI to do. GAIA and OSWorld advanced realism, but the task source is still the researcher's perspective. The consequence is an unmeasured gap between leaderboard scores and "can this deliver something a user would adopt directly."

This paper inverts the pipeline: start from AI startup products already validated in the market (over USD 1M raised, paid usage or real traction), interview 30+ deep users to recover the actual workflows and deliverables, then have 50+ domain experts convert those workflows into reproducible evaluation tasks. One filter matters: tasks that frontier models breeze through in pilots are excluded, preserving discriminative power. The result is 97 tasks across six domains (Medical 21, Finance 18, Legal 16, Business 19, STEM & CS 16, Education & Humanities 7), with deliverables in real formats like DOCX, XLSX, and PPTX.

Method

Three evaluation decisions stand out:

Results

Nine models (closed: GPT-5.6-sol, GPT-5.5, Gemini-3.1-Pro, Seed-2.1-Pro, Qwen-3.6-Max; open: Kimi-K3, Kimi-K2.6, GLM-5.1, DeepSeek-V4-Pro), one harness, three runs each:

ModelAvg scoreSuccess rate (≥90)
Kimi-K373.6729.55%
GPT-5.6-sol73.6131.27%
GPT-5.572.7926.80%
Seed-2.1-Pro67.1922.34%
DeepSeek-V4-Pro61.1116.49%
GLM-5.160.7916.49%
Kimi-K2.659.9513.06%
Qwen-3.6-Max59.4615.12%
Gemini-3.1-Pro49.736.53%

Structural findings:

Why it matters

For agent product builders, this benchmark answers a commercial question: how much of the vertical-agent market can general models absorb today? The data says far from all of it. Specialized systems hold an 11.75-point edge over the oracle general setting. The moat is real, and the paper decomposes it into an actionable capability list: end-to-end compliance with complex instructions, domain compliance, and independent verification of finished artifacts. That list is more useful to R&D than any ranking.

For anyone choosing models, the 90-point pass line is worth borrowing: continuous averages make models look more capable than they are, and "delivered correctly in one shot" is the user's actual bar.

Limitations

Terms

Source

Related papers

All paper explainers