Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-Gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, Yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wanzhe Liao, Chengzhi Liu, Junbo Peng, Haoran Sun, Zechen Xu, Bo Chen, Jiayi Cheng, Yi Jiang, Keying Kuang, Yuan Li, Youbang Pan, Ziyan Rao, Alexander Schubert, Yifan Shen, Vincent Siu, Xiatao Sun, Kangqi Zhang, Xiaopan Zhang, Yuchen Zhu, Ishaan Singh Chandok, Lei Ding, Jingxuan Fan, Andrew Glover, Jiaming Hu, Yiran Hu, Wenbo Huang, Zixin Jiang, Haoran Jin, Lukas Kim, Ming Liu, Yang Liu, Alireza Rafiei, Xuhuan Shen, Kunyang Sun, Sophia Sun, Ting Sun, Eric Wang, Yixin Wang, Hanwen Xing, Sihan Xu, Yuzheng Xu, Zhongxing Xu, Zhiling Yan, Boqin Yuan, Ruiqi Zhang, Yifan Zhang, Zibo Zhao, Liana, Santanu Bosu Antu, Haoyue Bai, Carlo Bosio, Joseph Cavanagh, Patricia Cavazos-Rehg, Tianxing Chen, Xuewen Chen, Yipu Chen, Chenyu Zhu, Chen Dai, Stefano De Castro, Yunfu Deng, Kaustubh Dhole, Jiayuan Ding, Chenchen Du, Zhehang Du, Hao Fan, Run-Ze Fan, Hengyu Fu, Shi Gu, Yifan Gu, Charlie Guo, Baihe Huang, Baixiang Huang, Rimika Jaiswal, Zhihan Jiang, Ran Jin, Erin Kasson, Xin Lan, Joseph Lee, Deren Lei, Chenyu Li, Daofeng Li, Haitao Li, Hongwei Li, Jingyan Li, Xiao Li, Yi Li, Yinsheng Li, Yuangang Li, Zhixu Li, Wenyu Liang, Longtai Liao, Kevin Qinghong Lin, Andy Zeyi Liu, Che Liu, Jiaming Liu, Kaiyuan Liu, Xuan Liu, Pan Lu, Wenbo Lv, Yicheng Lyu, Qiuyang Mang, Kyle Montgomery, Yuzhou Nie, Ruoxi Ning, Jorin Overwiening, Xu Pan, Layna Paraboschi, Core Francisco Park, Justin Purnomo, Swati Rajwal, Scott Rankin, Bixuan Ren, Yiren Rong, HaoYang Shang, Ventus Shaw, Fiona Shen, Jiawei Shen, Minqi Shi, Shi Qiu, Huaxiu Yao, Tianneng Shi, Jonah So, Vladislav Susoy, Hannah Szlyk, Haocheng Wang, Jialu Wang, Wei Wang, Xinyu Wang, Zehao Wang, Dowling Wong, Angela Wu, Dehao Wu, Fangyu Wu, Mengyuan "Millie" Wu, Yu Wu, Yuchen Wu, Yuhao Wu, Qingpo Wuwu, Weihang Xiao, Yongyi Xiong, Fan Xu, Ruiling Xu, Mingxuan Yan, Benjamin Yang, Jirong Yang, Sen Yang, Xiaoli Yang, Yushi Yang, Haoran Ye, Xiaohu Yu, Zhengming Yu, Chenlong Zhang, Chi Zhang, Hanning Zhang, Hanwen Zhang, Junge Zhang, Kunpeng Zhang, Song Zhang, Wenjin Zhang, Wenshuo Zhang, Ying Zhang, Yizhi Zhang, Brian Zhao, Qijian Zhao, Yimin Zhao, Yuhaohua Zheng, Liwei Zhou, Tianyue Zhou, Sichen Zhu, Siqi Zhu, Yan Zhu, Yishu Zhu, Jierui Zuo, Chonghao Cai, Helena Casademunt, Wenjia Chen, Cheng Cheng, Nawen Deng, Rao Fu, Tianfu Fu, Yifan Han, He Ren, Zhenyu He, Qiao Jin, Langlang Li, Yuetai Li, Sylvia Liu, Lu Lu, Luqing Zhou, Subhabrata Mukherjee, Yunqi Ouyang, Yin Ren, Dawei Shi, Haoran Wu, Zhiyue Wu, Hannah Yao, Zhuoran Yi, Jenny Yu, Rhea Zhan, Hang Zhou, Blake Zhu, Junfan Zhu, Alan Yuille, Yang Liu, Russell Alan Poldrack, Jiachen Li, Zhenglu Li, Molei Tao, Jing Huang, Wenqi Shi, Costas Spanos, Lichao Sun, Chenguang Wang, Orson Xu, Zhen Dong, Hector Gomez, Aylin Caliskan, Ali Emami, Haimin Hu, Zhi Li, Lihui Liu, Murphy Niu, Yi Shao, Jianxin Sun, Mikko Tolonen, Ting Wang, Sanjiv Das, Yanjun Gao, Wenbo Guo, Erika J Schneider, Zhiyong Lu, Yian Ma, Mark Mueller, Radha Poovendran, Somayeh Sojoudi, Yinglun Zhu, Dawn Song
cs.AI, cs.CL, cs.LG
2026-06-04
UC Berkeley's Agents' Last Exam (ALE): 1,490 verifiable tasks sourced from real professional workflows with 250+ experts; top agents pass under 1% on the hardest tier.
AI systems have cleared world-champion games, olympiad mathematics, and competitive programming, yet measurable change in core industries lags far behind. ALE's diagnosis is that this is an evaluation problem: widely used benchmarks do not measure whether an agent can sustain long-horizon, economically valuable work in a real professional environment. The stakes are real because the field's attention follows its benchmarks. ImageNet played exactly this role in computer vision; once a domain gets a verifiable public evaluation, progress and deployment tend to accelerate.
Three structural obstacles explain why nobody has built this yet. Authentic long-horizon workflows are expensive to collect, so existing benchmarks fall back on shorter computer-use tasks, synthetic environments, or plain QA. Broad industry coverage requires sustained access to experts, which most benchmark teams lack. Verification is genuinely hard because real deliverables are heterogeneous: a CAM toolpath, a financial workbook, a 3D mesh, a game world state. GDPval and the Remote Labor Index, the two closest predecessors, both rely on human grading, which is slow and hard to scale. The net effect is that existing benchmarks each give up one of realism, breadth, or verifiability. Mapping 16 prior benchmarks onto a shared 55-subdomain coordinate system shows their union still leaves 13 subdomains with zero coverage; GDPval covers 16/55, RLI 14/55.
Tasks are not invented; they are uploaded. Domain experts contribute projects they have already shipped, work that took them days to weeks at the time. Each submission passes five gates: targeted recruitment through an advisory committee; a submission portal with AI-assisted editing that forces five components to be fully specified (description, input files, target software, expected deliverable, evaluation spec); a conference-style first-pass review (major/minor revision, borderline, accept, strong accept); engineering implementation with containers and dry-runs; and a final expert-committee peer review that calibrates scoring bounds. The pool is 960 external submissions plus 530 commissioned tasks, 1,490 task instances total. The taxonomy is anchored in SOC 2018 / ONET, the U.S. federal occupational taxonomy, yielding 55 subdomains across 13 domains, including 4 frontier subdomains SOC has not yet catalogued.
At runtime, three components are decoupled: a task specification (an executable main.py exposing load(), start(), evaluate()), an agent (harness plus model, given only the task description and metadata), and a remote VM with a four-directory layout: read-only input/, pre-installed software/, the sole writable output/, and a hidden reference/ used only for scoring. Any agent that conforms to the action interface can run any task.
Scoring is where the design effort went. Comparison forms range from exact hashes and tolerance-bounded numeric tables to geometric distances and behavioral world states. The most common composition is gate-and-score: a binary precondition such as "no toolpath collision" or "file parses" must pass before any quality metric counts. The stated principle is to never use an LLM judge where a deterministic alternative exists, and for the minority of tasks that genuinely need visual judgment, to use narrow evidence-anchored yes/no probes instead of holistic impressions.
The evaluation target is explicit: the GCUA (Generalist Computer-Use Agent), an agent with all five capability layers (Brain, Eyes, Body, Hands, Feet). CLI agents lack Eyes; GUI agents have shallow orchestration and tool access. Mainstream evaluation uses GUI-as-Tool mode: 14 desktop actions exposed as ordinary tools through an MCP bridge, so one model reasons over shell output and screenshots in a single loop. Only about 10% of tasks (150) are public; the rest stay private and rotate periodically to resist contamination.
Main results run on the public set, capped at five hours per run, with an overall timeout rate of 3.8%.
| Configuration | Near-Term pass | Last-Exam pass | Overall |
| Codex (GPT-5.5) | 38.1% | 0.0% | 24.0% |
| ALE-Claw (GPT-5.5) | 32.8% | 2.6% | 23.0% |
| Claude Code (Fable 5) | 34.3% | 0.0% | 22.0% |
| Claude Code (Opus 4.7) | 20.9% | 0.0% | 13.2% |
The reference point is Terminal-Bench: Codex with GPT-5.5 scores 82% there but only 23.3% overall on ALE-CLI, the 105-task Linux-only subset, and 0% on its hardest tier. Across mainstream configurations the average pass rate on the hardest tier is below 1%, which is where the name comes from. The three tiers serve different purposes: Near-Term (67 tasks) is where frontier agents reach 38.1% and leaderboards can actually move; Full-Spectrum (55 tasks) guarantees at least one instance per subdomain; Last-Exam (38 tasks) is where most agents sit at 0%. The team's stripped-down reference harness ALE-Claw, built from basic GCUA components, lands within a point of the mainstream harnesses.
The failure analysis is the most informative part. For failed Claude Code + Opus 4.7 runs, Understanding and Approach failures together account for roughly three quarters, so the bottleneck is domain knowledge, not execution. Agents lacking specialized knowledge bypass the intended professional software and improvise ad-hoc scripts. Tool traces confirm it: 34% of tasks designate GUI software as the primary tool, yet GUI calls stay a small share across most configurations as agents substitute Bash. Domain spread is wide, with computational mathematics and agriculture/environment at 55–85% and education below 25%. Decomposing performance variance, the choice of foundation model accounts for roughly 3× the spread of the choice of harness. A single run of one frontier agent on one task costs $3–10 and tens of minutes to hours.
For agent builders the signal is blunt: swapping among well-engineered harnesses buys less than swapping models, and since most failures trace to missing domain knowledge, further harness polish has limited headroom. On the evaluation side, ALE is the first benchmark covering all 55 SOC/ONET digital industries with fully automated deterministic scoring, sidestepping the human-grading bottleneck that caps GDPval. The three-tier design is pragmatic about evaluation budgets and gives different goals different entry points.
Be clear about what this is: a benchmark-construction paper. The contribution is the sourcing and verification protocol, not algorithmic novelty. Whether it becomes the next Terminal-Bench depends on community adoption and sustained task-pool maintenance, neither of which is demonstrated yet.
Stated by the paper: evaluation is expensive ($3–10 per task per run), so tasks are tiered and main results use the public subset; only some configurations got three repeated runs, so variance estimates are incomplete; 323 tasks are still pending QC.
Reading closely raises more. Ninety percent of tasks are private, so outsiders cannot audit task quality or scoring implementations, and cross-time comparability rests on the organizing team's execution of the rotation. Mapping other benchmarks onto the 55 subdomains used an LLM-assisted classifier, and the ONET screening used GPT-4o mini; no full manual audit of classification error is described. Calibration of scoring bounds is expert discretion. Main results cover only the 150 public tasks, with representativeness of the full pool argued in an appendix. The five-hour cap means genuinely long workflows can be truncated; OpenClaw already times out on 5.7% of runs.