OfficeVal Benchmark: LLMs Cheaper Than Humans on Office Tasks, But Lag in Quality
Jingbo Zhou · hf · 2026-07-30
Researchers introduced OmegaUse-OfficeVal, a benchmark designed to evaluate LLM agents on long-horizon office-suite tasks with task-level economic grounding.
- Task Composition: The benchmark consists of 100 tasks derived from real practitioner requests, requiring an average of 2.32 hours of human labor to complete.
- Economic Signals: Each task is paired with human labor time and a task price proxy. These signals enable direct comparisons between LLM inference costs and human costs, as well as value-weighted evaluations.
- Findings: Although all evaluated frontier LLMs are substantially cheaper and faster than human workers, their deliverable quality has not yet approached human-level performance.
More from Research
- GuidedRAG: Semantic Steering Improves Retrieval Precision and Slashes RAG Overhead — _reachsumit · 2026-07-30
- Yandex Study: ID Embeddings Outperform GNN Embeddings in Large-Scale Recommenders — _reachsumit · 2026-07-30
- UCSB & LinkedIn Research: Agents Can Speculate Their Own Tool Calls — dair_ai · 2026-07-30
- evalstats: Open-Source Tool for LLM Judge Bias-Corrected Stats Tests — IanArawjo · 2026-07-30
- Kuaishou's DIRECTOR Framework Uses Optimal Transport to Cut RecSys Compute by 66.7% — _reachsumit · 2026-07-30
- Kuaishao's PSG: Generative Reranking Decodes Ordered Item Pairs, Halving Steps Losslessly — _reachsumit · 2026-07-30