New Benchmark CorporateBench Tests LLMs on 230k Corporate Docs

_reachsumit · x · 2026-08-28

Sil Hamilton et al. present CorporateBench (CB), a human-validated large-scale Q&A benchmark simulating real-world enterprise communication networks. It features over 230,000 documents across four synthetic firms (12 to 10,000 employees). The data is sampled from a temporally evolving knowledge base ensuring cross-document logical consistency. Evaluations of five LLMs reveal significantly degraded performance as input size approaches realistic scales, filling a crucial gap in enterprise reasoning metrics.

Original post →

More from Research

Research channel →