New Benchmark CorporateBench Tests LLMs on 230k Corporate Docs
_reachsumit · x · 2026-08-28
Sil Hamilton et al. present CorporateBench (CB), a human-validated large-scale Q&A benchmark simulating real-world enterprise communication networks. It features over 230,000 documents across four synthetic firms (12 to 10,000 employees). The data is sampled from a temporally evolving knowledge base ensuring cross-document logical consistency. Evaluations of five LLMs reveal significantly degraded performance as input size approaches realistic scales, filling a crucial gap in enterprise reasoning metrics.
More from Research
- Miles now supports RL training for Qwen and GLM models with high-performance kernels — ying11231 · 2026-08-28
- Miles-diffusion introduces LoRA SFT for fast post-training of diffusion models — ying11231 · 2026-08-28
- SovietRxiv adds 7,000 translated Soviet scientific papers to archive — generativist · 2026-08-28
- 87% of "Quantum Supremacy" Claims Fail Under Real-World Testing, Physicist Says — AryHHAry · 2026-08-28
- FP-AMB: a first-person agent memory benchmark that tells you why each miss happened — LowDistribution3995 · 2026-08-28
- Penn & Yale Paper: Conformal Prediction Calibrations Diverge — ReCal Makes Them Reproducible — burkov · 2026-08-28