Harvard/Stanford paper proposes a 4-field reporting standard for knowledge-work benchmarks
Shahules786 · x · 2026-09-03
Thread (5/6): "Designing Benchmarks for Knowledge Work" proposes a clean reporting standard with four fields separating what a benchmark represents from what its metric can prove: represented activity, tested setting (tools, materials, role, authority, workflow state), required work product, and evaluated result.
More from Research
- Revera: Lean-verified POSIX regex engine with identical output in 6 languages — jedisct1 · 2026-09-03
- ChatGPT helps resolve 30-year-old stable forking conjecture in logic — Dr_Singularity · 2026-09-03
- Yale trio's new paper: mechanism design for AI agents with unknown alignment — Afinetheorem · 2026-09-03
- PufferLib 5.0 Self-Play Trains 10-Ship Duel in 5 Minutes on a $700 PC — yacineMTB · 2026-09-03
- Until Labs scales cryoprotectant discovery to 250,000+ candidate molecules with AI — lukaszkaiser · 2026-09-03
- Stanford's Michael Bernstein builds a "What-If Machine" for simulating decisions with AI — msbernst · 2026-09-03