Notion Launches Knowledge Board: Evaluating LLMs on Real-World Traffic Instead of Benchmarks
ivanhzhao · x · 2026-08-14
Notion has introduced the Knowledge Board, a model evaluation dashboard. Unlike traditional benchmarks using fixed tasks, it uses a small, random, and anonymized sample of live platform traffic, split evenly across models to handle actual knowledge work.
Developed in collaboration with researchers from Anthropic and OpenAI, the system uses an ensemble judge to measure resolution rate, time, and cost. Because many model differences fall within confidence intervals, Notion deliberately avoids creating a traditional leaderboard, instead plotting quality against cost to help users find the right model for their specific needs.
More from Models
- Anthropic's Unreleased Opus 5 Excels at Rust Optimization but Makes Stupid Mistakes — mertdumenci · 2026-08-14
- Anthropic's Suspicious Silence Hints at Upcoming 'Claudette' Drop — beffjezos · 2026-08-14
- DeepSeek V4-Pro slows to crawl on long projects; reloading reveals it's further ahead — teortaxesTex · 2026-08-14
- Google Focuses on Smaller, Faster AI Models Leveraging Massive Search Scale — haider1 · 2026-08-14
- DeepSeek V4 Pro Hits Baseten APIs: 1.7T Parameters, MIT License — baseten · 2026-08-14
- Grok Blocks Intimate Likeness Edits, Proving Musk's 'Legal = Allowed' Wrong — firasd · 2026-08-14