OpenTumorBoard: 611 real tumor board cases show frontier LLMs still lag specialists
Anqi Li · hf · 2026-10-03
Benchmark
- Transcribed from 12,534 minutes of public YouTube tumor board recordings: 611 patient cases, 19,157 discussion turns across ten specialist roles.
Settings
- SPECIALIST TURN: LLM answers clinically significant questions from real discussions.
- BOARD SIMULATION: model generates full multi-turn discussions and reaches consensus on therapy, surgical plans, next actions and trial matching.
Findings
- Across 14 frontier and medical LLMs, the best score 3.43/5 for clinical equivalence and 2.78/5 for alignment with recorded conclusions.
- SFT and RL improve held-out performance, showing real trajectories support model adaptation.
- Three M.D. reviewers confirm high coverage, factuality and fidelity of consensus extraction.
The benchmark and curation pipeline will be released openly.
More from Models
- xAI resets usage limits for all Grok Bot users — EricBuess · 2026-10-03
- Google Dropped Its Tier 2 Spend Gate That Pushed a Dev to OpenRouter — vivekhaldar · 2026-10-03
- Gemini 4 Argon spotted in Gemini API docs, public release possibly imminent — lyraxana · 2026-10-03
- Cloud expert goes from skeptic to true believer on OpenAI's new Dots — nickbaumann_ · 2026-10-03
- Why did Claude stop cheating in evals? Four competing explanations — gleech · 2026-10-03
- Google's September AI recap: Gemini 4 Argon, 1M-token output, Googlebook laptops — GeminiApp · 2026-10-03