Rubric-based medical AI benchmarks fail to penalize fabricated citations, paper finds
MaziyarPanahi · x · 2026-09-14
- 新论文《When Rubrics Fail》发现:在 3 个成熟医学基准上,临床上有意义的幻觉可以完全不改变 rubric 评分。
- 作者实测 31 道 HealthBench 题目明明写了禁止伪造引用的规则,但注入的虚假编造引用仍没有被评分器扣分。
- 结论:高分可能掩盖糟糕答案,呼吁公开查看模型实际输出,并把评测的 grader 本身也纳入测试。代码与数据已公开。
More from Safety
- Joshua Saxe bets an out-of-control AI botnet appears within 18 months — joshua_saxe · 2026-09-14
- US House to vote Tuesday on act making tech firms pay data center infrastructure costs — Polymarket · 2026-09-14
- AI worm debate: Saxe argues agents could hide in $100B inference stream, Halvar pushes back on cost and capability limits — joshua_saxe · 2026-09-14
- VC claims Anthropic destroys books for LLM training, calls it 'largest knowledge loss since dark ages' — StewartalsopIII · 2026-09-14
- Saxe proposes hybrid inference botnet strategy: purchased APIs, stolen cloud keys, and local models — joshua_saxe · 2026-09-14
- Halvar Flake: 'Conflicts like an attempt to kill humanity have no zero-risk moves' — compute cost is the key constraint — halvarflake · 2026-09-14