MedKIT Benchmark at NeurIPS 2026 Shows LLMs Can Recall Updated Facts But Fail to Use Them
zeynepakata · x · 2026-10-01
MedKIT (Medical Knowledge Integration and Transfer), a new NeurIPS 2026 benchmark from Zeynep Akata's team, evaluates whether LLMs can actually apply updated knowledge rather than just recall it. Across 12 knowledge-integration strategies on 5 models, most methods gain on lexical variation but show limited relational generalization, and none meaningfully improve compositional or operational tasks — exposing a fundamental gap between recall and usable knowledge in safety-critical medical settings.
More from Research
- How an internal 1945 EDVAC draft made von Neumann architecture famous — burny_tech · 2026-10-01
- Stanford's Kundaje calls out RNA model renaming: 'classical fine-tuning isn't a new model' — anshulkundaje · 2026-10-01
- SpikingBrain fuses linear attention with spiking neurons for zero-latency edge LLMs — gekobraa · 2026-10-01
- Ben Recht: utility maximization is inescapable—we must learn when it's misapplied — beenwrekt · 2026-10-01
- SCOPD self-distillation recovers 92% of full-context VLM accuracy with 90% fewer visual tokens — CSProfKGD · 2026-10-01
- Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro — SergioPaniego · 2026-10-01