Medical trial example: most frontier LLMs fail at flipped-eligibility claim check
deliprao · x · 2026-10-06
Deliprao demonstrates his COLM26 finding with a concrete case: a medical trial evidence document ends with "trial is open only to Japanese women". Keeping all details and flipping only the eligibility to "regardless of ethnicity" makes most frontier LLMs wrongly accept the claim. Benchmarks make the broken part always salient, so the shortcut scores well — and models trained on them generalize it badly in deployment.
More from Research
- PerturBot Breaks Shortcut Priors in Vision-Language-Action Models With Perturbative Training — Mingyu Liu · 2026-10-06
- Georgia Tech's TextReg Fixes Prompt Distributional Overfitting, Gains Up to +11.8% OOD — GeorgiaTech · 2026-10-06
- SourceLearn Builds Source-Specific Agent Competence, Wins 13 of 15 Benchmarks — GeorgiaTech · 2026-10-06
- Representation-Space MMD Post-Training Boosts Diffusion LMs, More Parallel Decoding at 16B — yresearch · 2026-10-06
- Attention Relay makes embedding models instruction-aware without training via LLM attention weights — _reachsumit · 2026-10-06
- Programmatic Search Agents boost task success by up to 7.56 points over query-based agents — _reachsumit · 2026-10-06