NYU study: incidental chatter contaminates 35% of LLM clinical notes and disrupts reasoning
NYU-OLAB · hf · 2026-10-10
NYU OLAB researchers quantified how sensitive LLMs are to incidental information in clinical settings.
- Across 576 real patient-clinician dialogues, frontier models inserted small-talk into 35% of notes, though mean quality scores shifted by at most 0.20 on five-point scales
- In 3.7% of frontier notes, models misattributed the asides or used them clinically
- In 57 mock consultations, background speech from another encounter at -10 dB leaked into 48.2% of transcripts; contamination was detected in 5.3% of notes from four open-weight models
- The authors propose a "dual encoding" hypothesis: the LLM components disrupted by incidental information are also those supporting clinical reasoning
They recommend evaluating resistance to incidental information before clinical deployment, with safeguards that block contamination without degrading reasoning.
More from Safety
- LiveOverflow asks why UUID-as-API-key is standard practice but UUID-as-ID is IDOR — rez0__ · 2026-10-10
- Anthropic accused of calling RSP 'commitments' while dodging legal binding force — Miles_Brundage · 2026-10-10
- Insider says real-time AI mass surveillance and profiling is already here — and it's just the tip — Graham_dePenros · 2026-10-10
- 16-Year-Old's Bug Report: 17 Trillion Microsoft Records Exposed, $5k Bounty — rez0__ · 2026-10-10
- Lawyer: OpenAI could legally disclose why it fired 3 safety researchers, 'trust me' isn't required — GarrisonLovely · 2026-10-10
- Paper decodes 315K encrypted reasoning blocks, recovers 367 PII and 182 credentials — DynamicWebPaige · 2026-10-10