Your Model May Be Learning Who Generated the Data — A Robotic Surgery Case
bravo_abad · x · 2026-09-30
AI-for-science writer Jorge Bravo Abad highlights an overlooked pitfall: models can pick up the habits of whoever generated the data. A model trained on scientific measurements may learn the operator's habits; if those habits also predict the label, a random train–test split makes the model look more useful than it is.
Ueki and colleagues provide a striking example in robotic surgery: analyzing 98 procedures by 16 surgeons with 3D hand-tracking data, they found exactly this kind of operator leakage.
The author explores these ideas further on Discovery at Scale, where he also writes a weekly AI for Science briefing.
Related event: Study Finds Models Can Learn the Habits of Data Generators(2 posts)→
More from Research
- Microsoft Research unveils Quine, a multimodal world model for biology — LesterMackey · 2026-10-01
- quallmer 0.5.0: R toolbox brings LLM-powered qualitative coding with reliability checks and audit trails — RexDouglass · 2026-10-01
- LLM Materials & Chemistry Hackathon expands to Open Scientific Intelligence — CatAstro_Piyush · 2026-10-01
- New philosophy paper probes the 'elusive author problem' of LLM co-authorship — SvenNyholm · 2026-10-01
- Sandia Labs partners with Radical AI's self-driving lab to discover hydrogen purification materials — CatAstro_Piyush · 2026-10-01
- Kosmos AI drafts IND filings in 4.2 hours vs 100 expert hours, cutting a drug program by 3 months — CatAstro_Piyush · 2026-10-01