Dataset signatures propagate from human data to user models to assistant evals, study finds
serinachang5 · x · 2026-10-08
Serina Chang shares research findings on user models trained and validated against human data: dataset "signatures" propagate from human-AI interaction datasets to user models and then to assistant evals, affecting conclusions throughout the pipeline.
The choice of dataset influences the user model's outputs, how user model quality is evaluated, and how a user model judges an LLM assistant. The author calls for critical thinking about what user models inherit from human datasets and what that will teach the next generation of LLM assistants.
Related event: Study: dataset signatures contaminate user models and assistant evaluations(3 posts)→
More from Research
- Kernaut: coding agents design Gaussian process kernels via program search — sirbayes · 2026-10-08
- Kernaut's discovered kernel beats tuned standard kernels on glucose prediction — sirbayes · 2026-10-08
- Kernaut's discovered kernels stay interpretable: 16 scalar functions, inner-product form — sirbayes · 2026-10-08
- iOSWorld brings computer-use agent benchmarking to iOS at COLM 2026 — kohjingyu · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- Dankrad: AI's inelegant math results reveal human bias from small context windows — CatAstro_Piyush · 2026-10-08