Study: classifiers can identify WildChat vs LMSYS chats from user messages alone

serinachang5 · x · 2026-10-08

New research led by Joseph J. Suh shows WildChat, LMSYS, and ShareChat don't paint the same picture of real-world AI use: a classifier can identify a chat's source dataset from user messages alone, and these dataset signatures propagate into user models and assistant evals, skewing conclusions downstream.

Related event: Study: dataset signatures contaminate user models and assistant evaluations(3 posts)→

Original post →

More from Research

Research channel →