PRISM: Language Models That Trace Next-Token Predictions to Training Data

keunwoochoi · x · 2026-10-06

Introduces PRISM (Prototype Language Models), a family of language models designed for interpretability: their next-token predictions can be traced back to specific pre-training data in a single forward pass. The work was done by Dan during his internship at Guide Labs and will be presented at COLM.

Related event: PRISM: Training Data Attribution in a Single Forward Pass(3 posts)→

Original post →

More from Research

Research channel →