PRISM: Language Models That Trace Next-Token Predictions to Training Data
keunwoochoi · x · 2026-10-06
Introduces PRISM (Prototype Language Models), a family of language models designed for interpretability: their next-token predictions can be traced back to specific pre-training data in a single forward pass. The work was done by Dan during his internship at Guide Labs and will be presented at COLM.
Related event: PRISM: Training Data Attribution in a Single Forward Pass(3 posts)→
More from Research
- HAIPS@COLM 2026 workshop on human-centered LM privacy and security opens call for papers — tianshi_li · 2026-10-07
- CUAWright: Terminal-Only Computer-Use Agent Beats GUI Harnesses, Cuts Cost 37.5% — ysu_nlp · 2026-10-07
- AI's Top 10 papers list: Rulin Shao lands two first-author picks — ShayneRedford · 2026-10-07
- Paradigm evals its math model across 7 hard benchmarks, releases full eval suite — tensorqt · 2026-10-07
- Paradigm: post-training gains hinge on combining procedural and LLM-based synthetic data — tensorqt · 2026-10-07
- Paradigm scales RL context from 65k to 131k tokens using a trained value model — tensorqt · 2026-10-07