PRISM: trace an LLM's outputs back to training data in a single forward pass

juliusadml · x · 2026-10-06

PRISM (Prototype Language Models) is a family of language models designed to make next-token predictions traceable to the pre-training data in a single forward pass. The work will be presented at COLM Poster Session 1, aiming to make model outputs interpretable and auditable rather than black-box generation.

Related event: PRISM Traces LLM Outputs to Training Data in a Single Forward Pass(3 posts)→

Original post →

More from Research

Research channel →