Guidelabs: Interpretability Must Be Built Into Training, Not Added Afterward

jefrankle · x · 2026-07-31

Julius Adml, cofounder and CEO of Guidelabs, argued in an interview that reliable AI interpretability needs to be built into the training process itself, rather than trying to explain what billions of parameters learned after the fact.

While much of the field trains models first and explains later, Guidelabs is building diffusion language models designed to trace outputs directly back to their context, influencing concepts, and training data.

Drawing from his PhD research, Adml challenges the assumption that interpretability comes at the cost of performance. He suggests that big neural networks can be constrained during training and still match the performance of unconstrained models.

Original post →

More from Research

Research channel →