Guidelabs: Interpretability Must Be Built Into Training, Not Added Afterward
jefrankle · x · 2026-07-31
Julius Adml, cofounder and CEO of Guidelabs, argued in an interview that reliable AI interpretability needs to be built into the training process itself, rather than trying to explain what billions of parameters learned after the fact.
While much of the field trains models first and explains later, Guidelabs is building diffusion language models designed to trace outputs directly back to their context, influencing concepts, and training data.
Drawing from his PhD research, Adml challenges the assumption that interpretability comes at the cost of performance. He suggests that big neural networks can be constrained during training and still match the performance of unconstrained models.
More from Research
- Percy Liang's Simile: Building Foundation Models to Predict Human Behavior — RishiBommasani · 2026-07-31
- SpecFirst Framework: Agents Write Specs First, Boosting Code Synthesis by 21% — centre-for-swe · 2026-07-31
- Study Shows Claude 3 Opus Attention Mechanism is Turing Complete — ctjlewis · 2026-07-31
- Physics-Based Data Augmentation for Quantum State Classification — bravo_abad · 2026-07-31
- OpenAI Responds to Benchmark Concerns: GDPval Nearing Saturation — emollick · 2026-07-31
- Tactile Sensing Emerges as Robotics Frontier with Origami Challenge — chris_j_paxton · 2026-07-31