New Paradigm: Scaling Inherently Interpretable Language Models Without Capability Tax

guidelabs · hf · 2026-08-11

Traditional LLMs are typically trained as opaque black boxes and explained post-hoc. This work challenges that premise by making interpretability a direct constraint within the training pipeline, optimized alongside the language modeling objective.

The authors instantiate this with Steerling-8B, a diffusion language model with a causal attention mask. Across three orders of magnitude of compute, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts at scale.

Steerling-8B attributes its outputs to input tokens, human concepts, and training data. This enables closed-loop intervention—diagnosing outputs, retrieving similar training data, and correcting behavior via concept steering without retraining. It remains competitive with open peer models trained on 2-16x more compute.

Original post →

More from Models

Models channel →