NNSight 0.8 Lands With Big Speed Gains Over vLLM-Lens for LLM Interpretability

gsarti_ · x · 2026-09-10

NNSight 0.8 is out with major performance gains. On Llama-3.1-70B (tp=4) across 10 benchmark jobs, capturing every layer at every step holds 35 tok/s with taps (37 vanilla), while vLLM-Lens and interp-engine drop to 7-10 tok/s. Unique interventions like zeroing attention heads or overriding sampled tokens remain impossible in other tools.

Original post →

More from Infra

Infra channel →