NNSight 0.8 Lands With Big Speed Gains Over vLLM-Lens for LLM Interpretability
gsarti_ · x · 2026-09-10
NNSight 0.8 is out with major performance gains. On Llama-3.1-70B (tp=4) across 10 benchmark jobs, capturing every layer at every step holds 35 tok/s with taps (37 vanilla), while vLLM-Lens and interp-engine drop to 7-10 tok/s. Unique interventions like zeroing attention heads or overriding sampled tokens remain impossible in other tools.
More from Infra
- Huawei hikes Ascend 950DT price 20-50% as black-market HBM costs multiply — teortaxesTex · 2026-09-10
- Google signs 22-year deal for half a Finnish nuclear plant's output in €13B AI push — Servola-Journal · 2026-09-10
- GPU scarcity's real cost: compute sold as 3-year blocks starting months out — kevinakwok · 2026-09-10
- West Bengal mulls opening 18 industrial parks for data centers as India capacity heads to 6 GW by 2029 — HimanshiET · 2026-09-10
- Same model, opposite curves: llama.cpp beats vLLM 7x at 90K context on DGX Spark — niacolhealth · 2026-09-10
- vLLM Inference Meetup lands in Bengaluru, co-hosted by AMD and Red Hat — dhruv2038 · 2026-09-10