NNsight 0.8 Pre-release Brings Near-Native Throughput White-Box Inference

NNsight 0.8 pre-release rewrites the execution engine while keeping the API unchanged, supporting more models including MoE and all HF Transformers tasks, with interpretability analysis on vLLM at near-native throughput.

2026-09-10 ~ 2026-09-10 · 2 related posts