NNsight 0.8 Pre-release Brings Near-Native Throughput White-Box Inference
NNsight 0.8 pre-release rewrites the execution engine while keeping the API unchanged, supporting more models including MoE and all HF Transformers tasks, with interpretability analysis on vLLM at near-native throughput.
2026-09-10 ~ 2026-09-10 · 2 related posts
- NNsight 0.8 pre-release ships faster engine, MoE and near-native vLLM support — davidbau · 2026-09-10
- NNsight 0.8 pre-release ships new engine for near-native vLLM interpretability — gsarti_ · 2026-09-10