Transformers models now run natively in vLLM with no port required
pcuenq · x · 2026-09-16
Hugging Face engineer Pedro Cuenqa notes that Transformers models can now run natively in vLLM with no port required. HF's Harry Mellor will detail how a Transformers model loads in vLLM in a talk at the AI Engineer conference in Paris (Sep 23–24, co-organised by MistralAI), scheduled for 10:30 on the 24th.
More from Infra
- Novita open-sources Chord W4A16 MoE kernels, up to 2.15x faster Kimi K2.x inference on B300 — vllm_project · 2026-09-16
- VCs say the GPU shortage is really a capital access problem: 30% down payment prints money — sarahdrinkwater · 2026-09-16
- Developer builds Linux subsystem for Haiku with OCI runtime and container support — unixterminal · 2026-09-16
- Anthropic signs lease at Queensland data center park costing ~$30B, online 2027 — DigitalColmer · 2026-09-16
- TSMC builds 20 fabs yet can't meet AI demand as labor shortage slows expansion — emmanuelvivier · 2026-09-16
- SemiAnalysis: Nvidia Vera Rubin NVL72 Delivers Up to 30x Higher Throughput per MW Than Blackwell for Agentic Inference — emmanuelvivier · 2026-09-16