vLLM team launches Inferact, powers new HUMAIN-M3 Arabic frontier model's inference
woosuk_k · x · 2026-09-04
The core vLLM team has unveiled their new company Inferact, focused on the hardest problems in AI inference: optimizing token quality with model vendors and deployment for sovereign AI partners, claiming large throughput gains even under strict constraints. First public win: HUMAIN-M3, a frontier Arabic model from HUMAIN and MiniMax, runs on vLLM under the hood and is live on HUMAIN Node. The team is hiring and open to compute partnerships.
More from Infra
- Open-source voice pipeline adds Smart Turn end-of-turn gate before LLM calls — ivan_digital · 2026-09-04
- Pedro Domingos: Double LLM efficiency and you should be worth $100B, given Nvidia's math — pmddomingos · 2026-09-04
- AMD MI355x beats Nvidia B300 on tokens-per-dollar TCO in AgentX, SemiAnalysis says — AnushElangovan · 2026-09-04
- Liquid AI launches Nanos: task-specific 350M-2.6B models that run on-device — JosephJacks_ · 2026-09-04
- Miles Brundage: AI outage cascade likely caused by enterprises mass-switching to new models — Miles_Brundage · 2026-09-04
- OpenAI's Astra can now layout and route PCBs, sparking hardware engineering debate — MikePFrank · 2026-09-04