ZML launches a free inference server to run LLMs across Nvidia, AMD, TPU and more
emmanuelvivier · x · 2026-07-28
French startup ZML has released LLMD, a free inference server designed to run open-source LLMs across multiple chip families, including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc.
The company says the goal is to break hardware silos, reduce vendor lock-in, and let models run at peak speed on whichever chips are available. TechCrunch frames the launch as part of the growing push to optimize inference rather than training as AI usage scales.
More from Companies & People
- Together AI and Moonshot AI set July 30 webinar on Kimi K3 architecture — togethercompute · 2026-07-29
- Koch v1.1 open-source robot arm joins Hugging Face LeRobot as Tau launches $30/hour humanoid cleaning in San Francisco — RemiCadene · 2026-07-29
- Pause AI’s communications strategy is being blamed for turning 150 million Americans against data centers — wordgrammer · 2026-07-29
- Sam Altman says founders should look for the next wave, not copy the hot trend — heyshrutimishra · 2026-07-29
- Crew launches Crew Studio after studying agents in production at big companies — MatthewBerman · 2026-07-29
- Reactorfield opens a 4-week AI fellowship for scientists and deep-tech startups — MaxUnfried · 2026-07-29