ZML launches a free inference server to run LLMs across Nvidia, AMD, TPU and more

emmanuelvivier · x · 2026-07-28

French startup ZML has released LLMD, a free inference server designed to run open-source LLMs across multiple chip families, including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc.

The company says the goal is to break hardware silos, reduce vendor lock-in, and let models run at peak speed on whichever chips are available. TechCrunch frames the launch as part of the growing push to optimize inference rather than training as AI usage scales.

Original post →

More from Companies & People

Companies & People channel →