ZML launches a free inference server for open-source LLMs across Nvidia, AMD, TPU and more

emmanuelvivier · x · 2026-07-29

French startup ZML has launched a free inference server, ZML/LLMD, that speeds up open-source LLMs across many different chips, including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc.

The company says its goal is to break chip silos and deliver peak performance regardless of hardware, reducing vendor lock-in at inference time. Founder Steeve Morin argues that inference is becoming more important than training and is still burdened by software and architecture fragmentation.

The thread also mentions SambaNova’s reported $1 billion Series F at an $11 billion valuation, underscoring how aggressively infrastructure players are still being funded.

Related event: ZML Releases Free Inference Server Across Multiple AI Chips(3 posts)→

Original post →

More from Venture

Venture channel →