ZML launches a free inference server for open-source LLMs across Nvidia, AMD, TPU and more
emmanuelvivier · x · 2026-07-29
French startup ZML has launched a free inference server, ZML/LLMD, that speeds up open-source LLMs across many different chips, including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc.
The company says its goal is to break chip silos and deliver peak performance regardless of hardware, reducing vendor lock-in at inference time. Founder Steeve Morin argues that inference is becoming more important than training and is still burdened by software and architecture fragmentation.
The thread also mentions SambaNova’s reported $1 billion Series F at an $11 billion valuation, underscoring how aggressively infrastructure players are still being funded.
Related event: ZML Releases Free Inference Server Across Multiple AI Chips(3 posts)→
More from Venture
- RCS Traffic Jumps 1298%: Infobip Builds Moat With 14K AI Agent Registrations — shashib · 2026-07-29
- AI capital is still flowing, but smaller open-weights and inference bets may be easier to fund — vaibhavbetter · 2026-07-29
- AI should help skilled workers run companies, not just shrink management — Helpful-West8007 · 2026-07-29
- AI Chip Startup SambaNova Raises $1B at $11B Valuation — emmanuelvivier · 2026-07-29
- SF Startup Toxicity: Bragging About Seed Rounds Exceeding Series B — iScienceLuvr · 2026-07-29
- South Korea’s AI-chip rally unwinds as KOSPI drops 41% from its peak — SumitGup · 2026-07-29