Nebius Engineers Break Down How to Serve Open LLMs Fast and Cheap in Production

AI Engineer · youtube · 2026-10-04

At AI Engineer World's Fair 2026, Nebius' Dylan Bristot and Sujee Maniyam walk through what it takes to run open LLMs in production as they near parity with proprietary models.

Key points:

A layer-by-layer roadmap (hardware → engine → routing → decoding → caching) for teams self-hosting open models, with Nebius Token Factory docs linked.

Original post →

More from Infra

Infra channel →