Nebius Engineers Break Down How to Serve Open LLMs Fast and Cheap in Production
AI Engineer · youtube · 2026-10-04
At AI Engineer World's Fair 2026, Nebius' Dylan Bristot and Sujee Maniyam walk through what it takes to run open LLMs in production as they near parity with proprietary models.
Key points:
- Closed APIs vs self-hosting: open models are now competitive, cutting cost and lock-in; the talk frames the full inference → data → post-training → deployment loop.
- Hardware: NVFP4 quantization on latest NVIDIA hardware.
- Engine selection and cache-aware routing, which beats naive load balancing for LLM serving.
- Speculative decoding, including training custom draft models on your own traffic.
- KV cache offloading and disaggregated prefill/decode.
- Finding the quantization sweet spot between accuracy and throughput.
A layer-by-layer roadmap (hardware → engine → routing → decoding → caching) for teams self-hosting open models, with Nebius Token Factory docs linked.
More from Infra
- Qwen 3.8 Flash Next local tests hit 80-110 tok/s, poised to replace the 27b model — inthesearchof · 2026-10-04
- Simon Willison: pay-by-usage services need default hard budget caps in the agent era — Simon Willison · 2026-10-04
- The Future of Compute Is Fungible: Hyperscaler Economics Hinge on More Than Chip Prices — BenBajarin · 2026-10-04
- local-codex-proxy: run self-hosted models inside Codex alongside ChatGPT — TheZachMueller · 2026-10-04
- Investors want to fund "insurance for AI" startups wrapping data center leases — katieruthmishra · 2026-10-04
- Data tower waste heat could warm 100,000 homes: 134 MW at 85°C, ~1 TWh/year — IgorCarron · 2026-10-04