Baseten's AI-generated inference engine beats vLLM by up to 90% on decode speed
baseten · x · 2026-10-05
Baseten engineer Shawn Rushefsky put the MetaInfer paper to the test: a skills-only framework (no post-training, no specialized harnesses) that generates custom LLM inference engines from scratch, built via a mostly autonomous week-long agent workflow.
- On a single B200 running Qwen-3.6-35B-A3B in NVFP4, the AI-generated engine beat vLLM 0.25.1 by up to 90% on single-stream decode speed and 2.33x on TTFT.
- The writeup highlights four converging trends — smarter long-horizon models, agent coordination, measurable optimization targets, and inference demand — enabling AI-assisted inference optimization.
- Takeaway: ultra-low-latency use cases that weren't practical before are now viable, and a skills-only approach can outperform SoTA open-source serving stacks.
More from coding & agent
- Is a missing 'capability layer' needed to stop agents reinventing the same tools? — Prestigious-Run-1954 · 2026-10-05
- Revera independently tests agent recovery after ambiguous MCP tool results — Street-Chest2270 · 2026-10-05
- 11 agents run my full trading pipeline autonomously — but none can place a trade — Glad-Ranger1879 · 2026-10-05
- HuggingFace tackles harness overfitting with multi-harness RL across Claude Code, Codex and more — huggingface · 2026-10-05
- Garry Tan: Lab-built harnesses burn tokens, and that's why startup harnesses like Grep matter — garrytan · 2026-10-05
- Dev builds a color-bar planning board to visualize long-running agent sessions — kevinkern · 2026-10-05