VitalOps launches agentic inference optimization, 2.6x median speedup
abhijithneil · x · 2026-10-12
VitalOps announced it is going all in on inference optimization, offering serving, edge and GPU sim services.
The core is an agentic optimizer that searches the stack from the serving engine down and verifies correctness at the SASS level, backed by a server and engineers. It targets open-weight LLMs/VLMs on H100, B200, Jetson, Apple and Qualcomm hardware, with an enforced accuracy floor.
The product UI shows 23 runs this month with a 2.6× median speedup across verified runs: Qwen3-32B at 3.3× (6,120 tok/s) on 2×H100, DeepSeek-R1-32B at 2.9×, Llama-3.3-70B at 2.1× (searching), plus Whisper-large-v3 and bge-m3; kernel-level verification like fusedrmsnormquant at 18.4 µs with max err 3.1e-4.
More from Infra
- Engineer joins NVIDIA's Groq LPU compilers team, working on multi-chip partitioning — blelbach · 2026-10-12
- Linus Ekenstam wants nothing less than a 100B-param model running on your phone — LinusEkenstam · 2026-10-12
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Fireworks: open models plus fine-tuning match closed ones — Cursor gets 13x faster inference — AI Engineer · 2026-10-12
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12
- Running Qwen3.8 Flash-Next locally on AMD 7900 XTX at 500k context, 105-160 tok/s — human_in_the_looop · 2026-10-12