Intel Ships OpenVINO 2026.4: Qwen3-VL, FLUX.2 on NPU, MTP Speculative Decoding
jacek2023 · reddit · 2026-09-17
Intel released OpenVINO 2026.4 with broad new model support (Qwen3-VL-4B with EAGLE3, Kokoro-82M, DeepSeek OCR-2, Granite 4.0 on CPU/GPU; FLUX.2-Klein 4B and Kokoro 82M on NPU), Multi-Token Prediction speculative decoding for Gemma 4 and Qwen models, Tree Drafting (Top-K) for EAGLE3, Xe3 iGPU optimizations on Core Ultra Series 3, ASRPipeline for Node.js, NPU profiling via VTune, and preview features like bounded dynamic shapes on NPU and idle model unloading in Model Server.
More from Infra
- One CPU Core Inspects 28.8M Packets/Sec: How XDP Rewrites Network Processing — blaizedsouza · 2026-09-17
- Dev's Rust/Wasm AAC Encoder Beats FFMPEG by 4-6x in Benchmarks — wavefnx · 2026-09-17
- OpenRouter weekly tokens surge 25,000% to 126.2 trillion — the AI bubble debate in one chart — The Decoder · 2026-09-17
- How do you verify an untrusted GPU host actually ran the model? Gonka's design notes on 3 cheating vectors — autoimago · 2026-09-17
- Lenovo launches ThinkAgile VX850 V4 servers that let apps stay on VMware or Hyper-V — shashib · 2026-09-17
- Running an Opus-level coding agent locally: MTPLX claims 126.5 TPS on Mac — julianharris · 2026-09-17