XeBoostLM: native C++ local LLMs on Intel NPUs and iGPUs, zero Python
Spiritual-Ad-5916 · reddit · 2026-09-17
A developer released XeBoostLM v0.1.2, a pure C++ CLI built on OpenVINO GenAI for running local LLMs on Intel Core Ultra NPUs, Arc iGPUs, and CPUs with zero Python overhead. It offers manual or HYBRID hardware routing, an OpenAI-compatible SSE streaming server that plugs into Open WebUI and LibreChat, INT4/INT8 model pulls (Qwen 2.5, Phi-3.5, Llama 3.2, DeepSeek R1), and a terminal REPL. Open source, feedback welcome.
More from Infra
- Do AI companies lose money on every token served? Krueger asks for accounting — DavidSKrueger · 2026-09-17
- After Iran Bans and 20,000 Leaked Private Repos, Should Devs Migrate Off GitHub? — FlolightC · 2026-09-17
- Running Qwen3.8 Flash on 12GB VRAM at 15 tokens/s with 3bpw quantization — KnownAd4832 · 2026-09-17
- DeepSeek V4 Pro parsing bug said to hit ~60% of OpenRouter providers; fix upstreamed to sglang — michellechen · 2026-09-17
- Unconventional AI Open-Sources Un-0, an Image Generator Built on Coupled Oscillators — NaveenGRao · 2026-09-17
- Naveen Rao's Startup Taped Out Custom AI Chip in 5 Months, Demonstrates On-Chip Causal Dynamics — NaveenGRao · 2026-09-17