Microsoft goes all-in on local AI: hybrid intelligence Windows, 1.6-bit DeepSeek V4 Flash
Sam Witteveen · youtube · 2026-10-08
Sam Witteveen breaks down Microsoft's Windows/Surface event, centered on a "hybrid intelligence" strategy: a Copilot router decides per-task whether to run locally or in the cloud, local first.
Key points:
- Extreme quantization for local models: MAI Code 1.1 Flash at 3-bit, a new Nemotron at 2-bit, and DeepSeek V4 Flash at just 1.6-bit
- llama.cpp embedded inside Windows ML; memory cost of 256K context discussed
- MXC: agent sandboxing in Windows
- Hardware: Surface Laptop Ultra with RTX Spark, plus Desktop and DGX Station for Windows, with specs and pricing
More from Embodied
- multimodalart gets hands on a microduck prototype hardware — multimodalart · 2026-10-09
- Interface, a $99 Handheld Controller for AI Agents, Opens Waitlist Ahead of January Ship — cneuralnetwork · 2026-10-09
- Apple announces surprise Oct 13 'Welcome home' event: smart home hub, Apple TV, LG devices — tomwarren · 2026-10-09
- Tesla Robotaxi team chased storms 24/7 to validate all-weather autonomous driving — yunta_tsai · 2026-10-09
- Agent Buddy: an ESP32 desk pet that shows what your coding agents are up to — DanWahlin · 2026-10-09
- Viture Beast on iPhone 18: reviewer says audio-only beats glasses for ambient compute — curious_vii · 2026-10-08