Windows demo routes coding tasks to local model with GPU spinning, llama.cpp lands on Windows ML
ryanshrout · x · 2026-10-08
AMD's Ryan Shrout live-tweeted details from a Microsoft event demo featuring Satya Nadella and Jensen Huang:
- A coding task in AUTO mode was routed to the local model, with the GPU visibly spinning up on screen—showing hybrid cloud/edge routing can automatically decide where tasks run
- Officially confirmed: llama.cpp is now integrated with Windows ML, bridging the open-source inference stack and Windows' on-device AI framework
- A later demo showed a more complex prompt splitting work between GPT-6-Luna and locally launched agentic tasks
More from coding & agent
- Jeffrey Emanuel's "say no to process" agent skill kills Codex ceremony output — used hundreds of times a day — doodlestein · 2026-10-08
- a16z backs Preference Model, which open-sources Karotte RL environment framework battle-tested by 1M+ evals — a16z · 2026-10-08
- Every's agent skims meeting notes and only pings you when your name comes up — here's the 4-step setup — every · 2026-10-08
- Exa's setup page swaps dev docs for a copy-paste prompt your coding agent runs — josh_bickett · 2026-10-08
- Haiku 5.5 targets high-volume tasks, works as a coding subagent with Opus/Sonnet — claudeai · 2026-10-08
- Non-coder runs his entire business on an army of Claude Opus 5.5 agents — EXM7777 · 2026-10-08