Lenovo Yoga Pro 9n runs a 120B-parameter model locally in a 1.65kg Windows laptop, with 128GB unified memory
APPSO · wechat · 2026-09-07
At IFA 2026, Lenovo unveiled the Yoga Pro 9n with NVIDIA's RTX Spark superchip (20-core Grace CPU fused with a 6,144-CUDA-core Blackwell GPU via NVLink-C2C), delivering 1 PFLOP (FP4) in a 1.65kg chassis with up to 128GB unified memory — the first mass-produced Windows laptop able to run a 120B-parameter model with 1M-token context locally.
Unified memory breaks the VRAM bottleneck, echoing Apple's playbook (M5 Ultra Mac Studio at 512GB; 4-unit Thunderbolt clusters running trillion-parameter models). The significance is bringing this into the Windows+CUDA ecosystem, challenging Mac and Linux dominance of local LLM workstations.
Lenovo also showed the ThinkCentre X Ultra (a 1.6L DGX Spark rival with 128GB unified memory, clusterable up to four units), dropped Qira's requirements to 16GB-RAM PCs, and previewed rollable-screen and fanless solid-state-cooling concepts. The AI PC metric is shifting from feature count to how much cloud work can stay on your machine.
More from Infra
- Microsoft open-sources tgrep, a trigram-indexed grep up to 52x faster than ripgrep — jedisct1 · 2026-09-07
- How should billing work when an AI system auto-selects the model? — Colddew-YJ · 2026-09-07
- Can you run Qwen Next on a 3090 + 64GB CMP 170HX? Local deployment help — JustinPooDough · 2026-09-07
- SmolVM: open-source microVM sandbox runs OpenClaw 2.0 in isolation, boots in milliseconds — aniketmaurya · 2026-09-07
- Hesamation recommends the best technical book on training LLMs at scale — free to read — Hesamation · 2026-09-07
- After His OpenAI Key Was Stolen, He Found Stratum: a Docker-Layer Secret Scanner Crunching 700K Layers Daily — Ubunta · 2026-09-07