Running Qwen 27B and DeepSeek v4 Flash together on one heterogeneous machine
samsja19 · x · 2026-09-16
A demo shows Qwen 27B and DeepSeek v4 Flash running simultaneously on a single Lucebox machine, pairing a discrete GPU with unified access memory for ample capacity plus decent FLOPS and bandwidth. The team is working to improve concurrency performance on heterogeneous setups.
More from Infra
- SentencePiece Lite ships: 50KB binary, 20-30x faster tokenization for edge devices — heiga_zen · 2026-09-16
- Leaked 10,000-word Huawei memo: become the NVIDIA for any LLM, pivot around Ascend 950 — pstAsiatech · 2026-09-16
- Doubao AI phone tops 360K pre-orders; RTX 5090 scalped at $9,500 — 创业邦 · 2026-09-16
- JPMorgan sees 25M+ GPU/ASIC shipments by 2028, ASICs dominate — a 'narrative violation' — bookwormengr · 2026-09-16
- iamtrask: The Endgame Is a Trust Web of Personal LLM Servers, Not One AGI — iamtrask · 2026-09-16
- Program-as-Weights: 0.6B interpreter matches Qwen3-32B prompting with 1/50 the memory — yuntiandeng · 2026-09-16