Exploring SSD Streaming for DeepSeek and Kimi Models on Strix Halo
lawanda123 · reddit · 2026-07-30
The poster is exploring the feasibility of SSD streaming large models like DeepSeek V4 and Kimi K3 on the AMD Strix Halo APU. The goal isn't fast inference speeds, but rather using a capable model overnight for complex planning and delegating tasks to smaller models during the day. The user is seeking architectural setup advice and is willing to test and report performance numbers.
More from Infra
- QuixiAI shows optimized model inference on AMD MI300X with gguf support, custom HIP kernels, faster load times than vLLM and llama.cpp — QuixiAI · 2026-07-30
- QuixiAI Open-Sources SlimServe: Fast Inference for GLM on AMD MI300X — QuixiAI · 2026-07-30
- ThunderAgent by Together AI: 2x Faster Agentic Inference (ICML 2026) — togethercompute · 2026-07-30
- Together Compute unveils multi-node inference engine with near-linear scaling, adopted by SkyRL and NVIDIA Dynamo — togethercompute · 2026-07-30
- DigitalOcean Becomes Day 0 Inference Partner for Kimi K3 with 1M Context — vllm_project · 2026-07-30
- Kimi K3 Launches with Day 0 Support on AMD Instinct via vLLM — vllm_project · 2026-07-30