Two 300B MoE models on one 128GB Strix Halo: Kyojin engine hits 44 tok/s decode
Yaniss916 · reddit · 2026-10-03
Yamz-Labs released Kyojin, an inference engine built on ExLlamaV3 for AMD Strix Halo (gfx1151, ROCm), packing a 300B MoE onto each 128GB Ryzen AI Max+ 395 mini PC:
- GLM-5.3-Flash (99.7GB EXL3): 580 tok/s prefill at 3.5K context, 26-30 tok/s decode with MTP.
- MiMo-V2.6-Flash-MOPD (105GB, their own quantization): 650 tok/s prefill at 4K; 29 tok/s plain decode, up to 44 tok/s on code with speculative decoding.
- Quality: KLD vs official FP8 of 0.151 / 0.0713, top-1 agreement 89.3% / 92.0%. Their custom GLM layer mix beats turboderp's public 2.05bpw pack on KLD (0.190 vs 0.275), though his is smaller and 10% faster to decode.
- Separate -Uncensored repos via a single load-time file.
- Quickstart: clone, ./build.sh, hf download, serve.py — OpenAI-style API. Engine on GitHub, weights on Hugging Face; benchmarks and PRs welcome.
More from Infra
- Open-source Strata runs a 125B-param Qwen model on a 12GB consumer GPU — lxfater · 2026-10-03
- Chip design is a ~10^2,632,341 search problem — AI and agents are turning hardware into search — ai · 2026-10-03
- Redis Creator antirez Ships DwarfStar 4, an MIT-Licensed Local Inference Engine — petrusenko_max · 2026-10-03
- The $899 Mac Mini M6 was built for 24/7 agentic computing — most owners use it to check email — hey_abusiddik · 2026-10-03
- Modal VM Sandboxes hit GA: demo runs Docker Compose apps, tests and coding agents — charles_irl · 2026-10-03
- Community poll of 793 MLX users: oMLX wins at 55.5%, dominating Ultra chips — HankYeomans · 2026-10-03