Colibrì Engine Update: Runs 2.8T Param Models, Boosts Speed via Expert Caching

solyarisoftware · x · 2026-08-27

Colibrì engine has evolved into a full system AI memory hierarchy: VRAM for hot experts, System RAM for warm experts, and NVMe SSD for everything else. It now supports models like Qwen3.6, DeepSeek V4 Flash, GLM-5.2, and the massive Kimi K3 (2.8T parameters). New features include GPU expert caching, prefetching, dual-NVMe striping, and support for CUDA, ROCm, Metal, Vulkan, and distributed expert workers. A benchmark showed expert caching boosting Qwen3.6-35B-A3B performance from 1.44 to 10.05 tok/s on two 8GB GPUs.

Original post →

More from Infra

Infra channel →