Colibrì Engine Update: Runs 2.8T Param Models, Boosts Speed via Expert Caching
solyarisoftware · x · 2026-08-27
Colibrì engine has evolved into a full system AI memory hierarchy: VRAM for hot experts, System RAM for warm experts, and NVMe SSD for everything else. It now supports models like Qwen3.6, DeepSeek V4 Flash, GLM-5.2, and the massive Kimi K3 (2.8T parameters). New features include GPU expert caching, prefetching, dual-NVMe striping, and support for CUDA, ROCm, Metal, Vulkan, and distributed expert workers. A benchmark showed expert caching boosting Qwen3.6-35B-A3B performance from 1.44 to 10.05 tok/s on two 8GB GPUs.
More from Infra
- Warning: $7T in Data Center Financing Repackaged into Complex Derivatives — SatelliteNetSec · 2026-08-27
- Why Is There No Fully Free Open Source AI Gateway? — tristanbob · 2026-08-27
- AWS to Deploy 2 Million Additional NVIDIA GPUs in 2027-2028 — BenBajarin · 2026-08-27
- Customers prefer performance over max AI cost savings — aronchick · 2026-08-27
- Zai achieves 3x cluster performance boost, matching NVIDIA GPU efficiency and cost — pstAsiatech · 2026-08-27
- View: Cloud compute service exploited; issues stem from time limits, not "bad agents" — dfrsrchtwts · 2026-08-27