KTransformers: LLM Inference and Fine-Tuning Framework

thesupermanmx · x · 2026-07-19

KTransformers is a research project focused on large model inference and fine-tuning, featuring CPU-GPU heterogeneous computing optimization. It currently offers two ready-to-use entry points: Inference and SFT.

Recent updates listed include Day0 support for models like MiniMax-M3, GLM-5.2, DeepSeek-V4-Flash, and Kimi-K2.5. It also adds an AVX2-only CPU backend, CPU-GPU Expert Scheduling, Native BF16/FP8 per-channel precision, and a unified AutoDL fine-tuning/inference workflow.

Related event: ktransformers Enables Local Inference of Massive Models on 24GB VRAM(4 posts)→

Original post →