KTransformers: LLM Inference and Fine-Tuning Framework
thesupermanmx · x · 2026-07-19
KTransformers is a research project focused on large model **inference** and **fine-tuning**, featuring CPU-GPU heterogeneous computing optimization. It currently offers two ready-to-use entry points: `Inference` and `SFT`. Recent updates listed include Day0 support for models like MiniMax-M3, GLM-5.2, DeepSeek-V4-Flash, and Kimi-K2.5. It also adds an AVX2-only CPU backend, CPU-GPU Expert Scheduling, Native BF16/FP8 per-channel precision, and a unified AutoDL fine-tuning/inference workflow.
Related event: ktransformers Enables Local Inference of Massive Models on 24GB VRAM(4 posts)→