K2-Horizon-7B gets 8-bit quant, 4-bit and VRAM-saving plan in progress

karminski3 · x · 2026-09-16

Community developer @waterluffy shared quantization progress on IFM/K2-Horizon-7B: the 8-bit version (IFM/K2-Horizon-7B-FP8) is already up on Hugging Face, a 4-bit build is in the works, and a separate "VRAM rescue plan" is underway. A commenter suggests mixed attention could be next. The HF model card shows a full tool-calling chat template supporting /xml formats.

Original post →

More from Models

Models channel →