Trending HF dataset: Qwen/GLM/Kimi multi-model distillation mix

lhoestq · x · 2026-08-21

The dataset r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation is trending #1 on Hugging Face. It contains approximately 22.8M rows and serves as a multi-teacher distillation dataset, combining outputs from models like Qwen, GLM, and Kimi. The data covers diverse tasks including SFT, reasoning, tool-use, long context, and math, with formats ranging from dialogue to agent interactions.

Original post →

More from Research

Research channel →