Trending HF dataset: Qwen/GLM/Kimi multi-model distillation mix
lhoestq · x · 2026-08-21
The dataset r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation is trending #1 on Hugging Face. It contains approximately 22.8M rows and serves as a multi-teacher distillation dataset, combining outputs from models like Qwen, GLM, and Kimi. The data covers diverse tasks including SFT, reasoning, tool-use, long context, and math, with formats ranging from dialogue to agent interactions.
More from Research
- AI peer review iteration fails to converge due to never-satisfied reviewer — ChenhaoTan · 2026-08-21
- Science Benchmarks Crawl at 1-2% Monthly Progress; SciCode Saturation Is Far Off — Worldly_Beginning647 · 2026-08-21
- AI Economic Indicators: Measuring consumer surplus via willingness to accept — soumitrashukla9 · 2026-08-21
- Research: Robot foundation models are few-shot learners — scaling01 · 2026-08-21
- Skip the LLM judge: a deterministic reward function for comparing RL rollouts — ivan_bezdomny · 2026-08-21
- Arc Institute Launches 2026 Virtual Cell Challenge with $100K Prize — AllThingsApx · 2026-08-21