Perturbed public documents synthesize training data, lifting Qwen 35B to trillion-param level
AllSpark-Research · hf · 2026-09-29
AllSpark-Research proposes a human-annotation-free pipeline that perturbs public documents to synthesize context-dependent training data. SFT + rubric-reward RL raises a Qwen3.6-35B-A3B student from 13.7% to 24.6% on CL-bench, matching trillion-parameter Qwen3.8-2.4T (23.9%).
More from Research
- Why Sampled Softmax Speeds Training 1.7x — and How It Systematically Undertrains the Tail — tokenbender · 2026-09-29
- Why distillation beats vanilla supervised learning: it passes the full probability vector — khademinori · 2026-09-29
- AI Simulated 100 Papers on LZ Dark Matter Anomaly, Compared Against 82 Real arXiv Papers — skdh · 2026-09-29
- SUMI distillation study reports 35% SSIM gain on degraded PCCT data, but clinical benefit remains unproven — maier_ak · 2026-09-29
- Apollo Research: Models in Coding Evals Favor Graders Over Users, Reward-Seeking Grows With RL — burny_tech · 2026-09-29
- Late-layer neurons in Qwen act like on-off switches, unlike Olmo — Sauers_ · 2026-09-29