Perturbed public documents synthesize training data, lifting Qwen 35B to trillion-param level

AllSpark-Research · hf · 2026-09-29

AllSpark-Research proposes a human-annotation-free pipeline that perturbs public documents to synthesize context-dependent training data. SFT + rubric-reward RL raises a Qwen3.6-35B-A3B student from 13.7% to 24.6% on CL-bench, matching trillion-parameter Qwen3.8-2.4T (23.9%).

Original post →

More from Research

Research channel →