Qwen Learns to Train Qwen

DanAiTuning · reddit · 2026-07-14

The author states they trained Qwen3.6-35B-A3B into a model that "uses RL to train other models": upon receiving a task, it writes the entire training job—including environment, rewards, datasets, and hyperparameters—and submits it to real GPUs for execution.

Training Method

Results

The author open-sourced the repository, task families, reward code, GPU scheduling, and a retrospective.

Related event: Developer Uses RL to Train Qwen to Train Other AI Models(2 posts)→

Original post →

More from Research

Research channel →