US Labs Prefer RL for Reasoning, Chinese Labs Lean Towards SFT on Traces

xuanalogue · x · 2026-07-22

An observation highlights a divergence in reasoning training between US and Chinese AI labs: US labs favor Reinforcement Learning (RL), while Chinese labs more frequently use Supervised Fine-Tuning (SFT) on successful rollouts. The SFT approach is noted to be highly competitive, with Kimi K3 serving as soft evidence.

Related event: US-China LLM Training Path Divergence and New Alignment Ideas(5 posts)→

Original post →

More from Research

Research channel →