US Labs Prefer RL for Reasoning, Chinese Labs Lean Towards SFT on Traces
xuanalogue · x · 2026-07-22
An observation highlights a divergence in reasoning training between US and Chinese AI labs: US labs favor Reinforcement Learning (RL), while Chinese labs more frequently use Supervised Fine-Tuning (SFT) on successful rollouts. The SFT approach is noted to be highly competitive, with Kimi K3 serving as soft evidence.
Related event: US-China LLM Training Path Divergence and New Alignment Ideas(5 posts)→
More from Research
- Niantic Spatial and Flexion push humanoid sim2real training with NVIDIA Isaac — ExtensionEcho3 · 2026-07-22
- ASCIITermDraw-Bench tests whether VLMs can generate and edit diagrams in ASCII — East-Muffin-6472 · 2026-07-22
- David Silver and Richard Sutton say AI is entering an era of experience — willccbb · 2026-07-22
- Muon nearly doubles agentic RL success in GiGPO, but only with the right setup — rohanpaul_ai · 2026-07-22
- Letta pitches OS-style long-term memory for local LLM agents — thisguyknowsai · 2026-07-22
- LoRA step counts can differ across AI-Toolkit and OneTrainer — BelowSubway · 2026-07-22