U.S. labs lean on RL while Chinese labs favor SFT on successful traces
goyalshaliniuk · x · 2026-07-22
A short observation on training trends across labs:
- U.S. AI labs appear to rely on RL more often.
- Chinese labs appear to use SFT on successful traces more often, and the author says this can be very competitive.
- Kimi K3 is cited as soft evidence for that claim.
The post is not a formal result, but a concise industry-level hypothesis about how reasoning models may be trained.
Related event: Diverging Paths: US Labs Favor RL for LLMs While Chinese Labs Prefer SFT(4 posts)→
More from Research
- RAND publishes first roadmap for protecting valuable algorithmic know-how — Scobleizer · 2026-07-22
- SWE-Pruner Pro Shows Coding Agents Already Know What Context to Drop — rohanpaul_ai · 2026-07-22
- SkewAdam cuts MoE optimizer memory by 97.4% and fits 6.78B on a 40GB GPU — Kooky-Ad-4124 · 2026-07-22
- New paper defines four conditions that turn an LLM into a coding agent — alex_verem · 2026-07-22
- Spectral clustering method groups Markov chains via P² eigenvectors and weighted k-means — michaelchchoi · 2026-07-22
- NVIDIA and ETH Zürich cut small-message AllReduce latency by deleting barriers — thoefler · 2026-07-22