U.S. labs lean on RL while Chinese labs favor SFT on successful traces
goyalshaliniuk · x · 2026-07-22
A short observation on training trends across labs:
- U.S. AI labs appear to rely on RL more often.
- Chinese labs appear to use SFT on successful traces more often, and the author says this can be very competitive.
- Kimi K3 is cited as soft evidence for that claim.
The post is not a formal result, but a concise industry-level hypothesis about how reasoning models may be trained.
Related event: Divergent Sino-US Training Paradigms and New SFT Alignment Ideas(5 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11