U.S. labs lean on RL while Chinese labs favor SFT on successful traces

goyalshaliniuk · x · 2026-07-22

A short observation on training trends across labs:

The post is not a formal result, but a concise industry-level hypothesis about how reasoning models may be trained.

Related event: Divergent Sino-US Training Paradigms and New SFT Alignment Ideas(5 posts)→

Original post →

More from Research

Research channel →