U.S. labs lean on RL while Chinese labs favor SFT on successful traces

goyalshaliniuk · x · 2026-07-22

A short observation on training trends across labs:

The post is not a formal result, but a concise industry-level hypothesis about how reasoning models may be trained.

Related event: Diverging Paths: US Labs Favor RL for LLMs While Chinese Labs Prefer SFT(4 posts)→

Original post →

More from Research

Research channel →