Yacine on Chinese RL SOTA models: it's just on-policy synthetic data
yacineMTB · x · 2026-09-23
Yacine MTB pokes fun at the hype around Chinese RL models hitting SOTA, noting that the underlying recipe is simply on-policy synthetic data rather than a novel breakthrough.
More from Models
- User calls Opus 5.5 biggest jump since 4.6, fears launch-then-distill downgrade — haider1 · 2026-09-23
- Claude Opus 5.5 Appears on Anonymous AI Platform Venice Before Official Release — juanbenet · 2026-09-23
- Open source is catching up — unless frontier labs' internal iteration widens the gap — skorusARK · 2026-09-23
- Rails agent evals: OpenAI stays ahead, Luna Max the dark horse at 18% completion for just $11 — npew · 2026-09-23
- Pangram v4 flags pre-2020 thousand-line texts as 100% AI-generated — birchlse · 2026-09-23
- AI model release curve looks scarily exponential on a phone screen — StrategicHarmony · 2026-09-23