Researchers split on whether scaled self-play with capped APM can break model ceilings
Kangwook_Lee · x · 2026-09-26
Rémi Leblond argues a game AI trained purely via self-play from scratch, no human data, would improve further if capped at 500 APM peak / 300 APM average and trained with 10x more compute. Prof. Kangwook Lee is skeptical: unlikely, with no evidence supporting it — citing AlphaStar. A concise disagreement on whether scaled self-play keeps paying off.
Related event: Self-Play StarCraft AI Sparks AGI Benchmark Debate(4 posts)→
More from Research
- InternLM open-sources Intern-Decision 4B/0.8B: structured decisions in one forward pass — jacek2023 · 2026-09-26
- Stanford's Noah Goodman uses philosophy to improve LLM pretraining, jokes ASI achieved — xuanalogue · 2026-09-26
- Researchers surface spurious probes across models: Sonnet 5 recommends green tea in evals, oolong in production — jankulveit · 2026-09-26
- Three papers, one warning: 1% synthetic data can trigger strong model collapse — suchenzang · 2026-09-26
- Simulation beats distillation: the real story of synthetic data is post-training worlds — realsohamparekh · 2026-09-26
- Experts Rise Where LLMs Disagree: rationale labeling cuts codebook revision from months to days — windx0303 · 2026-09-26