Weekly must-read papers: MoE looped transformer scaling laws, world model physics evals
TheTuringPost · x · 2026-09-09
TheTuringPost curates this week's must-read papers:
- Architecture & scaling: SMELT on compute-matched scaling laws for MoE looped transformers; a paper arguing modern transformers are implicit hybrids
- Reasoning & training: RISE (recursive improvement via self-extrapolating policy distillation), extremely sparse supervision incentivizing reasoning, and a follow-up rethinking on-policy distillation of LLMs
- World models (the bulk): WeatherNext 3 improving global weather model resolution from raw observations; World-Coherent Decoding for self-verifying test-time planning; VeriPhy and Principia on physics-based evaluation of world/video models; WISE for world-model-guided imagination scheduling in VLA post-training; TourPhysics on physics from a single image
More from Research
- EPFL talk slides on fractal maps in Lenia, citing Yevenko and Davis 2024 papers — BertChakovsky · 2026-09-09
- Cerebras paper: layer dropout saves up to 25% training FLOPs and yields 1.55x faster decoding — burny_tech · 2026-09-09
- Tencent paper: continuously harder task environments beat co-evolution, +8.6pp on Terminal-Bench — rohanpaul_ai · 2026-09-09
- Pretraining gains come mostly from data: experiments show 12x vs 3.7x compute multipliers — eliebakouch · 2026-09-09
- Mechanize's Tamay Besiroglu: why labs achieve the same breakthroughs at the same time — tamaybes · 2026-09-09
- DriveZero: End-to-End Autonomous Driving Beyond Human Demonstrations — Hao He · 2026-09-09