Qwen3.8-Max-0902 coding training lifts RSI-Exam recursive self-improvement score 22%
HuaxiuYaoML · x · 2026-09-04
Alibaba's Qwen team says comprehensive training on coding and cowork for Qwen3.8-Max-0902, targeting complex long-horizon tasks, generalized to a 22% jump on the RSI-Exam 0.1 leaderboard — from 0.322 to 0.392 on recursive self-improvement tasks, which researcher Huaxiu Yao calls 'a small step toward RSI.'
More from Models
- Qwen3.8-27b Is the First Local Model This User Can Blindly Trust for 8+ Hour Agent Runs — Express_Quail_1493 · 2026-09-04
- GLM runs four free-token events, unlimited coding use in daily window — Zai_org · 2026-09-04
- Muse Spark 1.3 3D Voxel Scenes Nearly Match GLM-5.3 and GPT-5.6 Sol — cedric_chee · 2026-09-04
- OpenAI API still burns reasoning tokens during context compaction even with effort set to none — mhmazur · 2026-09-04
- Speculation: OpenAI may internalize reasoning via synthetic CoT rewriting and FiM — cephaloform · 2026-09-04
- Scale alone can't explain why OpenAI and Anthropic stay ahead of same-scale rivals — haider1 · 2026-09-04