Alibaba says Qwen has run recursive self-improvement for a month ahead of Qwen 4.0

teortaxesTex · x · 2026-10-11

At the Yunqi conference, Qwen lead Liu Dayiheng revealed Alibaba has already been running RSI (recursive self-improvement) continuously for over a month on Qwen 3.8 Max, and the entire Qwen 4.0 lineup (Max/Plus/Flash/27B) will adopt it deeply.

The pipeline: the model mines real logs for weaknesses, generates its own training data, validates via smaller models over multiple rounds, then feeds results into flagship training — with near-zero human intervention. Reported results: 33 effective iterations over a month, with Qwen3.8-Max's AA intelligence index climbing from 40 to 45 in 30 days.

Poster teortaxesTex remains skeptical, quipping "yeah… after a manner…", suggesting the marketing framing may overstate the actual gains.

Original post →

More from Models

Models channel →