Qwen 3.8 Max reaches 42% on the hard INDUCTION benchmark, taking second place
DeryaTR_ · x · 2026-08-04
Qwen 3.8 Max jumps to 42% on the INDUCTION benchmark
A GitHub leaderboard shared in the thread shows Qwen 3.8 Max reaching 42.2% on the challenging INDUCTION benchmark, up from 0% for Qwen 3.7, and taking 2nd place behind GPT-5.6 Sol.
What the leaderboard says
- Qwen 3.8 Max: 63/64 evaluable, 27/64 correct (42.2%)
- GPT-5.6 Sol: 37/64 correct (57.8%)
- Fable 5: 26/64 correct (40.6%)
- The post highlights that Qwen’s result is much cheaper while closing the gap quickly.
Why it matters
The benchmark is presented as a difficult induction task from ICML 2026, so the result is framed as evidence of rapid progress in Chinese open-source models over the last few months.
More from Models
- Qwen3.8-Max fixes 19 of 105 hidden bugs in blind coding benchmark — breath_mirror · 2026-08-04
- xAI updates Grok Build with Grok 4.5, skills, MCP, and plan mode — elonmusk · 2026-08-04
- RL on custom search harnesses may beat the “one big model” idea — shangbinfeng · 2026-08-04
- OpenAI hires the creator of WebRTC as its GPT-Live voice system gets a deep dive — bookwormengr · 2026-08-04
- Code Arena WebDev puts four open-weight Chinese models near the frontier — floriandotorg · 2026-08-04
- GPT-5.6 Misspells Email Address, Then Hallucinates a Post-Hoc Excuse — WolframRvnwlf · 2026-08-04