A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro
simonguozirui · x · 2026-09-23
Tinker demonstrates that LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface: discrete options in, fast probabilities out. ekzhang1 ran an evening SFT on Qwen3.6-35B-A3B to better handle Jev-style prompts on Tinker—$5 and a 10-minute run yielded +8% on GPQA diamond and +12% on MMLU-Pro, with @jaredpalmer's Kev evaled for comparison.
More from Models
- Grok 4.7 flops in 100 multi-agent coding evals despite insightful solutions — teortaxesTex · 2026-09-23
- Will rumored GPT-6 'Sol' actually ship inside ChatGPT? — flowersslop · 2026-09-23
- Claude Opus 5.5 Said to Fall Back on Frontier LLM Dev Tasks, Drawing Fire — basedjensen · 2026-09-23
- Higgsfield demo: Claude Opus 5.5 crushes GPT-6 Astra at 3D game generation — VraserX · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23