A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro

simonguozirui · x · 2026-09-23

Tinker demonstrates that LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface: discrete options in, fast probabilities out. ekzhang1 ran an evening SFT on Qwen3.6-35B-A3B to better handle Jev-style prompts on Tinker—$5 and a 10-minute run yielded +8% on GPQA diamond and +12% on MMLU-Pro, with @jaredpalmer's Kev evaled for comparison.

Original post →

More from Models

Models channel →