$5, 10-minute SFT on Qwen3.6-35B-A3B lifts GPQA +8% and MMLU-Pro +12%
josh_wills · x · 2026-09-21
Developer ekzhang1 ran an evening SFT on Qwen3.6-35B-A3B via Tinker to better handle "Jev-y" prompts: cost $5, 10 minutes of training, +8% on GPQA diamond and +12% on MMLU-Pro, with Jared Palmer's Kev evaled for comparison.
He also open-sourced openjev-sglang: a TypeSafe/Jev-compatible HTTP API served by Qwen3.6-35B-A3B on SGLang, one B200 per container leveraging SGLang 0.5.19's Rust frontend, radix caching and breakable prefill CUDA graphs, with a FastAPI/uvloop async API layer and one-command deploy on Modal.
More from Infra
- Turn any local LLM into a confidence-scored classifier via logprobs, full llama.cpp recipe included — DivideHorror3217 · 2026-09-21
- Redditor spins up a 4x32GB V100 vLLM server, says local setup covers 90% of work — TrailFeatures · 2026-09-21
- Agent runtime promises billions of agents and 10-20x sandbox density — astralmatrix · 2026-09-21
- SGLang team helps user debug hicache crash, earning community praise — TheZachMueller · 2026-09-21
- M5 Max vs $2500 desktop upgrade: real-world local LLM speed comparison — vitamins1000 · 2026-09-21
- Two years after 'intelligence too cheap to meter', $10/$50 models are the new norm — teortaxesTex · 2026-09-21