Open-weight 0.8B/2B System 1 decision models match Jev at 83.1%, trained fully on local hardware
Usual_Maximum7673 · reddit · 2026-09-29
A developer replicated TypeSafe's Jev experiment with open-weight Apache 2.0 "Jeff" fine-tunes of Qwen3.5 and Gemma — zero-shot classifiers that output calibrated probabilities per option in one forward pass, no text generation.
- Jeff-Qwen3.5-2B scores 83.1% on a five-benchmark panel (Jev: 83.0%, AutoJev-27B: 84.9%); the 0.8B hits 79.1%
- Training used 31k synthetic questions from Qwen3.8-Flash-Next on two DGX Sparks plus 271k dataset-derived questions; a single RTX PRO 6000 trains the 0.8B in 2 hours — no cloud, no closed-model data
- Speed: 28 ms per decision on an M4 Max vs Jev's 212 ms API call; fine-tuned to 24 ms at near-100% accuracy in a voice navigation app
- Zero-shot gaming: Jeff 0.8B matches Jev's Doom score (6.55) without an aiming rule in the prompt, though multi-step reasoning (BBH 64–68%) lags Jev's 94%
Takeaway: tiny models aren't reasoners, but they're extremely fast judgement-callers. Weights and code are public, with a Jev-compatible API.
More from Infra
- Swift 1.5 + HyperQwen cuts task time 37% at 100+ tok/s on a single RTX 3090 — KingGongzilla · 2026-09-29
- Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe — DimeRhyme · 2026-09-29
- Redditor predicts sub-$1000 device running SOTA models will spawn the next big company — Robert__Sinclair · 2026-09-29
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29