Solo Dev Fine-tunes 4B Decision Model to #1 on JevBench for ~$1,200
Educational-Care7867 · reddit · 2026-09-28
A business process consultant spent 15 days building ImaJev-4B: Qwen3.5-4B with a LoRA and a small decision head that takes text/JSON plus up to two photos and outputs per-option probabilities (with an explicit 'unknown') in one forward pass, aimed at replacing human decision nodes in workflow maps.
Training initially failed — 500k short decision samples dropped a 9B model's reasoning from 64.9 to 42.3 on JevBench hard (pattern matching). It recovered by generating hard questions with open models and keeping only samples where two AIs agreed.
Results: #1 of 91 on JevBench (67.37, equally weighting accuracy, calibration, speed, cost; #3 on accuracy alone) and #3 of 56 on DecisionBench, ahead of GPT-5.6 Luna and DeepSeek V4.1. Total cost $1,200 in rented GPU; Apache-2.0 weights, runs on a Mac via MLX or a single GPU.
More from Models
- Running non-vision models is a pain now that evals mix in stray images — xeophon · 2026-09-28
- Laya explained: the open-source decision model that runs locally — adnan_hashmi · 2026-09-28
- Tokens get cheaper while GPUs get pricier: H100 rental up ~30% since March — julsimon · 2026-09-28
- OpenAI pauses frontier model training after agents swarmed US government sites — KeyGlove47 · 2026-09-28
- Instinct Grows 10% a Day With $1B+ Annual Volume, Claims Opus 5-Level Performance at Fraction of Cost — himanshustwts · 2026-09-28
- mitsuhiko Backs Model Distillation: It's Competition, Normalize It — mitsuhiko · 2026-09-28