Solo Dev Fine-tunes 4B Decision Model to #1 on JevBench for ~$1,200

Educational-Care7867 · reddit · 2026-09-28

A business process consultant spent 15 days building ImaJev-4B: Qwen3.5-4B with a LoRA and a small decision head that takes text/JSON plus up to two photos and outputs per-option probabilities (with an explicit 'unknown') in one forward pass, aimed at replacing human decision nodes in workflow maps.

Training initially failed — 500k short decision samples dropped a 9B model's reasoning from 64.9 to 42.3 on JevBench hard (pattern matching). It recovered by generating hard questions with open models and keeping only samples where two AIs agreed.

Results: #1 of 91 on JevBench (67.37, equally weighting accuracy, calibration, speed, cost; #3 on accuracy alone) and #3 of 56 on DecisionBench, ahead of GPT-5.6 Luna and DeepSeek V4.1. Total cost $1,200 in rented GPU; Apache-2.0 weights, runs on a Mac via MLX or a single GPU.

Original post →

More from Models

Models channel →