Kev 4B matches Jev on accuracy but Jev bills ~257 extra tokens per request, up to 12x cost

facethef · reddit · 2026-09-26

The Opper team benchmarked Kev 4B, Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B, against Jev side by side on the same endpoint, using a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues) to avoid data contamination.

Findings

Benchmark code, test items and results are open-sourced on GitHub. The takeaway: an open 4B fine-tune already matches the commercial model on capability, while billing quirks may matter more than accuracy for real-world cost.

Original post →

More from Models

Models channel →