Kev 4B matches Jev on accuracy but Jev bills ~257 extra tokens per request, up to 12x cost
facethef · reddit · 2026-09-26
The Opper team benchmarked Kev 4B, Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B, against Jev side by side on the same endpoint, using a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues) to avoid data contamination.
Findings
- Accuracy lands within 2 points on every task — inside the noise at this sample size
- Jev is better calibrated and wins on paraphrase detection (PAWS 87.0% vs 74.5%)
- Same list price, but Jev counts a fixed 257 extra input tokens per request (also when calling TypeSafe directly), making short requests cost up to 12x more
Benchmark code, test items and results are open-sourced on GitHub. The takeaway: an open 4B fine-tune already matches the commercial model on capability, while billing quirks may matter more than accuracy for real-world cost.
More from Models
- Anthropic and OpenAI shipped cheaper models 101 minutes apart — altryne · 2026-09-26
- ImageJevBench: image decision benchmark ranks top models, full eval costs $0.02 — airesearch12 · 2026-09-26
- Gemini 4 Pro Spotted in AI Arena: Leak Claims Google's Comeback Is Coming — mark_k · 2026-09-26
- Xiaomi's MIT-Licensed MiMo-V2.6-Pro Hits OpenRouter at $0.87/M Output Tokens — VraserX · 2026-09-26
- Pangram AI detector only catches sloppy AI text, human-edited output passes — soumitrashukla9 · 2026-09-26
- PhD student one-shots 3Blue1Brown-style paper animations with Opus 5.5 in one hour — hsu_byron · 2026-09-26