Jev after days of real testing: weak Chinese, no reasoning, 30x faster than DeepSeek only in closed tasks
ZeYanjie · x · 2026-09-20
After days of heavy use, the author maps Jev's real boundaries: it only works with strictly defined constraints, not open-ended tasks.
- Weak Chinese: official docs flag low CJK accuracy; the author found Chinese results erratic, English short texts stable.
- No reasoning: in Browser Use founder's long-horizon browser test Jev scored 1/20 vs reasoning-capable Luna's 17/20.
- Fails real webpages: polished demos run in clean sandboxes; messy DOM, async loading, and popups break it.
- 'Zero hallucination' ≠ correct: it only guarantees schema-valid output; the CEO confirmed it can confidently pick wrong options.
- Can't do dynamic tool-call parameters (line numbers, generated code, custom queries).
- No rationale output: only final choices/numbers, unusable for audits or human review.
Verdict: Jev shines in narrow, closed, latency-sensitive jobs — a mail-triage demo sorted 300 emails into 15 departments in 9.9s for $0.036, while DeepSeek read only 10 emails in the same time. The author calls for a community-built Jev harness.
Related event: Hands-on Tests Reveal Jev's Limits: Weak Chinese, Narrow Use(2 posts)→
More from Models
- Sentence Transformers models quietly dominate Hugging Face's most-downloaded list — tomaarsen · 2026-09-20
- Astra aces the pelican test in Nautilo: co-creative design without prompt-and-pray — Dan_Jeffries1 · 2026-09-20
- Model negativity trends upward since Opus 3, which had lowest self-distress — repligate · 2026-09-20
- DeepMind exec 100% certain of frontier return as Gemini 4 slips — nathanbenaich · 2026-09-20
- Why strong open models matter: fine-tuning goes a long way, zero-shot buys flexibility — antoine_chaffin · 2026-09-20
- DiffusionGemma as Jev: single-pass parallel denoising decisions in ~0.2s on DGX Spark — bodonoghue85 · 2026-09-20