First Jev Test Shows Only 80% Agreement with Verified Gemini 3.5 Flash Workflow
mayfer · x · 2026-09-16
mayfer shares a first test of Jev: results are decent but it only reaches about 80% agreement with a verified Gemini 3.5 Flash workflow. After some tuning it might hit acceptable levels, but it's not yet a drop-in replacement.
More from Models
- MindTopo benchmark: best MLLM scores 54.1% vs 97.4% human on topological reasoning; GPT-6 Astra needs 8+ hours per task — yining_hong · 2026-09-16
- all-MiniLM-L6-v2 still trending on Hugging Face, its creator says stop using it — tomaarsen · 2026-09-16
- HF researcher publicly questions Claude limits: monthly exhausted but weekly 84% left — NielsRogge · 2026-09-16
- Meta's promised Muse Spark open weights are over a month late and counting — RishiFurfox · 2026-09-16
- Ex-Huawei researcher: AI is great at small-step optimization, but taste and system design remain out of reach — yangyi · 2026-09-16
- New model Jev plays Super Mario Bros in real time on fast inference — hardimanjames · 2026-09-16