Local vs cloud decision models: 13.7ms on-device speed but coin-flip 50% accuracy
sven_ai · x · 2026-09-21
A benchmarker ran two rounds of 100-question decision-model duels on an M2 Pro, pitting cloud Jev against local Laya-MLX:
- Speed: local wins big — 13.7ms per question, under 2 seconds for 10 concurrent requests, QPS 58; cloud Jev needs 5s for 10 concurrent (385ms network RTT).
- Accuracy: local flops — Laya-MLX hits only 50%, coin-flip territory, while cloud Jev is slower but highly accurate.
Verdict: speed without accuracy is useless for production. The pragmatic play is a hybrid architecture — local for low-latency, high-frequency, privacy-sensitive workloads; cloud for complex reasoning and critical decisions.
More from Models
- Emad Mostaque: US firms can serve Kimi K3 at 1/10 the cost thanks to Nvidia, AMD chips — rohanpaul_ai · 2026-09-21
- Dev benchmarks Jev vs a local 7B model for LLM routing: Jev faster, 7B more accurate — tinyfool · 2026-09-21
- GPT-6 Astra claims Terminal-Bench Science lead at 65.7%, 31.4 points clear of second place — DeryaTR_ · 2026-09-21
- One Cent to 'Solve' Navier-Stokes: Jev Just Missed the Window — AAAzzam · 2026-09-21
- User hit with red warning bubble and 10-second lag after edgy ChatGPT questions — Acrobatic_Squid111 · 2026-09-21
- Tesla FSD v14.3.10 swerves around a giant mattress, showcasing end-to-end adaptability — xiaosun86 · 2026-09-21