Local Qwen matches Jev at 96.53% accuracy, 239 ms vs 368 ms median latency
Ok-Development6070 · reddit · 2026-09-19
A developer who didn't want to send data to third-party servers built choosekit, a small TypeScript package that gets choices and probabilities from a model running locally in llama.cpp.
Benchmarked with Qwen3.8 27B Q4 on SemIf's 144-task benchmark against Jev's hosted API, both achieved identical 96.53% accuracy — and the local setup was faster, with a 239 ms median response time versus 368 ms for the hosted API. Full setup and results are open-sourced on GitHub.
More from Infra
- VanEck: NVDA's bigger risk is customers can't get power; powered-land base case implies ~83% upside — menhguin · 2026-09-19
- Reverse-engineering Claude's subscription limits from unrounded floats: Max 5× is the real sweet spot — RexDouglass · 2026-09-19
- Hyperscaler ROIIC Peaked Near 40% vs 8% Cost of Capital, AI Capex Math Shows — menhguin · 2026-09-19
- ByteDance said to dominate Malaysia datacenter capacity as China buildout sparks debate — jwt0625 · 2026-09-19
- Databricks brings Unity Gateway to Neon, billed as fastest AI gateway for Kimi K3 — Yuchenj_UW · 2026-09-19
- $135 MI50 + RX 7900 GRE dual-GPU Vulkan benchmarks across 19 GGUF models — tabletuser_blogspot · 2026-09-19