Local Qwen matches Jev at 96.53% accuracy, 239 ms vs 368 ms median latency

Ok-Development6070 · reddit · 2026-09-19

A developer who didn't want to send data to third-party servers built choosekit, a small TypeScript package that gets choices and probabilities from a model running locally in llama.cpp.

Benchmarked with Qwen3.8 27B Q4 on SemIf's 144-task benchmark against Jev's hosted API, both achieved identical 96.53% accuracy — and the local setup was faster, with a 239 ms median response time versus 368 ms for the hosted API. Full setup and results are open-sourced on GitHub.

Original post →

More from Infra

Infra channel →