Jev replaces LLMs in edge service admission: 0.91 exact on-time completion where LLMs fall below 0.1

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration

Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu

cs.DC, cs.NI

2026-09-19

The Jev decision model replaces the admission-path LLM: median decision latency down 22.7–64.5%, exact on-time completion 0.91–0.95 where LLMs fall below 0.1.

What problem this solves

A request arriving at an edge node reads like this: read the text in this photo, keep the data on site, it's urgent. Turning it into a runnable job takes a handful of decisions first: service type, whether the payload may leave the site, quality tier, urgency. Those decisions sit on the admission path, ahead of queueing, transfer and execution. The experiments use a typical deadline of 2 s, while a single LLM interpretation call runs from a few hundred milliseconds to seconds and occupies one of a few admission slots, so slow calls back up the queue and completion collapses under load.

The premise here is that with a bounded service catalog, this interpretation is four to eight bounded decisions, not free-form generation. The question is whether a decision-oriented model can wholesale replace the LLM, and what that does to latency, accuracy and the bill.

Method

The roster: Jev, the self-hosted SemIf-Qwen3.5-4B and Laya (421M parameters), against three hosted LLMs (DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash), with a rule parser and a DistilBERT classifier as references.

Results

MetricJev-1.13.0Fastest LLM (DeepSeek-V4.1-Flash)
Exact match, clean requests0.9500.987
Exact match, 8-field contracts0.527–0.5770.900–0.930
Exact on-time completion, 1–16 req/s0.907–0.9530.093 at 16 req/s
Fees at 16 req/s (USD per 1k correct)0.0380.58

Across the 33 conditions, Jev's median decision latency sits 22.7–64.5% below the fastest LLM and barely moves: 0.27 s on clean requests and 0.42 s at 16,384 tokens, where DeepSeek climbs from 0.38 to 0.64 s. Bundle eight requests into one message and Jev takes 0.29 s for the whole message while Qwen3.8-Flash reaches 3.79 s. On four-field contracts, Jev's API fees per correct decision are 59.7–80.9% lower, at the price of a few exact-match points. Tails point the same way: 0.7% of Jev's clean-request decisions exceed 0.5 s, against 19.0% for DeepSeek and 100% for Qwen3.8-Flash.

End to end is where it counts. In the real OCR service, Tesseract itself reads only 92 of 180 images correctly, and Jev's correct completion of 0.511 exactly matches that recognizer ceiling; no interpreter ever sent an image to a node its locality field forbade. Under bursty arrivals GLM-5.3-Flash completes 0.167. Catalog churn is the decision model's strongest card: Jev names services it has never seen at 0.995–1.000 accuracy, indistinguishable from known services, while a DistilBERT classifier retrained on 92 labelled examples reaches 0.142 on new services.

Caching narrows everything. With eight recurring descriptions cached, every interpreter except Laya lands at 0.06–0.09 s median request latency and the completion gaps close. Jev's advantage belongs to fresh decisions only.

Why it matters

The paper turns an intuition into an engineering result with numbers: bounded classification on a hot path does not need generation. Flat latency, lower fees and catalogs that grow without retraining are directly usable for agent gateways, intent routing and function-call pre-parsing, where the schema often holds four or five fields and is a multiple-choice question in disguise. The boundary is drawn just as clearly: wide contracts and unsupported-request detection still favor LLMs. The honest summary is a trade, a few exact-match points for latency and throughput, valid on narrow contracts.

Limitations

The authors' own admissions: eight-field contracts mark the limit of the substitution, with Jev 32.3–40.3 points behind DeepSeek; unsupported-request detection sits at F1 0.723–0.929 against 0.889–1.000 for the LLMs; and fees grow with the option count, so from 64 services on Jev costs more per correct decision than DeepSeek.

Points that look weaker on a close read:

Terms

Source

Related papers

All paper explainers