16 semantic questions route LLMs at 72.6%, beating dense embeddings and the best fixed model

SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis

cs.AI, cs.CL, cs.LG

2026-09-28

SeLMRoute routes each query via 16 interpretable semantic judgments kept as probabilities, reaching 72.64% AvgAcc on LLMRouterBench vs 69.23% for the best fixed model.

What problem this solves

LLM routing asks a concrete question: given an incoming request and a pool of candidate models, which model should handle it? Existing routers learn this decision directly from query embeddings, learned model representations, or clusters of similar queries. Those representations capture similarity, and similarity is not requirement. Two Python requests, one asking for dictionary-comprehension syntax and one asking to diagnose a race condition across async functions, sit near each other in embedding space while demanding very different capabilities. The LLMRouterBench evaluation added an uncomfortable finding: leading routers cluster in a narrow band around 70-72 AvgAcc, all approaching the Dataset Oracle that picks one best model per dataset. Most measurable routing gain comes from coarse domain structure, leaving query-level information on the table. SeLMRoute asks whether a set of explicit, readable task-requirement judgments can do the job without returning to a large representation.

Method

The routing function is factorized into three stages, each solving a different problem.

The factorization buys practical things. Swapping the candidate pool or the deployment objective does not require re-extracting semantics; adding a model requires performance calibration data but no re-encoding of queries; every routing decision can show its evidence, such as high code reasoning, high exactness, high constraint density. The evaluation protocol also groups duplicate normalized queries so the same query never straddles train and test, collapsing 11,481 instances into 11,423 groups.

Results

Main setting: the LLMRouterBench performance-oriented pool, 15 datasets, 20 lightweight models, 11,481 queries.

RepresentationAvgAccGain@B
SeLMRoute probability mass (40-dim)72.08 ± 0.455.72%
SeLMRoute full (88-dim)71.30 ± 0.184.58%
SeLMRoute hard labels (16-dim)71.20 ± 0.594.41%
GTE-Qwen2 dense embeddings + CatBoost71.16 ± 0.744.34%
Domain label only71.08 ± 1.034.45%
TF-IDF70.87 ± 0.783.82%

References on the same splits: Best Single 68.81, Dataset Oracle 73.94, Instance Oracle 91.91. Under grouped five-fold out-of-fold evaluation, probability mass reaches 72.64% against 69.23% for the strongest fixed candidate (Qwen3-8B), a 3.4-point margin, with only 1.86 points left to the Dataset Oracle.

The most informative ablations:

Distribution shift splits the field. Holding out single datasets, the semantic state stays competitive (67.02% vs 67.29% for GTE-Qwen2). Holding out an entire domain, dense embeddings win clearly (67.31% vs 65.66%).

In the performance-cost setting (13 flagship models, 10 datasets, 12,446 queries, GPT-5 as reference), PerfGain is positive in all five grouped splits with a mean of +2.66% ± 1.85%, but no split achieves strict monetary savings; the two qualifying configurations cost 0.10% and 1.58% more than GPT-5.

Overhead: semantic extraction costs about 1,607 input tokens and 350 ms per query (roughly $0.0675 per 1,000 queries at evaluation-time JEV pricing); the downstream CatBoost predictor takes 1.23 ms per query.

Why it matters

Forty human-readable features matching a 7B-parameter embedding model is the headline number for practitioners. Routing evidence becomes auditable: when the router errs, you can trace which semantic judgment was wrong. The factorization is deployable, since objectives and candidate pools can change without re-encoding queries, and new models need only performance calibration. The direct-routing ablation is a useful negative result for anyone tempted to skip the machinery and just ask an LLM which model to use.

Keep the position in view, though. The router sits 1.86 points below the Dataset Oracle and nearly 20 points below the Instance Oracle. Like the leading routers in LLMRouterBench, most of what it captures is dataset-level regularity. This is a solid representation improvement, not a jump in what routing extracts from individual queries.

Limitations

The paper discloses most of these itself:

Two reservations from reading it. JEV, the primary extractor, comes from the authors' own company (Indigma), and the open-weight substitute loses 2.1 points, so the interface ports but the performance does not. And the 350 ms extraction latency per query is a real cost for latency-sensitive serving, which the paper mentions only in passing in the overhead section.

Terms

Source

Related papers

All paper explainers