RouteFM pretrains LLM routing once for anywhere: +2.23 quality points over strongest baseline

nanjinguniv · hf · 2026-10-02

LAMDA (Nanjing University) proposes RouteFM, shifting LLM routing from repeated local fitting—optimized for one query workload and candidate pool—toward a foundation-model paradigm: pretrain a reusable routing capability once, then generalize across tasks, candidate models, and deployment conditions.

How it works:

Experiments show transfer across domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is scarce: on held-out MMR-Bench, RouteFM beats the strongest baseline by 2.23 quality points with only eight observations per candidate.

Original post →

More from Infra

Infra channel →