RouteFM pretrains LLM routing once for anywhere: +2.23 quality points over strongest baseline
nanjinguniv · hf · 2026-10-02
LAMDA (Nanjing University) proposes RouteFM, shifting LLM routing from repeated local fitting—optimized for one query workload and candidate pool—toward a foundation-model paradigm: pretrain a reusable routing capability once, then generalize across tasks, candidate models, and deployment conditions.
How it works:
- Characterizes anonymous candidate models from behavioral context and infers target-specific capabilities, rather than binding routing to fixed model identities.
- Episodic pretraining across heterogeneous routing environments; the frozen router adapts to new environments via context alone.
- Code open-sourced: github.com/LAMDA-Model-Reuse/RouteFM
Experiments show transfer across domains, modalities, candidate pools, and context budgets, with the largest gains when behavioral evidence is scarce: on held-out MMR-Bench, RouteFM beats the strongest baseline by 2.23 quality points with only eight observations per candidate.
More from Infra
- Morgan Stanley: Meta won't buy new chips to scale its Muse agent — AccBalanced · 2026-10-02
- Cerebras insiders dump stock as shares plunge $18 in a day, no word on lost GPT-6.1 deal — firstadopter · 2026-10-02
- Data Center Bottleneck Isn't GPUs: Transformer Lead Times Hit 115 Weeks — AccBalanced · 2026-10-02
- MLPerf Adds DLRMv4: HSTU Sequence Modeling Meets 560GB Embeddings for Production Recommenders — TheKanter · 2026-10-02
- Cerebras-Linked Team Launches Detailed Series Explaining Disaggregated Inference — AccBalanced · 2026-10-02
- Fei-Fei Li now with Team AMD; CoreWeave keynote: "No one knows the right way" to build on model APIs — AccBalanced · 2026-10-02