MetroLLM-Bench shows small fine-tuned models can match larger LLMs on transit-kiosk tasks
continker · hf · 2026-09-11
MetroLLM-Bench is a new benchmark evaluating LLMs as transit-kiosk runtimes, focusing on structured tool-use and fare-quoting with policy reasoning.
Key finding: small, parameter-efficiently fine-tuned models can match much larger models on these structured tasks, suggesting vertical scenarios don't require large general-purpose models.
More from Research
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11
- EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026 — jiqizhixin · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11