SCATR: a lightweight calibrated scorer finds the right LLM answer cheaply
pliang279 · x · 2026-09-30
SCATR (Simple Calibrated Test-Time Ranking), accepted to the COLM 2026 Efficient Reasoning Workshop, trains a lightweight scorer over an LLM's own representations using a small calibration set to rank candidate answers — no heavyweight verifier needed. As test-time compute and agentic systems generate more candidates, cheap answer selection becomes the bottleneck; empirically this simple recipe meaningfully improves best-of-n on coding and math.
More from Research
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30
- François Fleuret: AI math will dwarf human math, focus on lean proofs not explainability — francoisfleuret · 2026-09-30
- SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning — burkov · 2026-09-30
- UMass professor Luc Rey-Bellet's stochastic processes lecture notes on Markov chains and MCMC — michaelchchoi · 2026-09-30
- VoxMem benchmark: none of 15 audio LLMs top 40% on spoken multi-session memory — unimelb-hf · 2026-09-30
- EpiCon: shared multimodal memory lifts agent scores 1.7-4.9 points across 11 benchmarks — Ziyun Zeng · 2026-09-30