Huawei Noah's Tail-Influence Sampling Cuts CVaR Policy Evaluation MSE by Up to 76%
huawei-noah · hf · 2026-10-06
Huawei Noah's Ark Lab introduces Tail-Influence Sampling (TIS) for CVaR policy evaluation: policies with similar mean returns can differ sharply in rare failures, yet accurate lower-tail estimation usually demands many costly rollouts.
Method
- Derives a "tail influence" per queryable conditional law aggregating its uncertainty's effect on CVaR, yielding the oracle Neyman allocation.
- TIS estimates influences from a pilot model and reallocates budget toward kernels that matter most for the tail; a visitation-anchored variant guards against pilot underallocation.
Results
- TIS attains oracle asymptotic variance; the anchored variant stays within 2x of oracle.
- On CliffWalking, MSE drops 41% vs learned occupancy and 76% vs full rollouts at equal budget.
- In frozen LLM review workflows, anchored TIS wins 23 of 24 MMLU-Pro settings and achieves 2.4-3.4x lower MSE than rollouts on six-call FinQA reviews.
More from Research
- Cambridge team's new work on pullback geometry for data on mixtures of manifolds — skoularidou · 2026-10-06
- CMU L3 Lab brings GradAlign RL data selection and sim2real papers to COLM 2026 — wellecks · 2026-10-06
- A statistical framework for LLM watermarks: optimal detection rules via hypothesis testing — weijie444 · 2026-10-06
- CMDB-1500: open-source multimodal benchmark with 1,500 decision-making tasks — Powerful_Buy_4616 · 2026-10-06
- 4DCodeBench: new benchmark tests coding agents on reconstructing dynamic 3D scenes from video — CSProfKGD · 2026-10-06
- EMNLP paper: LLMs learn novel tasks more reliably from rules than from in-context examples — najoungkim · 2026-10-06