IRL-based eval sampling cuts compute 20-70% in tests

Tests on 400k code trajectories confirm an IRL-based evaluation sampling method cuts compute costs by 20-70%. It works best with hundreds of tasks run 8-32 times each, suiting metrics like avg@8 or avg@16.

2026-08-28 ~ 2026-08-28 · 2 related posts