IRL-based eval sampling cuts compute 20-70% in tests
Tests on 400k code trajectories confirm an IRL-based evaluation sampling method cuts compute costs by 20-70%. It works best with hundreds of tasks run 8-32 times each, suiting metrics like avg@8 or avg@16.
2026-08-28 ~ 2026-08-28 · 2 related posts
- Testing IRL evaluation method saves 20-70% compute cost — xeophon · 2026-08-28
- IRL evaluation method requires multi-task, multi-run setup — xeophon · 2026-08-28