Pin Tool Stubs Before Trusting a Checkpoint Delta
conifer_v11 · reddit · 2026-08-26
When comparing model checkpoints, if the harness shuffles tool order or introduces retries, your comparison might be measuring the harness differences rather than the model's capabilities.
Recommendations before evaluation:
- Freeze the tool schema: Ensure tool definitions are identical.
- Lock the seed: Eliminate randomness.
- Pin the compact policy: Keep policy logic consistent.
Only by observing Eval results with these external variables fixed can you ensure you are scoring the model, not two different harnesses.
Related event: Model Checkpoints Must Be Evaluated Together with Their Test Harness(2 posts)→
More from Research
- Study finds Agent harness impacts benchmark scores more than the model itself — rohanpaul_ai · 2026-08-26
- NeurIPS mechanistic interpretability workshop extends deadline — ninamiolane · 2026-08-26
- NeurIPS Findings track deadline extended to Sept 7th — ninamiolane · 2026-08-26
- Nonprofit Sophron Research Launches to Develop AI Model Evaluations — ryan_t_lowe · 2026-08-26
- NeurIPS workshop CFP: Interpreting Agent Behavior — mdredze · 2026-08-26
- Reasoning Models Outperform via Higher Recovery Rates, Not Just "More Thinking" — Jeande_d · 2026-08-26