Zero Gap Is Not Restoration: SA-PPG Metric and RailCap for Benchmark Contamination
zju · hf · 2026-08-10
A paper from ZJU argues that current metrics for evaluating benchmark contamination mitigation (G-AP) are fundamentally flawed, as over- and under-suppression can cancel each other out.
The authors propose SA-PPG (Stratified Aggregate of Per-question Probability Gaps): it estimates each question's solve probability by sampling, differences it against a clean model per question, and aggregates within groups.
They also introduce a novel mitigation strategy called RailCap. Instead of pre-estimating contamination, RailCap judges it dynamically during generation: if a sample falls back onto the greedy trajectory, it caps the next token to the runner-up until the response distribution is sufficiently dispersed. Experiments reveal that prior strategies substantially overestimate restoration, while RailCap achieves the lowest SA-PPG.
More from Research
- Study: LLMs Write More Correct Code in JS Than TS, Questioning Type Guardrails — hichaelmart · 2026-08-10
- 269 Hours of EEG Data: Non-Invasive Speech Decoding Hits 61.3% Accuracy — kaixhin · 2026-08-10
- Why Speculative Decoding Exploded: Tri Dao's Paper Fuels an Inference Revolution — Ok-River5924 · 2026-08-10
- Training Humanoid Robots in 3D Scans: Zero Real-World Fine-Tuning — lukas_m_ziegler · 2026-08-10
- MOSS-Transcribe-Diarize: Transcription and Diarization in One 0.9B Model — vanstriendaniel · 2026-08-10
- Harvard and MIT Open-Source MatrAIx: Simulating 8.3B Global Personas — SRSchmidgall · 2026-08-10