New paper fixes contrastive RL blind spot by re-weighting InfoNCE with 1-bit failure signal
kastnerkyle · x · 2026-09-06
- A newly recommended contrastive RL (CRL) paper identifies a long-standing blind spot: InfoNCE only draws positive pairs from surviving steps, ignoring the missing probability mass from failure terminations
- This makes the critic delusionally optimistic near death traps, tricking the policy into reckless moves just before dying
- Instead of safety reward hacks, the authors use a basic 1-bit failure signal to re-weight InfoNCE and fold survival mass back into policy updates
- Across 12 benchmark tasks, results are similar on simple Point/Car setups, but the gap on harder Ant and Humanoid pitfall runs with hazards is night and day
- The sharer argues this fixes a problem that has troubled many researchers
More from Research
- Claude completes 13M-line Lean formalization of Fermat's Last Theorem, largest proof ever — burny_tech · 2026-09-06
- Single-metric robustness claims for LLMs can mislead, multi-level arXiv study finds — burny_tech · 2026-09-06
- Prove2Me: the Lean crowdsourcing platform behind Anthropic's Fermat's Last Theorem formalization — burny_tech · 2026-09-06
- Causal Foundation Models: pretrained nets estimate treatment effects in-context, no retraining — burny_tech · 2026-09-06
- Google and HHMI release largest cellular brain map: complete fly CNS with 166,000 neurons — udmrzn · 2026-09-06
- Unreal Engine pipeline yields 8.7K+ hours of action-conditioned video for world-model pretraining — udmrzn · 2026-09-06