KAIST Proposes PLC-DPO to Correct Noisy Preference Labels via Posterior Margins
kaist-ai · hf · 2026-09-14
KAIST AI released PLC-DPO (Posterior Label Correction DPO), which tackles noisy and ambiguous pairwise preference labels.
- Uses calibrated policy-reference margins to route each noisy pair into clean, flipped, or tied cases
- Adds posterior label correction on top of standard DPO for more robust preference optimization
- Code available on Hugging Face
More from Research
- Carl Feynman explains why Navier-Stokes became a Millennium Prize problem — burny_tech · 2026-09-14
- UMass AutoIndex Turns Chunking Into Programs an LLM Writes, Not Configs You Tune — mrdrozdov · 2026-09-14
- OpenAI says its AI agents solved Navier-Stokes blow-up, a Millennium Prize Problem — Chris_Armstrong · 2026-09-14
- Hugging Face engineer builds an interactive visual guide to Flow Matching — aritra_rg · 2026-09-14
- Tencent Hunyuan Open-Sources SAS: End-to-End Sparse Attention via Context Ranking — Tencent-Hunyuan · 2026-09-14
- NUS releases LIT to break vision-action shortcuts in robot foundation models — NationalUniversityofSingapore · 2026-09-14