Meta claims it can fix AI slop with RL-XAR: RL on expert-aligned rubrics for writing
jaseweston · x · 2026-09-28
- Jason Weston (Meta AI) announces RL-XAR (RL with eXpert-Aligned Rubrics), a training method targeting "AI slop" in non-verifiable tasks like writing.
- Pipeline: collect top human-written texts, learn LLM-judge rubrics that score experts above model generations, then run RL against those rubrics, iterating until meta-optimization finds no discernible gap.
- Motivation: pretraining reproduces context quality, and RLHF rewards inherit non-expert annotators' ceiling, so expert-level writing needs a new reward signal.
- Tested on scientific paper sections, Pulitzer-novel continuations and high-quality Wikipedia pages; multiple metrics show large gains (details in the thread's follow-up).
Related event: Meta's RL-XAR Method Aims to Eliminate 'AI Slop' Writing(3 posts)→
More from Research
- Mila hosts EcoHull hackathon Oct 26-27: ML to predict ship hull fouling and cut maritime carbon — Mila_Quebec · 2026-09-28
- A 24-line Python agent that pulls and ranks weekly arXiv papers in 10s — omarsar0 · 2026-09-28
- SmolDataEnvs launches: 5k+ verifiable RL tasks on real Kaggle data, fully open — SergioPaniego · 2026-09-28
- Linear mode connectivity appears on test loss but not train loss, notes researcher Arohan — _arohan_ · 2026-09-28
- John Langford ships polished Vowpal Wabbit 9.11.9, fixing the broken Maven Central release — JohnCLangford · 2026-09-28
- NeurIPS papers: CoT faithfulness metrics are near-random, plus MoE router geometry and LLM belief studies — megamor2 · 2026-09-28