Meta proposes RL-XAR: expert-aligned rubrics to push LLMs past AI slop in writing

jaseweston · x · 2026-09-28

Meta AI's Jason Weston shares RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics), a training method targeting "AI slop" in non-verifiable tasks: collect top human-written texts, learn LLM rubrics that score experts above model generations, then RL on those rubrics iteratively. Tested on scientific paper sections, Pulitzer-novel continuations, and Wikipedia pages using Qwen3.5-27B with a cross-family judge, showing strong gains evaluated by GPT5-6. Ablations show judge strength, meta-optimizer strength, and expert human baselines all matter.

Related event: Meta's RL-XAR Uses Expert-Aligned Rubrics to Fix AI Slop(4 posts)→

Original post →

More from Models

Models channel →