Meta's RL-XAR Method Aims to Eliminate 'AI Slop' Writing
Meta AI researchers led by Jason Weston introduced RL-XAR, a reinforcement learning method using expert-aligned rubrics as rewards to fix low-quality 'AI slop' writing. Trained on Qwen3.5-27B, it achieved 89% and 95% human-evaluation win rates on papers and stories respectively.
2026-09-28 ~ 2026-09-28 · 3 related posts
- Meta claims it can fix AI slop with RL-XAR: RL on expert-aligned rubrics for writing — jaseweston · 2026-09-28
- RL-XAR results: 89% and 95% human win rates on papers and stories, judged across model families — jaseweston · 2026-09-28
1 near-duplicate retellings: jaseweston