Meta's RL-XAR Method Aims to Eliminate 'AI Slop' Writing

Meta AI researchers led by Jason Weston introduced RL-XAR, a reinforcement learning method using expert-aligned rubrics as rewards to fix low-quality 'AI slop' writing. Trained on Qwen3.5-27B, it achieved 89% and 95% human-evaluation win rates on papers and stories respectively.

2026-09-28 ~ 2026-09-28 · 3 related posts

1 near-duplicate retellings: jaseweston