Why RL generalizes to reasoning but not literary writing, per AI researchers

phl43 · x · 2026-09-22

phl43 argues that even the latest models, while imperfect on deep conceptual issues, do far better there than on literary work — likely thanks to generalization from RL on easier goals. Doing the same for literature seems genuinely harder and costlier. The practical workaround discussed: rubrics, examples, and step-by-step bootstrapping, which works but is slow going.

Related event: Literary writing remains AI's hardest benchmark; agents need rubrics and guidance(2 posts)→

Original post →

More from Models

Models channel →