LLMs lack artistic intent, and writing may resist outcome-based RL
teortaxesTex · x · 2026-09-22
teortaxesTex argues 'style' isn't the bottleneck: pretraining on code improved all tasks, and R1's math/code RL made it a far better writer than V3, so LLMs are good style imitators. The real gap is artistic intent — a task that may be hostile to outcome-based reward, unlike the logically structured training signals that generalize reasoning.
More from Models
- Average users aren't throwing frontier models at open math problems, dev observes — felpix_ · 2026-09-22
- Zero-day hits Meta's Muse for Mac: local process can steal prompts, auth tokens and file access — MicahBerkley · 2026-09-22
- Grok 4.7 spotted in user chatter as users hope for usage reset — BWay124 · 2026-09-22
- Amassing a PhD team is exactly what OpenAI did, dev notes in AI research debate — felpix_ · 2026-09-22
- Dev argues prompting alone can't get AI to solve natural science problems — felpix_ · 2026-09-22
- Dev claims benchmarks are 'absolutely meaningless' — models only differ by vibe — gnukeith · 2026-09-22