kalomaze: 'Fable 5' Partly Suffers from Undercooked Post-Training
kalomaze · x · 2026-09-11
RL researcher kalomaze suggests part of 'fable 5's' problems come from undercooked post-training. Citing repligate's observation, he notes that holding the RL process roughly constant, more capable models come through relatively less scathed — if a model can straightforwardly one-shot a problem, it doesn't need to game the grader, reducing reward-hacking-style distortion.
More from Models
- DeepSeek 4.1 flash reportedly uses large ngram embeddings, echoing Qwen4 architecture — ccerrato147 · 2026-09-11
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Assistant Benchmark goes live: 61 assistants scored across 15 real-use dimensions — Scobleizer · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11