RL-optimized model writes news; deterministic reward function replaces LLM judge

ivan_bezdomny · x · 2026-08-21

The author shares progress on optimizing a model with RL for news writing via HuggingNews. A key difference from approaches like Harvey is the design of a deterministic reward function instead of using an LLM as a judge. The author argues that if judgments can be approximated by code, this method is more deterministic and much faster, avoiding the inherent difficulty LLMs face in ranking multiple options and converting them into scores.

Related event: Developer Replaces LLM Judges With Deterministic Reward Functions in RL(2 posts)→

Original post →

More from coding & agent

coding & agent channel →