Developer Replaces LLM Judges With Deterministic Reward Functions in RL
A developer shared that instead of using LLM judges for RL training as Harvey does, he built deterministic reward functions to compare rollouts—faster and more stable—and applied the approach to RL-optimized news writing.
2026-08-21 ~ 2026-08-21 · 2 related posts
- Skip the LLM judge: a deterministic reward function for comparing RL rollouts — ivan_bezdomny · 2026-08-21
- RL-optimized model writes news; deterministic reward function replaces LLM judge — ivan_bezdomny · 2026-08-21