Legal model training shifts from SFT to LLM-judge RL

ivan_bezdomny · x · 2026-08-21

It is observed that legal model training has moved from standard SFT to Reinforcement Learning using 'LLM as a judge' for rewards. This trend reflects a shift towards learned reward functions over traditional SFT libraries.

Original post →

More from Research

Research channel →