Legal model training shifts from SFT to LLM-judge RL
ivan_bezdomny · x · 2026-08-21
It is observed that legal model training has moved from standard SFT to Reinforcement Learning using 'LLM as a judge' for rewards. This trend reflects a shift towards learned reward functions over traditional SFT libraries.
More from Research
- AI Economic Indicators: Measuring consumer surplus via willingness to accept — soumitrashukla9 · 2026-08-21
- Research: Robot foundation models are few-shot learners — scaling01 · 2026-08-21
- Skip the LLM judge: a deterministic reward function for comparing RL rollouts — ivan_bezdomny · 2026-08-21
- Arc Institute Launches 2026 Virtual Cell Challenge with $100K Prize — AllThingsApx · 2026-08-21
- Analytical Connectionism Summer School 2026 Announced in Gothenburg — SaxeLab · 2026-08-21
- GitHub CEO shares his cancer case: AI-driven personalized medicine is now viable — garrytan · 2026-08-21