LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu
cs.CL
2026-09-30
LexReward scores Chinese legal answers on Style, Element, and Chain, then trains LexRM. On Qwen3-8B, GRPO lifts Legal∆ avg from 60.62 to 66.66; LexChain only from 55.99 to 56.40.
A legal answer can hit the right conclusion while dropping a decisive fact, citing the wrong statute, or dressing a broken inference in fluent judgment-style prose. An outcome reward that checks the charge, the statute, or a sum of money never sees that path. A general-purpose reward model folds style, length, and substance into one number. Which parts of a legal answer deserve reward is still undefined.
LexReward, from Peking University, Tsinghua University, and Northeastern University, writes that definition down before any training. Legal experts split response quality into Style, Element, and Chain. Rubrics turn the dimensions into scores, the scores into preference pairs, and the pairs into LexRM. The paper presents LexRM as the first family of reward models built for Chinese legal text. The same signals then drive DPO and GRPO.
Experts set the dimensions from ordinary requirements of legal writing, then added failure modes they found by comparing model outputs with reference answers.
Style ignores the question. Five attributes, word specificity, subjective-word control, sentence cohesion, sentence structure, and collocation, are scored against 1,000 real judgments from China Judgments Online. Another 986 judgments set the normalization scales. Term distributions and sentiment-word distributions use KL divergence. Conjunction rate uses an absolute gap. Sentence-length statistics and 3- to 6-gram coverage use Euclidean distance. Each gap is divided by its calibration mean, negated, and averaged with equal weight. A higher score means the text looks more like an authentic judgment.
DeepSeek-V4-Flash scores Element on subjects, facts, statutes, and decisions, each from 0 to 1, with no reference answer. The reward is the mean of the four. Chain scores order, completeness, correctness, and non-redundancy. If the task prescribes the steps, rules check them. If the path has to be inferred, the same judge scores it. The experimental prompts are Chinese.
Rubric scores often tie. Five generators, Qwen3-4B, Qwen3.5-4B, Qwen3.5-9B, Llama-3.1-8B-Instruct, and gemma-4-12B-it, each write one answer per query. Pairs are drawn across high, medium, and low scores, and queries where every candidate ties are dropped. Style preferences are built from all 4,000 CLASE training queries, Element yields 2,244 pairs, and Chain assigns 4,000 training instances to preference construction. Qwen3-8B is trained with a Bradley-Terry loss into one LexRM per dimension. Cosine similarity between the three task vectors averages at most 0.03 and never exceeds 0.14, so task arithmetic adds all three updates onto one backbone as LexRM-Merge.
DPO starts from Qwen3-8B. GRPO first runs SFT on 1,000 queries, using the highest-scoring answer for that dimension as the target, then continues from that checkpoint on another 1,000 queries. One run uses LexRM. The other uses a rule: a verifiable final answer for Element and Chain, and ROUGE against the reference for Style. Preferred answers are model generations selected by the rubric. Authentic judgments are not used as training text.
On Style, selection accuracy is pairwise: the scorer must prefer the gold answer over a model-written negative. On Element and Chain, the scorer picks the best of five model answers, and that answer is then scored on Legal∆ or LexChain.
| Setting | Metric | Result |
| Style pairs | Accuracy | LexRM-Style 81.75, rubric 79.00, Skywork-Qwen 78.50, random 50.00 |
| Element, pick 1 of 5 | Legal∆ avg | Skywork-Qwen 68.04, LexRM-Element 66.32, rubric 60.27, five-model mean 49.63 |
| Chain, pick 1 of 5 | LexChain overall | LexRM-Chain 50.46, Skywork-Qwen 50.19, rubric 47.80, five-model mean 45.41 |
| Generation, Style | CLASE-Mix (of 10) | Base 3.08, DPO 5.21, SFT 7.32, GRPO rule 8.43, GRPO+LexRM 8.54 |
| Generation, Element | Legal∆ avg | Base 57.78, DPO 58.13, SFT 60.62, GRPO rule 60.86, GRPO+LexRM 66.66 |
| Generation, Chain | LexChain overall | Base 53.55, DPO 54.47, SFT 55.99, GRPO rule 55.29, GRPO+LexRM 56.40 |
Element under GRPO+LexRM is where the scores separate. All five metrics rise over SFT, and the average rises 6.04, with sentence-length accuracy from 46.80 to 56.70 and financial calculation from 81.80 to 89.80. GRPO with an outcome reward stays at 60.86, and financial calculation falls to 75.80. At selection time on Element, Skywork-Qwen still leads LexRM, 68.04 to 66.32. On Style, the move from 3.08 to 7.32 happens at SFT; LexRM beats ROUGE by 0.11. On Chain, LexRM is 0.41 above SFT, while rule-based GRPO falls below SFT. Dispute focus drops from 39.00 to 37.30, and the statute score of 39.55 sits under the base model's 42.10. In the generation runs, plaintiff identification stays at or above 97.25 and defendant identification stays between 88.90 and 90.85.
Inside each dimension, the aggregated rubric has the best overall score. A subjects-only Element reward drops sentence-length selection to 7.60; the full rubric returns it to 33.10. LexRM-Merge scores 65.77 on Legal∆, close to the Element expert at 66.32. Style accuracy falls from 81.75 to 69.75, and LexChain overall falls from 50.46 to 46.84.
Open legal questions often have no charge or amount to check. Element and Chain are scored without the reference answer, so the reward still exists when no gold judgment is available.
For a team that already has a legal SFT model, the new supervision is mostly Element. Most of the Style score is collected during supervised fine-tuning. The Chain total barely moves. The useful piece is the split, and the dimension worth training against is Element.
The appendix states two limits. The taxonomy and the evaluation are tied to Chinese law and Chinese data; other languages and legal systems are untested. Combining the dimensions into one reward, including the weights, the mix of preference data, and the training recipe, is left open.
DPO is trained from Qwen3-8B. It beats the base model on every dimension and trails the matching SFT run on every dimension. The runs do not share an initialization, so they do not rank preference optimization against supervised fine-tuning. Half of CLASE-Mix compares textual features with real judgments, the same kind of signal the Style reward optimizes. Rubric judging is done by DeepSeek-V4-Flash, LexChain's final scores by GPT-4o, and the subjective half of CLASE by GPT-4o-mini. No correlation with lawyers is reported. On Chain selection, LexRM leads Skywork-Qwen by 0.27.