Frozen-base 34M logit correction module fixes 53.3% of Gemma errors losslessly
eulogik · hf · 2026-09-29
Can a small correction module fix a frozen language model's errors without degrading base capabilities? CRN v2 answers partially yes.
Key findings:
- CRN v2 is a 34M-parameter (0.73% of the 4.65B text module) logit-level correction module atop a fully frozen Gemma 4 E2B, trained via SFT plus reference-free DPO on 83,400 error-correction pairs;
- It corrects 53.3% of base-model errors on a 60-question domain exam with no degradation on MMLU/BoolQ benchmarks;
- A matched-budget LoRA baseline achieves 83.3% correction but loses 30–75% capability, illustrating the correction-capability tradeoff;
- Ablations show the KL preservation term (λ=0.1) is critical (dropping it to 0.01 cuts correction to 35.0%); hidden-state injection, multi-depth, and longer training variants all underperform;
- Code, weights, and eval scripts released.
More from Research
- Tsinghua humanoid robot plays badminton with one policy from just 30 min of human motion data — ChongZzZhang · 2026-09-29
- Open-source xvr AI aligns live X-ray with 3D CT at submillimeter accuracy, published in Nature — Dr_Alex_Crimi · 2026-09-29
- Simulating Human Consciousness: Paper Maps a New Frontier for AI and Robotics — ugail · 2026-09-29
- Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds — gaganghotra_ · 2026-09-29
- 176.9B MoE squeezed to ~1.89 effective bpw: GSQ-RCO GGUFs run Coder build in 29.6GB — Loginhe · 2026-09-29
- GRPO with a judge model biases toward longer answers — maybe why LLMs write essays to simple prompts — djcows · 2026-09-29