Meituan's ACL Outstanding Paper: GeoRA Solves RLVR Geometric Misalignment

美团技术团队 · wechat · 2026-08-27

The Meituan technical team won the ACL 2026 Outstanding Paper Award for "GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR," proposing a new solution to the geometric misalignment between RLVR (Reinforcement Learning from Verifiable Rewards) and existing Parameter-Efficient Fine-Tuning (PEFT) methods.

Background & Problem

RLVR has become a key paradigm for enhancing LLM reasoning capabilities, but its training cost is high. PEFT methods like LoRA, designed primarily for SFT, lead to suboptimal results, capability forgetting, or even training collapse when applied to RLVR. This is because RLVR updates tend to be sparse and avoid the principal directions of pre-trained weights, a geometric difference LoRA ignores.

GeoRA Method

Experimental Results

Business Application

The method has been deployed in Meituan's AI recruitment Agent (AgenticRL) scenario, achieving full-parameter training results at LoRA costs, with a 12% improvement over LoRA and significantly reduced memory usage.

Original post →

More from Companies & People

Companies & People channel →