Mistral 新论文:特权价值函数让 LLM 强化学习 critic 更强

jm_alexia · x · 2026-08-19

Mistral AI 联合 Mila、蒙特利尔大学发布论文 Le Critique: Privileged Value Functions for LLM Reinforcement Learning,针对 LLM 强化学习中价值函数被严重低估的问题提出两项改进:

论文与代码均已开源。

原文链接 →

「研究」频道最新

更多「研究」频道 AI 资讯 →