Prime Intellect 开放后训练全栈,Extropic 百步 RL 大幅提升 Qwen3.6
Prime Intellect 与量子计算公司 Extropic 合作,展示了完整的定制后训练流程:Extropic 负责设计任务与奖励,Prime Intellect 提供 RL 基础设施。基于该平台,Extropic 对 Qwen3.6-35B-A3B 进行约 100 步强化学习后训练,用于热力学机器学习研究,使评测成绩提升近三倍,验证了开放后训练全栈在定制 RL 模型上的效果。
2026-10-02 ~ 2026-10-02 · 2 条相关
- Prime Intellect 开放后训练全栈:Extropic 借此定制 RL 模型 — beffjezos · 2026-10-02
- Extropic 用 Prime Intellect 训练 Qwen3.6-35B,约百步 RL 使评测成绩近三倍 — MarvinTBaumann · 2026-10-02