LoRA Co-inventor Joins Mercor, Releases 397B RL Training Guide

himanshustwts · x · 2026-09-02

Edward Hu, co-inventor of LoRA and muP (ex-OpenAI), has joined Mercor to lead model training and research. Mercor released a detailed guide on post-training Qwen3.5-397B-A17B using reinforcement learning (RL) with SkyRL. Focusing on expert data for knowledge work agents, the project improved Pass@1 on the APEX-Agents benchmark from 16.11% to 27.29% (a 70% relative increase). The full training script, model weights, and evaluation traces have been open-sourced.

Original post →

More from coding & agent

coding & agent channel →