LoRA Co-inventor Joins Mercor, Releases 397B RL Training Guide
himanshustwts · x · 2026-09-02
Edward Hu, co-inventor of LoRA and muP (ex-OpenAI), has joined Mercor to lead model training and research. Mercor released a detailed guide on post-training Qwen3.5-397B-A17B using reinforcement learning (RL) with SkyRL. Focusing on expert data for knowledge work agents, the project improved Pass@1 on the APEX-Agents benchmark from 16.11% to 27.29% (a 70% relative increase). The full training script, model weights, and evaluation traces have been open-sourced.
More from coding & agent
- NextAdmin Launches: Open-Source Component Kit for Consistent AI Agents — Scobleizer · 2026-09-02
- METATRON: Local LLM-powered automated penetration testing assistant — tom_doerr · 2026-09-02
- Why AI Agents Write Billions of Lines: No Laziness or Time Constraints — intellectronica · 2026-09-02
- AI Agent Beats Civ6 Deity: Fable 5.1 via MCP Shows Promise — xeophon · 2026-09-02
- Script enables seamless switching between multiple Claude Max accounts — AaronBergman18 · 2026-09-02
- GitHub Copilot CLI v1.0.83-2: Adds Multi-Model Fallback, Claude Support — copilot-cli-release-app[bot] · 2026-09-02