LoRA Creator Edward Hu Publishes Guide on Post-Training Open-Source Models with RL
iamrobotbear · x · 2026-09-05
Edward Hu, the creator of LoRA fine-tuning, has published a blog post on how to post-train open-source models, drawing strong recommendations from practitioners including Instacart CEO Brendan Foody.
Those sharing it call it a must-read — reportedly one of the best open-source writeups on actually running RL training for knowledge-work agents at scale.
For anyone post-training open models, first-hand guidance from the technique's original author makes this especially worth reading.
More from coding & agent
- Tinker releases RL recipe for training calibrated future-event forecasting models — simonguozirui · 2026-09-05
- Grok Build v1.0.19 ships background loops, worktree support, now powered by Grok 4.6 — elonmusk · 2026-09-05
- OpenAI unveils GPT-6 Astra, a computer-use agent that does anything on your PC — OpenAIDevs · 2026-09-05
- PSAISuite: one string swap in -Model runs any new model, unchanged for months — dfinke · 2026-09-05
- YC-backed Pentagon teases cloud-native multiplayer rebuild for long-running AI agents — edgarpavlovsky · 2026-09-05
- Full-day coding with Fable 5.1 couldn't exhaust a Max 5x quota — here's why — joshwhiton · 2026-09-05