Miles v0.1: Open-Source Production-Grade RL Post-Training System with LoRA and Diffusion Support
RadixArk · hf · 2026-09-09
RadixArk released Miles v0.1 on Hugging Face, an open-source, production-ready system for large-scale reinforcement learning and post-training.
Key features:
- Multiple training backends
- Weight synchronization
- LoRA support
- Built-in distillation
- Diffusion model support
The project targets production environments and offers ready-made infrastructure for teams running large-scale RL post-training.
More from Research
- 2-billion-line Lean proofs: mathematicians debate proof without understanding — Singularitarian · 2026-09-09
- A digital fruit fly brain can play Beat Saber — cixliv · 2026-09-09
- KVMem pages agent context overflow to KV state, beating compaction on DeepSWE — omarsar0 · 2026-09-09
- Collision attacks on SHA-2 pushed to 39 steps in new paper — jedisct1 · 2026-09-09
- 20,000-qubit trapped-ion quantum computer could break 256-bit ECC in 26 days — jedisct1 · 2026-09-09
- Open Optimization Challenge: Cut TensorFrost Wave Equation Runtime Below 1e-6 Max Error — Michael_Moroz_ · 2026-09-09