AgileRL Arena v1.0: manifest-driven RL training with LoRA/GRPO finetuning

Balance- · reddit · 2026-09-11

AgileRL Arena v1.0 ships as an independent PyPI SDK/CLI that runs RL and LLM finetuning (LoRA + GRPO/SFT/preference) from Pydantic-validated YAML manifests, locally or on a managed cluster. It adds evolutionary HPO via population mutation, stricter manifest validation, and automatic exclusion of Mamba layers unsupported by PEFT 0.20.

Original post →

More from coding & agent

coding & agent channel →