AWS Walkthrough: SkyRL GRPO on HyperPod Lifts Qwen3-VL Maze Solve Rate From 43.75% to 95%+

AWS ML Blog · rss · 2026-09-26

AWS ML Blog details a full multimodal RL post-training walkthrough: running the open-source SkyRL framework with GRPO on SageMaker HyperPod to train Qwen3-VL-8B for visual maze navigation, improving solve rate from 43.75% to over 95% on a 64-maze eval set. The setup colocates FSDP policy shards and vLLM rollout engines on 6 Blackwell GPUs, syncs LoRA adapters via FSx for Lustre, and ships a reproducible Dockerfile.

Original post →

More from Infra

Infra channel →