AWS described the procedure for training the vision-language model Qwen3-VL-8B using RL on SageMaker HyperPod
The AWS Machine Learning Blog described how to use the open-source framework SkyRL and the GRPO method to train the model Qwen3-VL-8B (Alibaba) on SageMaker HyperPod infrastructure. The success rate of solving the maze increased from 43.75% to more than 95%.
Amazon published a guide on the AWS Machine Learning Blog on how to train the vision-language model Qwen3-VL-8B (from Alibaba) using the open-source reinforcement learning framework SkyRL on Amazon SageMaker HyperPod infrastructure. The model's task was to navigate a visual maze based on images. Training started from a VisGym SFT (supervised fine-tuning) checkpoint, and using the Group Relative Policy Optimization (GRPO) method, the success rate of solving the maze on a fixed set of 64 test mazes increased, according to the company, from 43.75% to more than 95%.
SageMaker HyperPod runs on Amazon EKS and, according to the company, provides automatic monitoring of node health and their replacement in case of hardware failure, combined with checkpointing, so that a long training run continues from the last saved step after a failure instead of starting from the beginning. According to the description, Ray clusters can be launched directly from SageMaker Studio, and training metrics are displayed in pre-built Grafana dashboards via the HyperPod Observability add-on. In the described setup, the cluster consisted of six GPUs across three worker nodes of type ml.g7e.12xlarge (NVIDIA RTX PRO 6000 Blackwell) and one CPU head node ml.r5d.16xlarge with 512 GB of RAM.
The model was trained with a LoRA adapter (rank 32) split across GPUs using Fully Sharded Data Parallel (FSDP). The SkyRL framework places episode generation (rollouts) via the vLLM engine on the same GPUs as the training itself, and the updated LoRA adapter weights are synchronized via shared storage Amazon FSx for Lustre. According to the description, GRPO compares multiple attempts at solving the maze within a single starting position against the group average, which means it does not need a separate critic model — suitable for tasks with sparse reward and multiple steps.
The source article is a longer technical guide and part of it was not available; see the source article for details.
Why it matters
The guide gives ML engineers a concrete reference procedure for multi-turn reinforcement learning with sparse reward on multimodal models, including cluster configuration, LoRA adapter synchronization, and recovery after a node failure. At the same time, it is a demonstration of the capabilities of Amazon's SageMaker HyperPod product when training models from another provider (Alibaba), which is relevant for decisions about choosing infrastructure for RL post-training.
Two audiences, two different impacts
What this means
For individuals
Anyone preparing multi-turn RL training of multimodal models gets a ready-made technical blueprint — a specific combination of SkyRL, GRPO, LoRA, FSDP and vLLM, as well as recommended instance types.
For a business
Companies considering training their own vision-language models using reinforcement learning have access to a documented example of cluster and tooling configuration on SageMaker HyperPod infrastructure, which makes it easier to estimate infrastructure requirements when planning similar projects.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.