AMD’s ROCm blog says verl 0.9.0 now runs on ROCm 10.0 as release/0.9.0.amd0. The same post tells you to build docker/rocm/Dockerfile.rocm and expect vLLM 0.27.0 and SGLang in that image. TechCrunch the same week said AMD is paying $8.2 billion for Fei-Fei Li’s World Labs. Those are not the same stack. One is a post-training container you can fail to schedule in Ray. The other is a world-model lab moving under a chip vendor.
We already wrote Transformers loading GGUF while llama.cpp keeps the no-Python box and vLLM as 39 percent of the speed of light. This week is the AMD fork of ByteDance’s verl, plus a laptop YAML trainer named Soup that still claims 8B on 4GB. If your mental model of “Python RL” is a Colab, you are not the audience for MI355X 4+4 splits. You might still be the audience for a config key that now errors on typo.
The branch name is the product
AMD-Ecosystem/verl, branch release/0.9.0.amd0. Host driver ROCm 10.0. Image from Dockerfile.rocm. Inside: vLLM 0.27.0, SGLang, Megatron-Core 0.18. Upstream 0.9.0 features the blog lists: V1 trainer as default, Megatron-Bridge, delta-sharded weight sync, Qwen3.5 GDN, DeepSeek-V4-Flash GRPO. PlatformROCm is native detection: HIP visibility, Ray GPU discovery, SGLang/vLLM ROCm defaults in the tree instead of out-of-tree patches.
The previous AMD walkthrough, “Scaling RL with verl on AMD Instinct MI355X,” is still the compatibility target. Fully async RL. Separate rollout and training fleets. RCCL weight sync. GRPO and DAPO. Same 4+4 scripts, same staleness / sync / partial-rollout knobs. What changes is the container, the driver, and a pile of bring-up fixes. Old pain the post names: Ray reporting zero GPUs, a missing rollout engine. If you have lived that error on a new image, you do not need it explained.
The 10.0 image pins SGLang to Triton attention and a patched non-AITER RMSNorm path. That sentence is a landmine if you swap libraries because a README on GitHub looked newer. Pin the image. Do not “just pip install” vLLM over it and assume 0.27.0 is 0.27.0 everywhere. We compared vLLM, TensorRT-LLM, and ONNX Runtime as engines. Here vLLM is a dependency frozen next to SGLang for rollouts. Different job.
MI300 and MI350 series are the hardware sentence. MI355X is the named walkthrough GPU. If you have a consumer Radeon, this blog is not your fine-tune guide. If you have Instinct and a Ray cluster that still thinks CUDA is the only platform, PlatformROCm is the patch you wanted last quarter.
I am not going to pretend I reran their GRPO job. The useful copy is: keep the scripts, change the image, stop losing a day to Ray seeing zero devices. That is operations. It is still Python ML. verl is a PyTorch RL post-training library. The AMD fork is how you run it without waiting for every upstream PR to notice HIP.
Soup is the other extreme, and it fails closed on typos
Soup’s GitHub is the anti-cluster. Search-week copy: fine-tune from one YAML, layer streaming, 8B on a 4GB laptop GPU. The README in the extract is messier and more useful. Unknown config keys used to validate and get discarded. v0.74 warned with a did-you-mean. v0.75 refuses to load. A typo like quantizaton no longer silently trains the wrong run.
Apache-2.0. Telemetry off unless SOUP_TELEMETRY=1. Maintainer says the project is built on a single 4GB laptop and the numbers are measured. Install extra: "soup-cli[train]" with double quotes so cmd.exe, PowerShell, bash, and zsh all survive the extras syntax. Light install still gets soup init and soup data. Donations go through Stripe as MePlay, Inc., which is the name on the statement.
Put Soup next to verl only as a scale check. verl 0.9.0.amd0 is MI-class async RL with vLLM in the cluster image. Soup is a YAML file on a machine that cannot run Cinebench 2024 if it were a laptop with 8GB of system RAM, except here the constraint is 4GB of VRAM. If you paste a verl Hydra config into Soup you will get a rejected key, which is the correct outcome.
I do not have Soup’s measured tokens-per-second in the extract. I will not invent them. What I will copy is the policy: unknown keys fail the run. That is a better default than most training CLIs. verl’s own config surface is large. The AMD post is telling you which knobs stayed compatible (async 4+4) and which substrate moved (ROCm 10, V1 trainer). Read both layers. Do not assume 0.7.1 scripts are 0.9.0 scripts just because GRPO is still named GRPO.
If your actual hardware is a 4GB laptop, you are not enabling Dockerfile.rocm. You might still read the AMD post to know what “real” post-training looks like when money exists. Then you go back to YAML. There is no shame in that split. There is shame in claiming a laptop run is MI355X async.
World Labs is the roadmap, not your import
TechCrunch: AMD acquires World Labs for $8.2 billion. Fei-Fei Li becomes EVP and chief scientist. The companies already had an inference-and-training partnership. Li was at AMD’s CES. She founded World Labs in 2024 after ImageNet. First product Marble is pitched for entertainment experiences and for simulated environments for robot training. Li’s post, as quoted: proof exists, now scale, get closer to the hardware. Nvidia already has open-weight world models like Cosmos. AMD, in that piece, had been stronger on text- and video-side models than on that world-model shelf.
None of that changes pip show verl this week. It changes why AMD is willing to staff ROCm enablement for a ByteDance RL library. Chip vendors fund the boring Dockerfiles when they want the next workload to land on their device. World models and RL post-training are the workloads. Your import path is still release/0.9.0.amd0.
Do not write a grant that says “we will use Marble inside verl.” Nobody in these three sources said that. Marble is a World Labs product. verl is a post-training framework. They may share a badge later. Today they share a buyer.
If you are choosing NVIDIA versus AMD for a new RL cluster, the $8.2 billion is strategy. The Dockerfile is evidence. Evidence wins. Can Ray see the GPUs. Does vLLM 0.27.0 serve the rollout. Does delta-sharded weight sync actually move. Those tests are cheaper than a take on Li’s title.
What actually changed since 0.7.1 on the AMD tree
The blog frames 0.9.0 against the last public AMD enablement, not against a generic PyPI wheel. Upstream 0.9.0 landed V1 trainer, Megatron-Bridge, vLLM 0.27, delta-sharded weight sync, DeepSeek-V4-Flash GRPO, Qwen3.5 GDN. The AMD tree then had to make those paths exist on HIP. PlatformROCm is how HIP visibility and Ray discovery stop being a fork you maintain in a gist.
If you last ran 0.7.1-era scripts, the trainer default is the silent break. V1 trainer as default means a flag you never set is now a different class. Read the 0.9.0 notes before you blame ROCm for a loss curve that moved. Megatron-Bridge is the other silent break if you were on a hand-rolled Megatron path. Core 0.18 is pinned in the image. Do not pip a newer Core over it on Friday because Twitter said so.
Delta-sharded weight sync is the async RL sentence. Separate rollout and training fleets only work if weights move. RCCL was already in the MI355X write-up. 0.9.0 adds a sharded delta path. I do not have a bandwidth chart in the extract. I have the feature name. Treat it as “try this if full sync is the stall,” not as a guaranteed speedup.
Qwen3.5 GDN and DeepSeek-V4-Flash GRPO are named so you know the model zoo moved. If your job is still a 7B you trained in July, you do not have to touch those paths. If your job is “run the new Flash GRPO recipe on MI350,” you do. The blog is a starting point for that recipe on Instinct, not a paper on the algorithm.
SGLang’s Triton attention pin and the patched non-AITER RMSNorm are the kind of lines that make a cluster non-reproducible when someone “upgrades” SGLang inside the container. Snapshot the image digest. Put the digest in the run.yaml next to the verl branch. Soup, for all its laptop energy, at least fails the run when a key is unknown. Your cluster should fail the run when the digest drifts. That is the same philosophy at two budgets.
Ray reporting zero GPUs is still the first ticket. PlatformROCm is supposed to make that ticket boring. If it is not boring on your node, you are not on the image they published, or the host driver is not 10.0, or the device is not visible to HIP before Ray starts. Fix visibility. Do not rewrite the trainer.
What to pin in the repo this week
Pin the branch name release/0.9.0.amd0, the file docker/rocm/Dockerfile.rocm, vLLM 0.27.0, Megatron-Core 0.18, ROCm 10.0 on the host. Write down that V1 trainer is the upstream default now. Keep the MI355X 4+4 scripts if they already ran. If Ray reports zero GPUs on a new node, read the PlatformROCm paragraph before you rewrite the cluster.
If you are not on Instinct, pin nothing from that Dockerfile except as documentation. Use Soup’s fail-on-unknown-key behavior as the laptop lesson, and keep your training YAML in git. Quote the extras string correctly. Do not copy 'soup-cli[train]' with single quotes from an old gist.
Inference engine choice is still the old post. Training-time rollout engine in this AMD image is vLLM 0.27.0 plus SGLang with a Triton attention pin. That pin is the kind of line that vanishes from a Slack screenshot. Put it in the runbook.
If you are evaluating Soup as a teaching tool, keep the 4GB claim in the same paragraph as the v0.75 fail-closed keys. A class that trains 8B from YAML on a laptop is a different course from a class that schedules GRPO on MI350. Do not cross-list them as “RL in Python” without naming the GPU. Students will believe the slide.
World Labs can wait until there is an API or a weight. verl 0.9.0 cannot wait if your cluster is already mid-upgrade. The blog is dated around the same $8.2 billion news cycle for a reason. You still have to build the image.
Discussion
Leave a comment
No comments yet
Be the first to start the conversation.