I want to try and give a little back about my experience with llm on R9700s, especially given that I finally managed to get my setup running and stable thanks to this forum.
The following is an AI generated summary of my experience:
vLLM 0.23.x / ROCm 7.14 upgrade on 4x Radeon AI PRO R9700: ~50% regression on one model, unaffected on another, but a free +9% via kernel tuning
Setup: Threadripper PRO 9965WX, 4x AMD Radeon AI PRO R9700 (32GB each, gfx1201/RDNA4), Ubuntu 24.04.4 LTS, kernel 7.0.0-28-generic, iommu=pt. Docker Compose stack (vLLM + LiteLLM + Open WebUI). Two models running in production, one per GPU pair (TP=2 each):
Qwen/Qwen3.6-35B-A3B-FP8— MoE, FP8 block-quantizedgoogle/gemma-4-26B-A4B-it— native precision, not quantized
Wanted to know if upgrading from vllm/vllm-openai-rocm:v0.20.2 (bundles ROCm 7.2.1, confirmed via rocm-core inside the container) to AMD’s day-0 ROCm 7.14 image (rocm/vllm:rocm7.14.0_rdna_ubuntu24.04_py3.14_pytorch_2.11.0_vllm_0.23.0, vLLM 0.23.1.dev1) was worth doing — so this is really a ROCm 7.2.1 → 7.14.0 jump, not just a minor vLLM point release. Built a small harness to A/B both models on identical hardware, isolating one variable at a time and repeating runs to average out noise. Sharing what came out of it since it took a while to nail down and might save someone else the time.
TL;DR
- Qwen3.6-35B-A3B-FP8: not worth it. ~45-50% throughput regression on the new image, reproducible, isolated from every config lever we could find.
- gemma-4-26B-A4B-it: no change at all (within a few percent either way). Same image, same host, different model.
- That contrast is actually the most useful finding here: it rules out “the new image is just generally slower on this hardware” and points specifically at the FP8 quantized kernel path.
- Found a free +7-9% for the FP8 model regardless of which image you run, from kernel configs vLLM ships pre-tuned per GPU but doesn’t have for this card yet. Generated our own. Details and where they came from below.
The benchmark
Same GPUs, same TP=2 topology, same model, image swapped. Concurrency 1 and 4, 3 repeats each (this GPU has a documented bimodal-throughput quirk depending on process spawn, so single runs aren’t reliable).
| Model | Old image | New image | Verdict |
|---|---|---|---|
| Qwen3.6-35B-A3B-FP8 | ~77 tok/s (conc=1), ~250 tok/s (conc=4) | ~39 tok/s (conc=1), ~140 tok/s (conc=4) | ~47-49% regression |
| gemma-4-26B-A4B-it | ~98 tok/s (conc=1), ~281 tok/s (conc=4) | ~95 tok/s (conc=1), ~287 tok/s (conc=4) | noise-level, no regression |
Dead ends (ruled out for the Qwen regression, in case it saves someone a rabbit hole)
- RCCL AllReduce protocol (
NCCL_PROTO=Simple, a workaround for a known gfx1201 deadlock) — removing it made no throughput difference. TP=2 also ran clean with zero hangs across every test on both images, so that deadlock bug may itself be fixed in the newer RCCL — but that’s a stability note, not a speed one. - AITER — turns out the
aiterpackage isn’t installed in the new image at all. SettingVLLM_ROCM_USE_AITER=1was a silent no-op for Qwen (no crash, no effect). For gemma specifically it’s not a no-op — it hard-crashes with aValueErrorfrom vLLM’s MoE backend oracle, a different failure mode per model architecture. Worth knowing if you try this on other models. - Attention backend — identical on both images (
ROCM_ATTN, falling back to a Triton paged-decode kernel either way, same log line, same reason). - MoE backend selection — identical (
TRITONFp8 MoE backend on both). fast_moe_cold_startdiffers between the two configs (True/False), but based on the vLLM source comments this is about compile/warmup ordering, not steady-state throughput — doesn’t fit the pattern of a slowdown that persists across many post-warmup requests.- IOMMU passthrough — legit ROCm recommendation for multi-GPU stability, already correctly set on this host, but it’s host-level and identical across every phase tested, so it can’t explain a relative difference between images.
So: not RCCL, not AITER, not attention, not MoE backend selection. Whatever’s actually regressing appears to live inside the FP8 GEMM/MoE Triton kernel codegen itself for this vLLM version — the gemma control group (same everything, just not FP8) reinforces that read, since it shares most of that surrounding machinery and shows no regression at all.
What can be gained: kernel tuning (real win, wrong lever for the regression)
vLLM ships pre-tuned Triton kernel configs (optimal tile sizes) per exact GPU model + matrix shape + dtype, as JSON files bundled in the package. Checked the shipped configs directory on the new image: plenty of entries for AMD_Instinct_MI300X/MI325X, zero for AMD_Radeon_R9700 at the shapes this model needs.
Generated our own using vLLM’s own tuning benchmarks, straight from the repo:
benchmarks/kernels/benchmark_w8a8_block_fp8.py(dense FP8 GEMM) — note: this script is hardcoded for DeepSeek-V3/R1’s weight shapes, not derived from--model. Had to patch itsget_weight_shapes()to return the actual shapes this model needs at TP=2 (read straight off the “Config file not found” warnings in the startup log).benchmarks/kernels/benchmark_moe.py(MoE) — this one is model-aware (--modelflag works as expected), no patching needed for shapes. Did hit an unrelated Ray/env-var bug here (Ray in this vLLM build wantsHIP_VISIBLE_DEVICES, but the script’s own compatibility shim translates it intoROCR_VISIBLE_DEVICESand deletes it, then crashes because Ray now rejects that — had to patch the shim out).
Applying these tuned configs:
| Old image | New image | |
|---|---|---|
| With tuning | +7% (conc=1), +9% (conc=4) | +2-3 percentage points only |
So on the new image, tuning barely moves the needle — confirms it’s not the dominant cause of the regression. But applied to the currently-running production image, it’s a genuine free win with no downside found (tool-calling still verified working, no correctness issues). That’s now deployed via a per-file read-only bind mount over the vLLM install path pointing at the generated JSONs — cheap, reversible, no image rebuild needed.
One gotcha if you try this yourself: the old and new images use different Python/site-packages paths internally (different base images entirely, not just a vLLM bump) — config file destination paths aren’t interchangeable between them.
MoE tuning itself is slow, by the way — full sweep (18 batch sizes) took ~8.5 hours on a single R9700, serial. Needs an otherwise-idle GPU too; running it against a GPU your production server is also using will OOM partway through.
Where this leaves things
Verdict for anyone on similar hardware asking “should I upgrade”: not yet, if you’re running FP8-quantized MoE models — you’ll eat a ~50% regression for whatever the new image’s Day-0 features buy you. If your models aren’t FP8-quantized, the upgrade looks like a wash either way based on our one data point.
Open questions we haven’t chased yet: whether an INT4/GPTQ quantization of the same model sidesteps the FP8-specific kernel path (structurally different kernel dispatch — worth testing before assuming), and rocprof-level profiling to pin down the exact kernel if we do chase this further. Might file this upstream given how clean the isolation ended up being — happy to share the actual scripts/compose files if anyone wants to reproduce or extend this on their own R9700/RDNA4 box.