vLLM 0.23.x / ROCm 7.14 upgrade on 4x Radeon AI PRO R9700: ~50% regression on one model, unaffected on another, but a free +9% via kernel tuning

I want to try and give a little back about my experience with llm on R9700s, especially given that I finally managed to get my setup running and stable thanks to this forum.

The following is an AI generated summary of my experience:

vLLM 0.23.x / ROCm 7.14 upgrade on 4x Radeon AI PRO R9700: ~50% regression on one model, unaffected on another, but a free +9% via kernel tuning

Setup: Threadripper PRO 9965WX, 4x AMD Radeon AI PRO R9700 (32GB each, gfx1201/RDNA4), Ubuntu 24.04.4 LTS, kernel 7.0.0-28-generic, iommu=pt. Docker Compose stack (vLLM + LiteLLM + Open WebUI). Two models running in production, one per GPU pair (TP=2 each):

  • Qwen/Qwen3.6-35B-A3B-FP8 — MoE, FP8 block-quantized
  • google/gemma-4-26B-A4B-it — native precision, not quantized

Wanted to know if upgrading from vllm/vllm-openai-rocm:v0.20.2 (bundles ROCm 7.2.1, confirmed via rocm-core inside the container) to AMD’s day-0 ROCm 7.14 image (rocm/vllm:rocm7.14.0_rdna_ubuntu24.04_py3.14_pytorch_2.11.0_vllm_0.23.0, vLLM 0.23.1.dev1) was worth doing — so this is really a ROCm 7.2.1 → 7.14.0 jump, not just a minor vLLM point release. Built a small harness to A/B both models on identical hardware, isolating one variable at a time and repeating runs to average out noise. Sharing what came out of it since it took a while to nail down and might save someone else the time.

TL;DR

  • Qwen3.6-35B-A3B-FP8: not worth it. ~45-50% throughput regression on the new image, reproducible, isolated from every config lever we could find.
  • gemma-4-26B-A4B-it: no change at all (within a few percent either way). Same image, same host, different model.
  • That contrast is actually the most useful finding here: it rules out “the new image is just generally slower on this hardware” and points specifically at the FP8 quantized kernel path.
  • Found a free +7-9% for the FP8 model regardless of which image you run, from kernel configs vLLM ships pre-tuned per GPU but doesn’t have for this card yet. Generated our own. Details and where they came from below.

The benchmark

Same GPUs, same TP=2 topology, same model, image swapped. Concurrency 1 and 4, 3 repeats each (this GPU has a documented bimodal-throughput quirk depending on process spawn, so single runs aren’t reliable).

Model Old image New image Verdict
Qwen3.6-35B-A3B-FP8 ~77 tok/s (conc=1), ~250 tok/s (conc=4) ~39 tok/s (conc=1), ~140 tok/s (conc=4) ~47-49% regression
gemma-4-26B-A4B-it ~98 tok/s (conc=1), ~281 tok/s (conc=4) ~95 tok/s (conc=1), ~287 tok/s (conc=4) noise-level, no regression

Dead ends (ruled out for the Qwen regression, in case it saves someone a rabbit hole)

  • RCCL AllReduce protocol (NCCL_PROTO=Simple, a workaround for a known gfx1201 deadlock) — removing it made no throughput difference. TP=2 also ran clean with zero hangs across every test on both images, so that deadlock bug may itself be fixed in the newer RCCL — but that’s a stability note, not a speed one.
  • AITER — turns out the aiter package isn’t installed in the new image at all. Setting VLLM_ROCM_USE_AITER=1 was a silent no-op for Qwen (no crash, no effect). For gemma specifically it’s not a no-op — it hard-crashes with a ValueError from vLLM’s MoE backend oracle, a different failure mode per model architecture. Worth knowing if you try this on other models.
  • Attention backend — identical on both images (ROCM_ATTN, falling back to a Triton paged-decode kernel either way, same log line, same reason).
  • MoE backend selection — identical (TRITON Fp8 MoE backend on both).
  • fast_moe_cold_start differs between the two configs (True/False), but based on the vLLM source comments this is about compile/warmup ordering, not steady-state throughput — doesn’t fit the pattern of a slowdown that persists across many post-warmup requests.
  • IOMMU passthrough — legit ROCm recommendation for multi-GPU stability, already correctly set on this host, but it’s host-level and identical across every phase tested, so it can’t explain a relative difference between images.

So: not RCCL, not AITER, not attention, not MoE backend selection. Whatever’s actually regressing appears to live inside the FP8 GEMM/MoE Triton kernel codegen itself for this vLLM version — the gemma control group (same everything, just not FP8) reinforces that read, since it shares most of that surrounding machinery and shows no regression at all.

What can be gained: kernel tuning (real win, wrong lever for the regression)

vLLM ships pre-tuned Triton kernel configs (optimal tile sizes) per exact GPU model + matrix shape + dtype, as JSON files bundled in the package. Checked the shipped configs directory on the new image: plenty of entries for AMD_Instinct_MI300X/MI325X, zero for AMD_Radeon_R9700 at the shapes this model needs.

Generated our own using vLLM’s own tuning benchmarks, straight from the repo:

  • benchmarks/kernels/benchmark_w8a8_block_fp8.py (dense FP8 GEMM) — note: this script is hardcoded for DeepSeek-V3/R1’s weight shapes, not derived from --model. Had to patch its get_weight_shapes() to return the actual shapes this model needs at TP=2 (read straight off the “Config file not found” warnings in the startup log).
  • benchmarks/kernels/benchmark_moe.py (MoE) — this one is model-aware (--model flag works as expected), no patching needed for shapes. Did hit an unrelated Ray/env-var bug here (Ray in this vLLM build wants HIP_VISIBLE_DEVICES, but the script’s own compatibility shim translates it into ROCR_VISIBLE_DEVICES and deletes it, then crashes because Ray now rejects that — had to patch the shim out).

Applying these tuned configs:

Old image New image
With tuning +7% (conc=1), +9% (conc=4) +2-3 percentage points only

So on the new image, tuning barely moves the needle — confirms it’s not the dominant cause of the regression. But applied to the currently-running production image, it’s a genuine free win with no downside found (tool-calling still verified working, no correctness issues). That’s now deployed via a per-file read-only bind mount over the vLLM install path pointing at the generated JSONs — cheap, reversible, no image rebuild needed.

One gotcha if you try this yourself: the old and new images use different Python/site-packages paths internally (different base images entirely, not just a vLLM bump) — config file destination paths aren’t interchangeable between them.

MoE tuning itself is slow, by the way — full sweep (18 batch sizes) took ~8.5 hours on a single R9700, serial. Needs an otherwise-idle GPU too; running it against a GPU your production server is also using will OOM partway through.

Where this leaves things

Verdict for anyone on similar hardware asking “should I upgrade”: not yet, if you’re running FP8-quantized MoE models — you’ll eat a ~50% regression for whatever the new image’s Day-0 features buy you. If your models aren’t FP8-quantized, the upgrade looks like a wash either way based on our one data point.

Open questions we haven’t chased yet: whether an INT4/GPTQ quantization of the same model sidesteps the FP8-specific kernel path (structurally different kernel dispatch — worth testing before assuming), and rocprof-level profiling to pin down the exact kernel if we do chase this further. Might file this upstream given how clean the isolation ended up being — happy to share the actual scripts/compose files if anyone wants to reproduce or extend this on their own R9700/RDNA4 box.

2 Likes

Are you being affected by the new ROCM bug affecting multiple GPUs where the prefill speed goes to nearly 0?

I’ve got 2 r9700s on order and want to try out 4 as well. Aiter is something I think this custom build addresses. Give this a try

vLLM inference server for the AMD Radeon AI PRO R9700 (gfx1201 / RDNA4). Bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack with the RDNA4 patches and custom kernels needed to run vLLM on this card, so you don’t have to build the stack yourself.

1 Like

From the other people I have spoken to, they have had issues scaling past two cards and having actually better performance, especially when using tensor parallelism.