AMD Radeon R9700 Linux TP2 and TP4 -- the Definitive Guide

Video Soon! Posting this early since many of you have messaged with this set of problems on 2 or 4 R9700s!

How to Run vLLM with Tensor Parallelism Across 4× AMD Radeon R9700 GPUs

todo: There is some really cool multi-token prediction stuff coming (MTP) .. I am still working on that. Lots of smart people here on the forum! If you’ve had success, or failure, with that share your stories and links below.

Last Updated Date: 2026-05-21
Main Test System: AMD Ryzen Threadripper 7980X, 64 GB RAM, 4× AMD Radeon AI PRO R9700 (RDNA4/gfx1201), 31.9 usable GiB VRAM each
ROCm: 6.4.3 kernel driver, ROCk module 6.12.12 (note this is a bit fungible – see context in guide)


Table of Contents

  1. The Problem
  2. Root Cause
  3. Two Working Approaches
  4. Quick Start
  5. Step-by-Step Guide
  6. Benchmark Results
  7. Troubleshooting
  8. Docker Command Reference
  9. References

1. The Problem

When starting vLLM with --tensor-parallel-size 2 or --tensor-parallel-size 4 on a system with 2+ AMD Radeon R9700 (gfx1201/RDNA4) GPUs, the server may appear to start (routes register, /v1/models responds), but:

  • Both GPUs immediately spin at 100% utilization
  • Inference requests hang indefinitely (no timeout, no response)
  • Ctrl+C may not kill the process — requires docker kill or kill -9
  • TP=1 works flawlessly on the same hardware

There are three distinct issues that affect RDNA4 multi-GPU:

Note: I also recommend disabling sleep/wake as this may be related to reports of RDNA4 fans that just… stop working or get locked to 30%. If you are experiencing these types of issues, I want to know and I want to know more about your system!

Issue A: RCCL Tuning Index Bug (affects vLLM ≥0.19)

This might be what led you here if you did a search. It can be quite frustrating to get more than 1 R9700 working, and it’s a series of innocent-ish bugs that, when strung together, cause RDNA for machine learning to fail in frustrating and catastrophic ways.

RCCL librccl.so.1.0.70201 (ROCm 7.2.x) is missing the rcclGetTuningIndexForArch() mapping for gfx1201 → tuning index 7. It falls back to tuning index 0 (designed for MI50/gfx906), selecting an incompatible AllReduce protocol that deadlocks.

Fix tracked in: ROCm/rccl PR #2166
Workaround (Approach A): NCCL_PROTO=Simple env var bypasses the broken LL protocol.
Workaround (Approach B): Use a different ROCm stack (vllm-dev image) that doesn’t have this bug.

Issue B: Shared-Memory Broadcast Hang (affects vLLM ≥0.21.0)

vLLM v0.21.0 (:latest tag) has an additional hang in shm_broadcast.py for larger models (Qwen3.6, gpt-oss-120B) on AMD.

Workaround: Use vLLM v0.20.2 (Approach A) or the vllm-dev image (Approach B).

Issue C: PCIe Topology Assumptions

Per the vLLM GitHub issue #40980 discussion, vLLM’s RCCL integration historically assumes GPUs are behind a PCIe bridge/switch (common in datacenter setups). Consumer RDNA4 cards connected directly to CPU PCIe lanes don’t have this topology, and the default tuning selects algorithms incompatible with direct CPU-attached GPUs. This causes problems on NVIDIA too, but the consequences are more catastrophic on AMD because RCCL’s LL protocol deadlocks rather than gracefully degrading.

There might actually be more bugs than these at work here – I will update as I learn more. This has been a tough one to get my head around.

What Doesn’t Work

These common workarounds from NVIDIA setups do NOT help on RDNA4:

Attempt Result
--enforce-eager :cross_mark: No effect
NCCL_P2P_DISABLE=1 :cross_mark: No effect (this works for rccl-tests but NOT vLLM)
--disable-custom-all-reduce :cross_mark: No effect
VLLM_WORKER_MULTIPROC_METHOD=spawn :cross_mark: No effect
Using vllm/vllm-openai-rocm:latest (v0.21.0) :cross_mark: Also deadlocks (ROCm 7.2.2 userspace + shm_broadcast hang)
Using rocm/vllm:latest (v0.10.1) :cross_mark: Also deadlocks (RCCL 2.22.3 too old)

2. Root Cause

The RCCL Tuning Index Bug

The definitive root cause was identified by srinivamd in vLLM Issue #40980:

RCCL librccl.so.1.0.70201 (shipped with ROCm 7.2.x) is missing the rcclGetTuningIndexForArch() function that maps gfx1201 (RDNA4/Navi4) to tuning index 7. Without this mapping, RCCL silently applies tuning index 0 — designed for MI50/gfx906 PCIe topology — to gfx1201’s PCIe-only dual-GPU configuration. The resulting algorithm/protocol selection is incompatible with RDNA4, causing the AllReduce operation to deadlock indefinitely.

Community Bisect

RCCL Version ROCm Version TP=2 Works?
librccl.so.1.0.70101 7.1.1 :white_check_mark: Yes
librccl.so.1.0.70201 7.2.x :cross_mark: Deadlocks

The Permanent Fix

ROCm/rccl PR #2166 adds {"gfx1201", 7} to src/graph/tuning.cc — not yet shipped in any release.

Why Two Approaches?

There are two independent ways to work around the bug:

  • Approach A (NCCL_PROTO=Simple) forces RCCL to use the Simple protocol instead of the broken LL protocol. This works with the standard vLLM v0.20.2 ROCm image.
  • Approach B (AITER + vllm-dev) uses a different ROCm stack (v7.0 pre-release) that doesn’t have the tuning index bug at all. It also enables AITER-accelerated MoE kernels, which is required for GPTQ-quantized models.

I think, but I am not sure, the second approach will be more forward-compatible with things like MTP?


3. Two Working Approaches

Approach A: NCCL_PROTO=Simple (vLLM v0.20.2)

Best for: Most models, simple setup, works with any vLLM-compatible model format
Docker image: vllm/vllm-openai-rocm:v0.20.2
Key env var: -e NCCL_PROTO=Simple
ROCm version in image: 7.2.2 (works despite host being 6.4.3)
Confirmed working: All Llama, DeepSeek, Qwen models

Approach B: AITER + vllm-dev Image (vLLM v0.17.0)

Best for: GPTQ-quantized models, no NCCL_PROTO needed, AITER-accelerated MoE
Docker image: rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
Key env vars: VLLM_ROCM_USE_AITER=1, HSA_OVERRIDE_GFX_VERSION=12.0.1
ROCm version in image: 7.0.51831 (pre-release)
Confirmed working: Qwen3.5-122B-A10B-GPTQ-Int4 (39.9 tok/s at TP=4)

Comparison

Aspect Approach A Approach B
Docker image vllm/vllm-openai-rocm:v0.20.2 rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
vLLM version 0.20.2 0.17.0-dev
ROCm version 7.2.2 7.0.51831 (pre-release)
NCCL_PROTO needed? Yes (=Simple) No
AITER support No Yes (VLLM_ROCM_USE_AITER=1)
GPTQ support No Yes
HSA override Not needed HSA_OVERRIDE_GFX_VERSION=12.0.1
Model compatibility Universal Limited to tested architectures
Stability Production-grade Pre-release

4. Quick Start

Note: make sure


GRUB_CMDLINE_LINUX_DEFAULT="quiet splash amd_iommu=on iommu=pt pcie_aspm=off"

Has already been done on your system.

Approach A: NCCL_PROTO=Simple (vLLM v0.20.2)

docker run --rm -d \
  --device=/dev/kfd --device=/dev/dri --group-add=video \
  -v /path/to/models:/models \
  -p 8000:8000 \
  --ipc=host \
  -e NCCL_PROTO=Simple \
  vllm/vllm-openai-rocm:v0.20.2 \
  /models/Llama-3.1-8B-Instruct \
  --tensor-parallel-size 4 \
  --dtype bfloat16 \
  --api-key token-abc123 \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.95 \
  --enforce-eager

The critical environment variable is -e NCCL_PROTO=Simple. Without this, TP≥2 will deadlock.

Approach B: AITER + vllm-dev (for GPTQ models)

docker run --rm -d \
  --name vllm-qwen-gptq \
  --ipc=host \
  --shm-size=128g \
  --device /dev/kfd:/dev/kfd \
  --device /dev/dri:/dev/dri \
  -e VLLM_ROCM_USE_AITER=1 \
  -e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
  -e VLLM_ROCM_USE_AITER_MOE=1 \
  -e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
  -e HSA_ENABLE_SDMA=0 \
  -v /path/to/models:/models \
  -p 8000:8000 \
  rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
  vllm serve /models/MODEL_NAME \
    --served-model-name my-model \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 56000 \
    --tensor-parallel-size 4 \
    --max-num-seqs 1 \
    --gpu-memory-utilization 0.95 \
    --dtype float16

5. Step-by-Step Guide

Step 1: Verify GPU Access

# Check GPUs are visible
rocm-smi
# Expected: 4x AMD Radeon AI PRO R9700, ~32 GiB VRAM each

# Verify Docker GPU access
docker run --rm --device=/dev/kfd --device=/dev/dri --group-add=video \
  rocm/dev-ubuntu-24.04:latest rocm-smi

Step 2: Pull the Docker Image

For Approach A:

docker pull vllm/vllm-openai-rocm:v0.20.2

For Approach B:

docker pull rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303

Step 3: Start the Server

Approach A — Standard models (Llama, DeepSeek, Qwen):

docker run --rm -d \
  --device=/dev/kfd --device=/dev/dri --group-add=video \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  --ipc=host \
  -e NCCL_PROTO=Simple \
  vllm/vllm-openai-rocm:v0.20.2 \
  /models/MODEL_NAME \
  --tensor-parallel-size 4 \
  --dtype bfloat16 \
  --api-key token-l1rocks \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.95 \
  --enforce-eager

Approach B — GPTQ models (Qwen3.5-122B, etc.):

docker run --rm -d \
  --name vllm-qwen-gptq \
  --ipc=host \
  --shm-size=128g \
  --device /dev/kfd:/dev/kfd \
  --device /dev/dri:/dev/dri \
  -e VLLM_ROCM_USE_AITER=1 \
  -e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
  -e VLLM_ROCM_USE_AITER_MOE=1 \
  -e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
  -e HSA_ENABLE_SDMA=0 \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
  vllm serve /models/MODEL_NAME \
    --served-model-name my-model \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 56000 \
    --tensor-parallel-size 4 \
    --max-num-seqs 1 \
    --gpu-memory-utilization 0.95 \
    --dtype float16

Step 4: Wait for Server Readiness

# Poll until the server responds
for i in $(seq 1 48); do
  sleep 5
  if curl -s http://localhost:8000/v1/models -H "Authorization: Bearer token-abc123" > /dev/null 2>&1; then
    echo "Ready after $((i*5))s"
    break
  fi
  if [ $i -eq 48 ]; then echo "TIMEOUT"; fi
done

Small models (1B-8B): 30–90 seconds
Large models (120B MoE): 4–15 minutes (kernel compilation for Approach A, longer for Approach B)

Step 5: Test Inference

# Completions endpoint (works for all models)
curl -s -X POST http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer token-abc123" \
  -d '{"model":"/models/Llama-3.1-8B-Instruct","prompt":"Hello world","max_tokens":10}'

# Chat endpoint (requires chat template)
curl -s -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer token-abc123" \
  -d '{"model":"/models/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"Hello"}],"max_tokens":10}'

Step 6: Run Benchmarks

python3 /nfs/spark-workspace/bench_intel.py \
  --api-url http://localhost:8000 \
  --api-key token-l1rocks \
  --model /models/MODEL_NAME \
  --concurrency 1 \
  --num-prompts 32 \
  --max-prompt-length 1024 \
  --max-new-tokens 256 \
  --tensor-parallel-size 4 \
  --warmup basic \
  --results-dir /nfs/spark-workspace/results

Step 7: Stop the Server

# Approach A:
docker kill $(docker ps -q --filter "ancestor=vllm/vllm-openai-rocm:v0.20.2")

# Approach B:
docker kill vllm-qwen-gptq

6. Benchmark Results

All results below are from a Threadripper 7980X with 4× AMD Radeon AI PRO R9700 (32 GB each), ROCm 6.4.3.

Benchmark parameters: 32 prompts, concurrency=1, max-prompt-length=1024, max-new-tokens=256, --enforce-eager.

Approach A: vLLM v0.20.2 + NCCL_PROTO=Simple

Model Size Quant TP tok/s TTFT avg Decode TPS Notes
Llama-3.1-8B-Instruct 8B BF16 4 61.8 0.067s 64.3 Best overall performer
DeepSeek-R1-Distill-Qwen-14B 14B BF16 4 41.5 0.092s 42.0 +42% from TP=2
Qwen3.6-35B-A3B-FP8 35B MoE FP8 4 15.6 0.106s 15.7 256 experts, 3B active
gemma-4-26B-A4B-it 26B MoE BF16 4 24.2 0.210s 26.0 128 experts, 8 active
gpt-oss-120B 120B MoE BF16 4 0.2 19.7s 5.9 No quantization; unusably slow
Llama-3.1-8B-Instruct 8B BF16 2 65.1 0.071s 67.8 Sweet spot for 8B
DeepSeek-R1-Distill-Qwen-14B 14B BF16 2 29.2 0.116s 29.4
Qwen3.6-35B-A3B-FP8 35B MoE FP8 2 17.2 0.147s 17.4
Qwen3.6-27B-bf16 27B BF16 2 17.0 0.214s 17.3 Dense alternative
Llama-3.1-8B-Instruct 8B BF16 1 38.1 0.114s 39.3 Baseline single GPU

Approach B: vLLM v0.17.0-dev + AITER

Model Size Quant TP tok/s TTFT avg Decode TPS Notes
Qwen3.5-122B-A10B-GPTQ-Int4 122B MoE GPTQ Int4 4 39.9 1.81s 53.8 256 experts, 10B active
Reddit post reference 122B MoE GPTQ Int4 4 41.0 34.9s 41.0 41k context, from u/grunt_monkey_

Scaling Observations

  • Llama-3.1-8B: TP=1→2 gives 1.7× (38→65 tok/s). TP=2→4 shows no gain (65→62) — too small for 4 GPUs.
  • DeepSeek-14B: TP=2→4 gives 1.4× (29→42 tok/s) — benefits from additional GPUs.
  • Qwen3.5-122B at TP=4: 39.9 tok/s is remarkable for a 122B MoE model. The GPTQ Int4 quantization makes it feasible on 32 GB GPUs.
  • MoE models (Qwen3.5, Qwen3.6, gemma-4): At low concurrency, TP=4 can be slightly slower than TP=2 due to communication overhead. Higher batch sizes would flip this.
  • gpt-oss-120B: Requires proper quantization (MXFP4/QUARK) to fit and run well on 4× R9700. Without quantization, it loads but produces <1 tok/s.

7. Troubleshooting

Server starts but inference hangs (GPUs at 100%)

Cause: RCCL tuning index mismatch — the LL protocol deadlocks on RDNA4.

Fix (Approach A): Ensure -e NCCL_PROTO=Simple is set. Verify with:

docker logs <container-id> | grep -i "NCCL_PROTO\|RCCL"

Fix (Approach B): The vllm-dev image doesn’t have this bug. Verify AITER is active:

docker logs vllm-qwen-gptq | grep -i "AITER\|aiter"

Container exits with “CUDA/ROCm GPU is required”

Cause: Docker doesn’t have GPU access.

Fix:

ls -la /dev/kfd /dev/dri/
getent group render
# Use: --device=/dev/kfd --device=/dev/dri --group-add=video

“No available shared memory broadcast block found in 60 seconds”

Cause: vLLM v0.21.0 shm_broadcast hang (Issue B).

Fix: Use vLLM v0.20.2 (Approach A) or the vllm-dev image (Approach B).

Model loads but produces zero output tokens

Cause: No chat template in tokenizer.

Fix: Use /v1/completions endpoint instead of /v1/chat/completions.

OOM during model loading

Fix: Reduce --max-model-len, lower --gpu-memory-utilization, or increase TP:

--max-model-len 4096
--gpu-memory-utilization 0.85

Host becomes unresponsive after failed TP attempts

Cause: RCCL deadlocks leave GPU processes in unrecoverable state.

Fix: Reboot the host, then remount NFS:

sudo reboot
# After reboot:
sudo mount -t nfs 10.200.0.129:/volume1/aidump /nfs

Orphan vLLM processes at 100% GPU with no container running

Fix:

ps aux | grep -E "vllm|LLM|EngineCore" | grep -v grep | awk '{print $2}' | xargs -r kill -9

GPUs running hot (90°C+)

Fix: Use fan control software (e.g., Afterburner) or set aggressive fan curves. Consider disabling ECC and setting perf-level=HIGH via amd-smi.

Approach B: AITER not found / import errors

Cause: Missing or incorrect env vars for the vllm-dev image.

Fix: Ensure all of these are set:

-e VLLM_ROCM_USE_AITER=1
-e HSA_OVERRIDE_GFX_VERSION=12.0.1
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE

Approach B: Container exits with “hipErrorNoBinaryForGpu”

Cause: HSA_OVERRIDE_GFX_VERSION not set or wrong value.

Fix: Set -e HSA_OVERRIDE_GFX_VERSION=12.0.1 (for gfx1201 / RDNA4).


8. Docker Command Reference

Approach A: vLLM v0.20.2 (universal)

Basic server (any model):

docker run --rm -d \
  --device=/dev/kfd --device=/dev/dri --group-add=video \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  --ipc=host \
  -e NCCL_PROTO=Simple \
  vllm/vllm-openai-rocm:v0.20.2 \
  /models/MODEL_NAME \
  --tensor-parallel-size N \
  --dtype bfloat16 \
  --api-key token-l1rocks \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.95 \
  --enforce-eager

Small model (TP=1, no NCCL issues):

docker run --rm -d \
  --device=/dev/kfd --device=/dev/dri --group-add=video \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  --ipc=host \
  vllm/vllm-openai-rocm:v0.20.2 \
  /models/Llama-3.2-1B \
  --tensor-parallel-size 1 \
  --dtype bfloat16 \
  --api-key token-l1rocks

Large model (120B MoE, needs TP=4, conservative settings):

docker run --rm -d \
  --device=/dev/kfd --device=/dev/dri --group-add=video \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  --ipc=host \
  -e NCCL_PROTO=Simple \
  vllm/vllm-openai-rocm:v0.20.2 \
  /models/gpt-oss-120B \
  --tensor-parallel-size 4 \
  --dtype bfloat16 \
  --api-key token-l1rocks \
  --max-model-len 4096 \
  --gpu-memory-utilization 0.85 \
  --enforce-eager

docker-compose.yml for Approach A:

version: "3.8"
services:
  vllm:
    image: vllm/vllm-openai-rocm:v0.20.2
    devices:
      - /dev/kfd:/dev/kfd
      - /dev/dri:/dev/dri
    volumes:
      - /nfs/spark-workspace/models:/models
    ports:
      - "8000:8000"
    environment:
      - NCCL_PROTO=Simple
    ipc: host
    command: >
      /models/Llama-3.1-8B-Instruct
      --tensor-parallel-size 4
      --dtype bfloat16
      --api-key token-abc123
      --max-model-len 16384
      --gpu-memory-utilization 0.95
      --enforce-eager

Approach B: vllm-dev + AITER (for GPTQ models)

Qwen3.5-122B-A10B-GPTQ-Int4 at TP=4:

docker run --rm -d \
  --name vllm-qwen-gptq \
  --ipc=host \
  --shm-size=128g \
  --device /dev/kfd:/dev/kfd \
  --device /dev/dri:/dev/dri \
  -e VLLM_ROCM_USE_AITER=1 \
  -e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
  -e VLLM_ROCM_USE_AITER_MOE=1 \
  -e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
  -e HSA_ENABLE_SDMA=0 \
  -v /nfs/spark-workspace/models:/models \
  -p 8000:8000 \
  rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
  vllm serve /models/Qwen3.5-122B-A10B-GPTQ-Int4 \
    --served-model-name Qwen3.5-122B \
    --host 0.0.0.0 \
    --port 8000 \
    --max-model-len 56000 \
    --tensor-parallel-size 4 \
    --max-num-seqs 1 \
    --gpu-memory-utilization 0.95 \
    --dtype float16

docker-compose.yml for Approach B:

version: "3.8"
services:
  vllm:
    image: rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
    devices:
      - /dev/kfd:/dev/kfd
      - /dev/dri:/dev/dri
    volumes:
      - /nfs/spark-workspace/models:/models
    ports:
      - "8000:8000"
    environment:
      - VLLM_ROCM_USE_AITER=1
      - HSA_OVERRIDE_GFX_VERSION=12.0.1
      - VLLM_ROCM_USE_AITER_MOE=1
      - FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
      - HSA_ENABLE_SDMA=0
    ipc: host
    shm_size: 128g
    command: >
      vllm serve /models/Qwen3.5-122B-A10B-GPTQ-Int4
      --served-model-name Qwen3.5-122B
      --host 0.0.0.0
      --port 8000
      --max-model-len 56000
      --tensor-parallel-size 4
      --max-num-seqs 1
      --gpu-memory-utilization 0.95
      --dtype float16

Shared: Benchmark & Visualization

TODO


9. References

GitHub Issues

  • vLLM #40980: TP=2 deadlock on dual AMD R9700 (gfx1201/RDNA4). Root cause analysis by srinivamd, NCCL_PROTO=Simple workaround by vllmellm.
  • ROCm/rccl PR #2166: “[cherry-pick] 7.2.1 Cherry pick changes [Navi4 default tuning enablement]” — the permanent fix (adds gfx1201→tuning_index_7). Not yet shipped.
  • rccl-tests #162: Radeon AI PRO R9700 4-card P2P failure — separate issue about hipIpcGetMemHandle on systems without PCIe switches.
  • vLLM #39010: Hang During CUDA Graph Capture on ROCM.
  • vLLM #40081: vLLM fails to start on RDNA 4 (gfx1201) inside containers.

Community Resources

Key People Working On This

  • srinivamd — Root cause analysis of the RCCL tuning index bug
  • vllmellm — Identified NCCL_PROTO=Simple workaround
  • 0xb3-err — Dockerfile for RCCL downgrade workaround
  • kyuz0 — Original reporter of vLLM #40980
  • grunt_monkey_ — Reddit post demonstrating Qwen3.5-122B at 41 tok/s on 4× R9700

Workarounds Summary

Workaround Applies To How To
NCCL_PROTO=Simple vLLM TP≥2 deadlock (Approach A) docker run -e NCCL_PROTO=Simple ...
AITER + vllm-dev vLLM TP≥2 deadlock (Approach B) Use rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
RCCL downgrade to 7.1.1 vLLM TP≥2 deadlock Replace librccl.so in container
Host reboot GPU state corruption sudo reboot
NCCL_P2P_DISABLE=1 rccl-tests only export NCCL_P2P_DISABLE=1

Generated from real benchmark runs on amdrocm (Threadripper 7980X, 4× R9700, ROCm 6.4.3)
Reddit reference: Qwen3.5-122B-A10B GPTQ Int4 on 4× Radeon AI PRO R9700

Huge Thanks

Thanks to the Level1Techs community for supporting these projects so I can figure stuff out.

What’s next??

MTP. We have to get MTP working. 4 GPUs? 128gb? MTP for 30-60% speedup and agentic and/or coding tasks? We need this!

12 Likes

It’s a good time to be an R9700 owner. Amazing work!

I know local agents are always getting better and more efficient.

But I think I’ll keep waiting until AMD comes out with a bigger die card that has 64GBs of VRAM, something right in the middle of an R9700 and the new MI 350P(hope you get one for testing from someone)

Unless I can find real world non-coder benchmarks and use cases, like if the new QWEN 3.7 or whatever can successfully decompile and reverse engineer retro games like GPT 5.5 can. Currently at 1000+ plan files, running 5 plan tooling upgrades/mapping efforts at a time. Wish I saved all the plan files from my gameboy game decomp, could share them and throw local agents at them just to rebuild the tooling and execute them.

GPT 5.5 / codex has gotten incredible at long running tasks, I haven’t switched agents in 100*5 goal runs and my token usage is barely impacted from all the automatic compacting.

Great writeup. Thanks Wendell!

I’m in the process of building a local LLM Server for inference, coding and agentic workloads.

  • Supermicro H12SSL + Epyc 7663 56core (450€ on eBay whooo)
  • 256gb RAM 3200MT/s
  • Silverstone RM4A Case, with some thik Artic 38mm server fans
  • Enermax 1650w psu
  • 2x Asrock R9700

Sure the Epyc is overkill. But for 450€ why not? Planning on using a huge ramdisk to quickly load and unload the models. Let’s see if this will work out.

Maybe I will add another pair R9700 later. Alphacool announced a waterblock for this gpu. Could be worth it to get thermals and noise under control.

Thanks for amazing post!

i am using 8xR9700 with MXFP4 Qwen3.5-397B got only 32-33 t/s how do you thing, is it ok decode speed or i have any chance to boost it? Context size 91k tp=8. mtp not enabled.

also have qwen3.5-122B MXFP4 with 110 t/s and mtp = 3 configuration.

vllm version is 0.18.1 dev

Well, what is a single R9700 getting you?

Also is it not worth switching over to Qwen 3.7? Or does that have some other requirements behind it?

He’s running some big models. 128gb is kinda the sweet spot and the r9700 can have some great perf in these configs. One r9700 with most layers on the GPU will have qwen3.6 not quite fit in 32gb. And you really don’t want less than q8 mostly for coding tasks

Boss, has anyone sent you an mi350p to at least borrow for testing?

I see, that is quite the set up to be sure. AMD better come out with at 64GB card for RDNA5, that seems like it’ll be the best sweet spot for local agents, where strix halo is apparently too slow on it’s fill rate, or you need multiple R9700s.

But in a broader sense how far are these set ups from say GPT 5.3 Codex levels of performance? GPT 5.5 has been beyond incredible for me, and produces only a few warnings or errors as I’m building a game engine in rust. It hasn’t made a single error in my Cura modifications either.

Where 5.3 Codex threw out many errors on my projects as a total non-coder.

I am attempting to make my Epyc 7663 + 4xR9700 build do something useful and this write-up pretty much summarises the maddeningly complex challenge. I am on attempt 4 (5?) and even building local container wheels keeps falling of a cliff based on evolved dependencies. Even with nudges by a SOTA model, there are pitfalls. Compared to my 4x3090 or 8x3090 of the past, this things is a pita. Silly me, I wanted a lower power setup with more VRAM then 4x3090.
Thank you for Option B, that has at least allowed me to leverage my setup with a model I happen to like and use on other setups. Here’s hoping the planets and stars line up to help deliver an even better experience soon.
Given my focus on tool use, Seraphim’s tool-eval-bench has been a constant validator.

@djdeniro I would love to see your vllm config to attempt a tp=4 setup. Thanks!

I’m still working on this so lmk what you ran into. I only had my 4 setup temporarily when I first did my videos but lots changed since then and only just got 4 again a couple weeks ago so trying to help makes sure I’ve got my bases covered

2 Likes

Thanks Wendell. If you need a validation of something, let me know. So far Option B is the only solid one. I have put some effort in checking [Matt’s efforts on SGlang] to no avail.

1 Like

This has been very helpful, would it be possible to have the bench_intel.py script to do a proper comparison and to submit back benchmarks.
I’ve been running some of the same models (on 4 x R9700) and seem to have a little higher total tps (NOT DOING ANYTHING SPECIAL) probably just taking advantage to the rapid update to rocm and vllm.

1 Like

Would love to see your vllm recipe and the specific model, to double check on my side. Thanks

? It’s the above, up there. in the op post. you can docker compose to exactly what I did. The python script isn’t anything super special though I’m constantly modifying it. if you run llama-bench with a largeish context (128k+) it should be same result .. the only reason I dont use llama-bench universally is that it doesn’t work in every scenario @beekeeper

Option A just hanged for me, I can re-test. I was actually trying to see if @beekeeper or others have some tidbit in the docker compose or vllm setup that “make it just work”. For me, only Option B works so far and nothing obvious jumps why. I assume I am making some dumb mistakes, but I do work with containers all day, so perhaps not as likely if B works :slight_smile:

I am running podman and on CachyOS base, so perhaps I’m seeing a weird interction, but I doubt that.

No, thats the behavior of option A if it isn’t picking up the NCCL proto as simple environment setting. It’ll just hang/crash. You may have to reboot after that.

can you do the troubleshooting step:

Fix (Approach A): Ensure -e NCCL_PROTO=Simple is set. Verify with:

docker logs <container-id> | grep -i "NCCL_PROTO\|RCCL"

and confirm?
anything useful in docker logs about the hang otherwise?

I had the NCCL_PROTO=Simple

But I realized that since I am re-using scripts of my own, perhaps something else gets in the way and I think I found it. I had caching setup for Triton, flashinfer and vllm in general. I deleted them and magically I got GPT 120b to deliver ~6 tps. I don’t know why the cache was the pain, but it was odd that sometimes a compile, gcc/clang or passt.avx2 would just sit there for 45 min and nothing. So I think it is on me, as tired debugging doesn’t yield clear results. Also, I know I had used sudo in anger a few times, and that could be that something had the wrong permissions in there. I don’t know.

Naturally, now I poke and prod to see what else can run, but since so many quantizations don’t work, I think Option B is the most logical for now.

Tried the recent nightly and that didn’t even work.

I just posted another how-to but its aimed at strix halo. However some of the learnings there may be applicable here. The performance was pretty good on strix halo for A35B qwen imho. And I actually got ROCm FP4 to work right. (FP4 is not expected to help here on RDNA though).

1 Like

Thanks Wendell, I guess I’ll dust of my Minisforum MS-S1 Max and check.

So far, single, dual and quad DGX is far more effective. Though 3090s are the most versatile, still. Who’d have thought that even a few years ago.