Video Soon! Posting this early since many of you have messaged with this set of problems on 2 or 4 R9700s!
How to Run vLLM with Tensor Parallelism Across 4× AMD Radeon R9700 GPUs
todo: There is some really cool multi-token prediction stuff coming (MTP) .. I am still working on that. Lots of smart people here on the forum! If you’ve had success, or failure, with that share your stories and links below.
Last Updated Date: 2026-05-21
Main Test System: AMD Ryzen Threadripper 7980X, 64 GB RAM, 4× AMD Radeon AI PRO R9700 (RDNA4/gfx1201), 31.9 usable GiB VRAM each
ROCm: 6.4.3 kernel driver, ROCk module 6.12.12 (note this is a bit fungible – see context in guide)
Table of Contents
- The Problem
- Root Cause
- Two Working Approaches
- Quick Start
- Step-by-Step Guide
- Benchmark Results
- Troubleshooting
- Docker Command Reference
- References
1. The Problem
When starting vLLM with --tensor-parallel-size 2 or --tensor-parallel-size 4 on a system with 2+ AMD Radeon R9700 (gfx1201/RDNA4) GPUs, the server may appear to start (routes register, /v1/models responds), but:
- Both GPUs immediately spin at 100% utilization
- Inference requests hang indefinitely (no timeout, no response)
- Ctrl+C may not kill the process — requires
docker killorkill -9 - TP=1 works flawlessly on the same hardware
There are three distinct issues that affect RDNA4 multi-GPU:
Note: I also recommend disabling sleep/wake as this may be related to reports of RDNA4 fans that just… stop working or get locked to 30%. If you are experiencing these types of issues, I want to know and I want to know more about your system!
Issue A: RCCL Tuning Index Bug (affects vLLM ≥0.19)
This might be what led you here if you did a search. It can be quite frustrating to get more than 1 R9700 working, and it’s a series of innocent-ish bugs that, when strung together, cause RDNA for machine learning to fail in frustrating and catastrophic ways.
RCCL librccl.so.1.0.70201 (ROCm 7.2.x) is missing the rcclGetTuningIndexForArch() mapping for gfx1201 → tuning index 7. It falls back to tuning index 0 (designed for MI50/gfx906), selecting an incompatible AllReduce protocol that deadlocks.
Fix tracked in: ROCm/rccl PR #2166
Workaround (Approach A): NCCL_PROTO=Simple env var bypasses the broken LL protocol.
Workaround (Approach B): Use a different ROCm stack (vllm-dev image) that doesn’t have this bug.
Issue B: Shared-Memory Broadcast Hang (affects vLLM ≥0.21.0)
vLLM v0.21.0 (:latest tag) has an additional hang in shm_broadcast.py for larger models (Qwen3.6, gpt-oss-120B) on AMD.
Workaround: Use vLLM v0.20.2 (Approach A) or the vllm-dev image (Approach B).
Issue C: PCIe Topology Assumptions
Per the vLLM GitHub issue #40980 discussion, vLLM’s RCCL integration historically assumes GPUs are behind a PCIe bridge/switch (common in datacenter setups). Consumer RDNA4 cards connected directly to CPU PCIe lanes don’t have this topology, and the default tuning selects algorithms incompatible with direct CPU-attached GPUs. This causes problems on NVIDIA too, but the consequences are more catastrophic on AMD because RCCL’s LL protocol deadlocks rather than gracefully degrading.
There might actually be more bugs than these at work here – I will update as I learn more. This has been a tough one to get my head around.
What Doesn’t Work
These common workarounds from NVIDIA setups do NOT help on RDNA4:
| Attempt | Result |
|---|---|
--enforce-eager |
|
NCCL_P2P_DISABLE=1 |
|
--disable-custom-all-reduce |
|
VLLM_WORKER_MULTIPROC_METHOD=spawn |
|
Using vllm/vllm-openai-rocm:latest (v0.21.0) |
|
Using rocm/vllm:latest (v0.10.1) |
2. Root Cause
The RCCL Tuning Index Bug
The definitive root cause was identified by srinivamd in vLLM Issue #40980:
RCCL
librccl.so.1.0.70201(shipped with ROCm 7.2.x) is missing thercclGetTuningIndexForArch()function that maps gfx1201 (RDNA4/Navi4) to tuning index 7. Without this mapping, RCCL silently applies tuning index 0 — designed for MI50/gfx906 PCIe topology — to gfx1201’s PCIe-only dual-GPU configuration. The resulting algorithm/protocol selection is incompatible with RDNA4, causing the AllReduce operation to deadlock indefinitely.
Community Bisect
| RCCL Version | ROCm Version | TP=2 Works? |
|---|---|---|
librccl.so.1.0.70101 |
7.1.1 | |
librccl.so.1.0.70201 |
7.2.x |
The Permanent Fix
ROCm/rccl PR #2166 adds {"gfx1201", 7} to src/graph/tuning.cc — not yet shipped in any release.
Why Two Approaches?
There are two independent ways to work around the bug:
- Approach A (
NCCL_PROTO=Simple) forces RCCL to use the Simple protocol instead of the broken LL protocol. This works with the standard vLLM v0.20.2 ROCm image. - Approach B (AITER + vllm-dev) uses a different ROCm stack (v7.0 pre-release) that doesn’t have the tuning index bug at all. It also enables AITER-accelerated MoE kernels, which is required for GPTQ-quantized models.
I think, but I am not sure, the second approach will be more forward-compatible with things like MTP?
3. Two Working Approaches
Approach A: NCCL_PROTO=Simple (vLLM v0.20.2)
Best for: Most models, simple setup, works with any vLLM-compatible model format
Docker image: vllm/vllm-openai-rocm:v0.20.2
Key env var: -e NCCL_PROTO=Simple
ROCm version in image: 7.2.2 (works despite host being 6.4.3)
Confirmed working: All Llama, DeepSeek, Qwen models
Approach B: AITER + vllm-dev Image (vLLM v0.17.0)
Best for: GPTQ-quantized models, no NCCL_PROTO needed, AITER-accelerated MoE
Docker image: rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
Key env vars: VLLM_ROCM_USE_AITER=1, HSA_OVERRIDE_GFX_VERSION=12.0.1
ROCm version in image: 7.0.51831 (pre-release)
Confirmed working: Qwen3.5-122B-A10B-GPTQ-Int4 (39.9 tok/s at TP=4)
Comparison
| Aspect | Approach A | Approach B |
|---|---|---|
| Docker image | vllm/vllm-openai-rocm:v0.20.2 |
rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 |
| vLLM version | 0.20.2 | 0.17.0-dev |
| ROCm version | 7.2.2 | 7.0.51831 (pre-release) |
| NCCL_PROTO needed? | Yes (=Simple) |
No |
| AITER support | No | Yes (VLLM_ROCM_USE_AITER=1) |
| GPTQ support | No | Yes |
| HSA override | Not needed | HSA_OVERRIDE_GFX_VERSION=12.0.1 |
| Model compatibility | Universal | Limited to tested architectures |
| Stability | Production-grade | Pre-release |
4. Quick Start
Note: make sure
GRUB_CMDLINE_LINUX_DEFAULT="quiet splash amd_iommu=on iommu=pt pcie_aspm=off"
Has already been done on your system.
Approach A: NCCL_PROTO=Simple (vLLM v0.20.2)
docker run --rm -d \
--device=/dev/kfd --device=/dev/dri --group-add=video \
-v /path/to/models:/models \
-p 8000:8000 \
--ipc=host \
-e NCCL_PROTO=Simple \
vllm/vllm-openai-rocm:v0.20.2 \
/models/Llama-3.1-8B-Instruct \
--tensor-parallel-size 4 \
--dtype bfloat16 \
--api-key token-abc123 \
--max-model-len 16384 \
--gpu-memory-utilization 0.95 \
--enforce-eager
The critical environment variable is -e NCCL_PROTO=Simple. Without this, TP≥2 will deadlock.
Approach B: AITER + vllm-dev (for GPTQ models)
docker run --rm -d \
--name vllm-qwen-gptq \
--ipc=host \
--shm-size=128g \
--device /dev/kfd:/dev/kfd \
--device /dev/dri:/dev/dri \
-e VLLM_ROCM_USE_AITER=1 \
-e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
-e VLLM_ROCM_USE_AITER_MOE=1 \
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
-e HSA_ENABLE_SDMA=0 \
-v /path/to/models:/models \
-p 8000:8000 \
rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
vllm serve /models/MODEL_NAME \
--served-model-name my-model \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 56000 \
--tensor-parallel-size 4 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.95 \
--dtype float16
5. Step-by-Step Guide
Step 1: Verify GPU Access
# Check GPUs are visible
rocm-smi
# Expected: 4x AMD Radeon AI PRO R9700, ~32 GiB VRAM each
# Verify Docker GPU access
docker run --rm --device=/dev/kfd --device=/dev/dri --group-add=video \
rocm/dev-ubuntu-24.04:latest rocm-smi
Step 2: Pull the Docker Image
For Approach A:
docker pull vllm/vllm-openai-rocm:v0.20.2
For Approach B:
docker pull rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
Step 3: Start the Server
Approach A — Standard models (Llama, DeepSeek, Qwen):
docker run --rm -d \
--device=/dev/kfd --device=/dev/dri --group-add=video \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
--ipc=host \
-e NCCL_PROTO=Simple \
vllm/vllm-openai-rocm:v0.20.2 \
/models/MODEL_NAME \
--tensor-parallel-size 4 \
--dtype bfloat16 \
--api-key token-l1rocks \
--max-model-len 16384 \
--gpu-memory-utilization 0.95 \
--enforce-eager
Approach B — GPTQ models (Qwen3.5-122B, etc.):
docker run --rm -d \
--name vllm-qwen-gptq \
--ipc=host \
--shm-size=128g \
--device /dev/kfd:/dev/kfd \
--device /dev/dri:/dev/dri \
-e VLLM_ROCM_USE_AITER=1 \
-e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
-e VLLM_ROCM_USE_AITER_MOE=1 \
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
-e HSA_ENABLE_SDMA=0 \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
vllm serve /models/MODEL_NAME \
--served-model-name my-model \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 56000 \
--tensor-parallel-size 4 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.95 \
--dtype float16
Step 4: Wait for Server Readiness
# Poll until the server responds
for i in $(seq 1 48); do
sleep 5
if curl -s http://localhost:8000/v1/models -H "Authorization: Bearer token-abc123" > /dev/null 2>&1; then
echo "Ready after $((i*5))s"
break
fi
if [ $i -eq 48 ]; then echo "TIMEOUT"; fi
done
Small models (1B-8B): 30–90 seconds
Large models (120B MoE): 4–15 minutes (kernel compilation for Approach A, longer for Approach B)
Step 5: Test Inference
# Completions endpoint (works for all models)
curl -s -X POST http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer token-abc123" \
-d '{"model":"/models/Llama-3.1-8B-Instruct","prompt":"Hello world","max_tokens":10}'
# Chat endpoint (requires chat template)
curl -s -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer token-abc123" \
-d '{"model":"/models/Llama-3.1-8B-Instruct","messages":[{"role":"user","content":"Hello"}],"max_tokens":10}'
Step 6: Run Benchmarks
python3 /nfs/spark-workspace/bench_intel.py \
--api-url http://localhost:8000 \
--api-key token-l1rocks \
--model /models/MODEL_NAME \
--concurrency 1 \
--num-prompts 32 \
--max-prompt-length 1024 \
--max-new-tokens 256 \
--tensor-parallel-size 4 \
--warmup basic \
--results-dir /nfs/spark-workspace/results
Step 7: Stop the Server
# Approach A:
docker kill $(docker ps -q --filter "ancestor=vllm/vllm-openai-rocm:v0.20.2")
# Approach B:
docker kill vllm-qwen-gptq
6. Benchmark Results
All results below are from a Threadripper 7980X with 4× AMD Radeon AI PRO R9700 (32 GB each), ROCm 6.4.3.
Benchmark parameters: 32 prompts, concurrency=1, max-prompt-length=1024, max-new-tokens=256, --enforce-eager.
Approach A: vLLM v0.20.2 + NCCL_PROTO=Simple
| Model | Size | Quant | TP | tok/s | TTFT avg | Decode TPS | Notes |
|---|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | 8B | BF16 | 4 | 61.8 | 0.067s | 64.3 | Best overall performer |
| DeepSeek-R1-Distill-Qwen-14B | 14B | BF16 | 4 | 41.5 | 0.092s | 42.0 | +42% from TP=2 |
| Qwen3.6-35B-A3B-FP8 | 35B MoE | FP8 | 4 | 15.6 | 0.106s | 15.7 | 256 experts, 3B active |
| gemma-4-26B-A4B-it | 26B MoE | BF16 | 4 | 24.2 | 0.210s | 26.0 | 128 experts, 8 active |
| gpt-oss-120B | 120B MoE | BF16 | 4 | 0.2 | 19.7s | 5.9 | No quantization; unusably slow |
| Llama-3.1-8B-Instruct | 8B | BF16 | 2 | 65.1 | 0.071s | 67.8 | Sweet spot for 8B |
| DeepSeek-R1-Distill-Qwen-14B | 14B | BF16 | 2 | 29.2 | 0.116s | 29.4 | |
| Qwen3.6-35B-A3B-FP8 | 35B MoE | FP8 | 2 | 17.2 | 0.147s | 17.4 | |
| Qwen3.6-27B-bf16 | 27B | BF16 | 2 | 17.0 | 0.214s | 17.3 | Dense alternative |
| Llama-3.1-8B-Instruct | 8B | BF16 | 1 | 38.1 | 0.114s | 39.3 | Baseline single GPU |
Approach B: vLLM v0.17.0-dev + AITER
| Model | Size | Quant | TP | tok/s | TTFT avg | Decode TPS | Notes |
|---|---|---|---|---|---|---|---|
| Qwen3.5-122B-A10B-GPTQ-Int4 | 122B MoE | GPTQ Int4 | 4 | 39.9 | 1.81s | 53.8 | 256 experts, 10B active |
| Reddit post reference | 122B MoE | GPTQ Int4 | 4 | 41.0 | 34.9s | 41.0 | 41k context, from u/grunt_monkey_ |
Scaling Observations
- Llama-3.1-8B: TP=1→2 gives 1.7× (38→65 tok/s). TP=2→4 shows no gain (65→62) — too small for 4 GPUs.
- DeepSeek-14B: TP=2→4 gives 1.4× (29→42 tok/s) — benefits from additional GPUs.
- Qwen3.5-122B at TP=4: 39.9 tok/s is remarkable for a 122B MoE model. The GPTQ Int4 quantization makes it feasible on 32 GB GPUs.
- MoE models (Qwen3.5, Qwen3.6, gemma-4): At low concurrency, TP=4 can be slightly slower than TP=2 due to communication overhead. Higher batch sizes would flip this.
- gpt-oss-120B: Requires proper quantization (MXFP4/QUARK) to fit and run well on 4× R9700. Without quantization, it loads but produces <1 tok/s.
7. Troubleshooting
Server starts but inference hangs (GPUs at 100%)
Cause: RCCL tuning index mismatch — the LL protocol deadlocks on RDNA4.
Fix (Approach A): Ensure -e NCCL_PROTO=Simple is set. Verify with:
docker logs <container-id> | grep -i "NCCL_PROTO\|RCCL"
Fix (Approach B): The vllm-dev image doesn’t have this bug. Verify AITER is active:
docker logs vllm-qwen-gptq | grep -i "AITER\|aiter"
Container exits with “CUDA/ROCm GPU is required”
Cause: Docker doesn’t have GPU access.
Fix:
ls -la /dev/kfd /dev/dri/
getent group render
# Use: --device=/dev/kfd --device=/dev/dri --group-add=video
“No available shared memory broadcast block found in 60 seconds”
Cause: vLLM v0.21.0 shm_broadcast hang (Issue B).
Fix: Use vLLM v0.20.2 (Approach A) or the vllm-dev image (Approach B).
Model loads but produces zero output tokens
Cause: No chat template in tokenizer.
Fix: Use /v1/completions endpoint instead of /v1/chat/completions.
OOM during model loading
Fix: Reduce --max-model-len, lower --gpu-memory-utilization, or increase TP:
--max-model-len 4096
--gpu-memory-utilization 0.85
Host becomes unresponsive after failed TP attempts
Cause: RCCL deadlocks leave GPU processes in unrecoverable state.
Fix: Reboot the host, then remount NFS:
sudo reboot
# After reboot:
sudo mount -t nfs 10.200.0.129:/volume1/aidump /nfs
Orphan vLLM processes at 100% GPU with no container running
Fix:
ps aux | grep -E "vllm|LLM|EngineCore" | grep -v grep | awk '{print $2}' | xargs -r kill -9
GPUs running hot (90°C+)
Fix: Use fan control software (e.g., Afterburner) or set aggressive fan curves. Consider disabling ECC and setting perf-level=HIGH via amd-smi.
Approach B: AITER not found / import errors
Cause: Missing or incorrect env vars for the vllm-dev image.
Fix: Ensure all of these are set:
-e VLLM_ROCM_USE_AITER=1
-e HSA_OVERRIDE_GFX_VERSION=12.0.1
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
Approach B: Container exits with “hipErrorNoBinaryForGpu”
Cause: HSA_OVERRIDE_GFX_VERSION not set or wrong value.
Fix: Set -e HSA_OVERRIDE_GFX_VERSION=12.0.1 (for gfx1201 / RDNA4).
8. Docker Command Reference
Approach A: vLLM v0.20.2 (universal)
Basic server (any model):
docker run --rm -d \
--device=/dev/kfd --device=/dev/dri --group-add=video \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
--ipc=host \
-e NCCL_PROTO=Simple \
vllm/vllm-openai-rocm:v0.20.2 \
/models/MODEL_NAME \
--tensor-parallel-size N \
--dtype bfloat16 \
--api-key token-l1rocks \
--max-model-len 16384 \
--gpu-memory-utilization 0.95 \
--enforce-eager
Small model (TP=1, no NCCL issues):
docker run --rm -d \
--device=/dev/kfd --device=/dev/dri --group-add=video \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai-rocm:v0.20.2 \
/models/Llama-3.2-1B \
--tensor-parallel-size 1 \
--dtype bfloat16 \
--api-key token-l1rocks
Large model (120B MoE, needs TP=4, conservative settings):
docker run --rm -d \
--device=/dev/kfd --device=/dev/dri --group-add=video \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
--ipc=host \
-e NCCL_PROTO=Simple \
vllm/vllm-openai-rocm:v0.20.2 \
/models/gpt-oss-120B \
--tensor-parallel-size 4 \
--dtype bfloat16 \
--api-key token-l1rocks \
--max-model-len 4096 \
--gpu-memory-utilization 0.85 \
--enforce-eager
docker-compose.yml for Approach A:
version: "3.8"
services:
vllm:
image: vllm/vllm-openai-rocm:v0.20.2
devices:
- /dev/kfd:/dev/kfd
- /dev/dri:/dev/dri
volumes:
- /nfs/spark-workspace/models:/models
ports:
- "8000:8000"
environment:
- NCCL_PROTO=Simple
ipc: host
command: >
/models/Llama-3.1-8B-Instruct
--tensor-parallel-size 4
--dtype bfloat16
--api-key token-abc123
--max-model-len 16384
--gpu-memory-utilization 0.95
--enforce-eager
Approach B: vllm-dev + AITER (for GPTQ models)
Qwen3.5-122B-A10B-GPTQ-Int4 at TP=4:
docker run --rm -d \
--name vllm-qwen-gptq \
--ipc=host \
--shm-size=128g \
--device /dev/kfd:/dev/kfd \
--device /dev/dri:/dev/dri \
-e VLLM_ROCM_USE_AITER=1 \
-e HSA_OVERRIDE_GFX_VERSION=12.0.1 \
-e VLLM_ROCM_USE_AITER_MOE=1 \
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE \
-e HSA_ENABLE_SDMA=0 \
-v /nfs/spark-workspace/models:/models \
-p 8000:8000 \
rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 \
vllm serve /models/Qwen3.5-122B-A10B-GPTQ-Int4 \
--served-model-name Qwen3.5-122B \
--host 0.0.0.0 \
--port 8000 \
--max-model-len 56000 \
--tensor-parallel-size 4 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.95 \
--dtype float16
docker-compose.yml for Approach B:
version: "3.8"
services:
vllm:
image: rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303
devices:
- /dev/kfd:/dev/kfd
- /dev/dri:/dev/dri
volumes:
- /nfs/spark-workspace/models:/models
ports:
- "8000:8000"
environment:
- VLLM_ROCM_USE_AITER=1
- HSA_OVERRIDE_GFX_VERSION=12.0.1
- VLLM_ROCM_USE_AITER_MOE=1
- FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
- HSA_ENABLE_SDMA=0
ipc: host
shm_size: 128g
command: >
vllm serve /models/Qwen3.5-122B-A10B-GPTQ-Int4
--served-model-name Qwen3.5-122B
--host 0.0.0.0
--port 8000
--max-model-len 56000
--tensor-parallel-size 4
--max-num-seqs 1
--gpu-memory-utilization 0.95
--dtype float16
Shared: Benchmark & Visualization
TODO
9. References
GitHub Issues
- vLLM #40980: TP=2 deadlock on dual AMD R9700 (gfx1201/RDNA4). Root cause analysis by
srinivamd,NCCL_PROTO=Simpleworkaround byvllmellm. - ROCm/rccl PR #2166: “[cherry-pick] 7.2.1 Cherry pick changes [Navi4 default tuning enablement]” — the permanent fix (adds gfx1201→tuning_index_7). Not yet shipped.
- rccl-tests #162: Radeon AI PRO R9700 4-card P2P failure — separate issue about
hipIpcGetMemHandleon systems without PCIe switches. - vLLM #39010: Hang During CUDA Graph Capture on ROCM.
- vLLM #40081: vLLM fails to start on RDNA 4 (gfx1201) inside containers.
Community Resources
- Reddit: Qwen3.5-122B-A10B GPTQ Int4 on 4× Radeon AI PRO R9700 — Post by u/grunt_monkey_ showing 41 tok/s on the same hardware. Credits to u/djdeniro, u/sloptimizer, and u/Ok-Ad-8976 for the AITER recipe.
Key People Working On This
srinivamd— Root cause analysis of the RCCL tuning index bugvllmellm— IdentifiedNCCL_PROTO=Simpleworkaround0xb3-err— Dockerfile for RCCL downgrade workaroundkyuz0— Original reporter of vLLM #40980grunt_monkey_— Reddit post demonstrating Qwen3.5-122B at 41 tok/s on 4× R9700
Workarounds Summary
| Workaround | Applies To | How To |
|---|---|---|
NCCL_PROTO=Simple |
vLLM TP≥2 deadlock (Approach A) | docker run -e NCCL_PROTO=Simple ... |
| AITER + vllm-dev | vLLM TP≥2 deadlock (Approach B) | Use rocm/vllm-dev:upstream_preview_releases_v0.17.0_20260303 |
| RCCL downgrade to 7.1.1 | vLLM TP≥2 deadlock | Replace librccl.so in container |
| Host reboot | GPU state corruption | sudo reboot |
NCCL_P2P_DISABLE=1 |
rccl-tests only | export NCCL_P2P_DISABLE=1 |
Generated from real benchmark runs on amdrocm (Threadripper 7980X, 4× R9700, ROCm 6.4.3)
Reddit reference: Qwen3.5-122B-A10B GPTQ Int4 on 4× Radeon AI PRO R9700
Huge Thanks
Thanks to the Level1Techs community for supporting these projects so I can figure stuff out.
What’s next??
MTP. We have to get MTP working. 4 GPUs? 128gb? MTP for 30-60% speedup and agentic and/or coding tasks? We need this!
