Ultimate 2x R9700 guide for Qwen 3.8 27b MXFP4: ~200t/s c=1; 500t/s at c=8


How-to: turn 2x R9700s into a fast local LLM rig (Qwen3.8-27B MXFP4, ~200 tok/s single, ~500 tok/s at 8 concurrent)

Radiance on 2x Radeon AI PRO R9700: A Setup and Benchmark Guide

The inference engine used here has an interesting lineage. The original project was vllm-radiance (StillDeadcode/vllm-radiance) – a patched fork of vLLM that added hand-written RDNA4 HIP kernels (libr4d) for attention, gated-delta-net, all-reduce, MXFP4, and DFlash speculative decoding. It worked well but seems like it was ultimately limited by being bolted onto vLLM’s Python/PyTorch stack. And there might have been some friction in the vLLM pull requests… haha.

The author then ditched vLLM entirely and rewrote everything from scratch as Radiance (StillDeadcode/radiance) – a modular C++/HIP inference engine with a plugin architecture. Model architectures, kernel libraries, and quantizers are all .so plugins loaded at runtime. The Docker image is just 423 MB (vs vLLM’s 13+ GB). The shared GPU kernel library libr4d which you might recognize from the vllm-radiance days (and is now bundled inside the new stand-alone Radiance).

The result is a lean, purpose-built engine for RDNA4 that delivers substantially better performance than any vLLM-based approach on these cards.

This is a step-by-step for a two-card AMD Radeon AI PRO R9700 host. By the end you will have a local OpenAI-compatible server running Qwen3.8-27B dense in MXFP4 with speculative decoding, and you will have run a real end-to-end coding benchmark (a Breakout clone from a 4,500-word spec) to prove it works. Everything here was run twice: once on my main box and once on a second, freshly set up 2x R9700 machine to make sure the steps actually replicate.

The short version of why this stack is fast: the R9700 has more compute than it can feed from its 32 GB of GDDR6, so the whole game is shrinking the bytes that move per token and getting more tokens out of each pass. MXFP4 quantization halves the weight bytes, a DFlash2 drafter gets ~3-4 accepted tokens per forward pass, and the radiance kernel stack keeps everything on the card’s own memory bus instead of crossing PCIe. Same model on llama.cpp or stock FP8 runs 17-23 tok/s single-stream; this config runs 150-200-280.

What you need

  • 2x AMD Radeon AI PRO R9700 (32 GB, gfx1201/RDNA4). One card works too (TP1, smaller context);
    the repo supports 1/2/4/8 foreshadowing.
  • A Linux host with the amdgpu kernel driver working (you can see /dev/kfd and /dev/dri). I
    validated on CachyOS (kernel 7.1.5) and Ubuntu 24.04. You do NOT install ROCm on the host; the
    ROCm userspace ships inside the container image.
  • Docker (or podman). Any recent version.
  • ~60 GB free disk: 19 GB source checkpoint + 19 GB built checkpoint + 2 GB drafter + ~10 GB
    image. The 19 GB source is deletable after setup and the script prints the command.
  • No host Python, no HuggingFace CLI, no build tools. Setup runs everything inside the image.
  • ROCm vs Vulkan? ROCm here. It’s getting super legit on RDNA.

Step 1: sanity-check the GPUs

ls -la /dev/kfd /dev/dri
groups   # you want render and video in the list

If /dev/kfd is missing, the amdgpu driver is not loaded; fix that first, this guide cannot help
you there. If you are not in the render/video groups, add yourself and re-login.

Step 2: Pull the Radiance Image

docker pull stilldeadcode/radiance:latest
# ~423 MB

Step 3: Download a Model Container

Radiance uses .rad container files – single-file model packages with weights, config, and metadata baked in. Pre-built containers are on Hugging Face under StillDeadcode.

Option A: Qwen3.8-27B FP8 Dense (29 GiB) – Best single-stream

curl -L -o qwen3.8-27b-fp8.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-27b-fp8/resolve/main/qwen3.8-27b-fp8.rad

Option B: Qwen3.8-27B MXFP4 Dense (18 GiB) – Smaller, similar speed

curl -L -o qwen3.8-27b-mxfp4.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-27b-mxfp4/resolve/main/qwen3.8-27b-mxfp4.rad

Option C: Qwen3.8-Flash-Next FP8+IQ4R MoE (114 GiB) – Largest model, highest quality

curl -L -o qwen3.8-next-flash-fp8-iq4r-moe.rad \
  https://huggingface.co/StillDeadcode/qwen3.8-next-flash-fp8-iq4r-moe/resolve/main/qwen3.8-next-flash-fp8-iq4r-moe.rad

Step 4: Launch the Server

Qwen3.8-27B FP8 Dense (fastest single-stream, ~130 tok/s sustained)

RENDER_GID=$(getent group render | cut -d: -f3)
VIDEO_GID=$(getent group video | cut -d: -f3)

docker run -d \
  --device /dev/kfd --device /dev/dri \
  --group-add ${RENDER_GID} --group-add ${VIDEO_GID} \
  --ipc host --network host --shm-size 32g \
  --security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
  --ulimit memlock=-1:-1 \
  -e HIP_VISIBLE_DEVICES=0,1 \
  -v /path/to/qwen3.8-27b-fp8.rad:/model.rad:ro \
  --name radiance \
  stilldeadcode/radiance:latest \
  --model /model.rad \
  --tp 2 --tp-wire wht6 \
  --max-model-len 262144 \
  --max-num-seqs 8 \
  --kv-cache-dtype fp8 \
  --gpu-headroom-mib 96 \
  --num-speculative-tokens 7 \
  --host 0.0.0.0 --port 8300

Qwen3.8-Flash-Next MoE (larger context, higher quality, ~125 tok/s)

RENDER_GID=$(getent group render | cut -d: -f3)
VIDEO_GID=$(getent group video | cut -d: -f3)

docker run -d \
  --device /dev/kfd --device /dev/dri \
  --group-add ${RENDER_GID} --group-add ${VIDEO_GID} \
  --ipc host --network host --shm-size 32g \
  --security-opt seccomp=unconfined --cap-add=SYS_PTRACE \
  --ulimit memlock=-1:-1 \
  -e HIP_VISIBLE_DEVICES=0,1 \
  -v /path/to/qwen3.8-next-flash-fp8-iq4r-moe.rad:/model.rad:ro \
  -v /home/w/.cache/huggingface:/root/.cache/huggingface \
  --name radiance-fn \
  stilldeadcode/radiance:latest \
  --model /model.rad \
  --tp 2 --tp-wire wht6 \
  --max-model-len 200000 \
  --max-num-seqs 8 \
  --kv-cache-dtype fp8 \
  --placement expert_tiered \
  --host-pool-mib 12288 \
  --gpu-headroom-mib 96 \
  --ngram-placement ram \
  --num-speculative-tokens 3 \
  --max-num-batched-tokens 2048 \
  --host 0.0.0.0 --port 8300

Note on first start: The first load reads the entire .rad file into VRAM and pinned host memory. For Flash-Next this is ~66 GiB of data at ~0.3 GB/s over NFS, taking about 4 minutes. Subsequent starts with the file cached by the OS are faster.


Step 5: Verify It’s Running

docker ps
curl http://localhost:8300/v1/models
# Expected: {"object":"list","data":[{"id":"qwen35","object":"model",...}]}

Step 6: Sanity Check

python3 -c "
import json, time, urllib.request
t0 = time.time()
req = urllib.request.Request(
  'http://localhost:8300/v1/chat/completions',
  data=json.dumps({
    'model': 'qwen35',
    'messages': [{'role': 'user', 'content': 'Write a short story about a sentient dot matrix printer.'}],
    'max_tokens': 512,
    'temperature': 0
  }).encode(),
  headers={'Content-Type': 'application/json'},
  method='POST'
)
resp = urllib.request.urlopen(req, timeout=300)
t1 = time.time()
result = json.loads(resp.read())
tokens = result['usage']['completion_tokens']
print(f'{tokens} tokens in {t1-t0:.1f}s = {tokens/(t1-t0):.1f} tok/s')
"

Reference Benchmarks (2x R9700, Radiance v1.2.3, 96 GiB DRAM)

All benchmarks use the OpenAI-compatible API with temperature 0. “Sustained decode” numbers use long generations (4096+ tokens) to amortize prefill overhead.

Qwen3.8-27B FP8 Dense (–num-speculative-tokens 7, DFlash2)

Test Result
Single stream, 512 gen 111 tok/s (includes prefill)
Sustained decode, 8192 gen 130 tok/s
Decode-only (1-token prompt) 169 tok/s
conc=2 177 tok/s agg
conc=4 291 tok/s agg
conc=8 596 tok/s agg
2K prompt decode 23 tok/s
200K prompt decode 1.7 tok/s
VRAM used 15.0 GiB/GPU

Qwen3.8-27B MXFP4 Dense (–num-speculative-tokens 7, DFlash2)

Test Result
Single stream, 512 gen 133 tok/s
conc=4 448 tok/s agg
conc=8 678 tok/s agg
VRAM used 9.7 GiB/GPU

Qwen3.8-Flash-Next FP8+IQ4R MoE (–num-speculative-tokens 3, MTP)

Test Result
Single stream, 512 gen 125 tok/s
conc=4 513 tok/s agg
conc=8 879 tok/s agg
conc=16 1103 tok/s agg
20K prompt decode 25 tok/s
VRAM used 21.9 GiB/GPU + 47.7 GiB host RAM (n-gram)

In some very specific benchmaxxed configs its possible to break out (hah, get it) past 200 t/s but for real-world usage these number sare much more usable.

Over time, I also found the FP8 dense version better than MXFP4 dense.

Breakout coding benchmark (FP8 dense)

The breakout test is a 5,899-token programming task asking the model to build a complete Breakout game in a single HTML file.

Model tok/s Wall time Result
FP8 dense (this guide) 125.7 260.8s Complete, 32,768 tokens, hit max_tokens still generating
MXFP4 dense (old vllm-radiance) 191.8 106.6s Complete, 20,443 tokens, early EOS, other quirks
Flash-Next FP8 (old vllm-radiance TP8) 42.4 N/A Failed – burned budget on planning

The FP8 dense result is strong: the model stayed engaged for the full 32K budget, meaning it was productively generating code the whole time.

Flash-next FP8 can be good, but it’s much slower, spends much more time reasoning. I just included it for comparison. It is a much larger model, and will be shown in the 8xR9600D video.


Step 7: Run the Breakout Test

TODO upload prompt zip here

Tuning Tips

Context length vs concurrency

The --max-model-len and --max-num-seqs flags trade off between deep context and concurrent users. The KV cache budget is roughly card_total - headroom - weights - activation_arena. For FP8 dense (15 GiB weights) at 262K context with 8 sequences, the pool backs about 37% of worst case. Reduce --max-model-len or --max-num-seqs to increase the backing ratio.

n-gram placement (Flash-Next only)

The 51B-parameter n-gram/PLE table is 47.68 GiB. Options:

  • --ngram-placement ram (recommended): Load into pinned host RAM. Requires ~48 GiB free DRAM. Fastest decode.
  • --ngram-placement disk: Stream from disk via io_uring. Uses no RAM but adds latency.
  • --ngram-placement vram: Put on GPU. Needs free VRAM (unlikely with 32 GB cards).

Lossy all-reduce

--tp-wire wht6 uses Walsh-Hadamard-rotated 6-bit quantization for cross-GPU communication. Lossy for messages >128 KiB (prefill and busy decode). Single-sequence decode is exact. Use --tp-wire exact for bit-exact reproducibility at ~5-10% throughput cost.


Troubleshooting

Symptom Likely cause Fix
E unknown option Typo in flag name in the radiance docs Check docker run --rm stilldeadcode/radiance:latest --help
libavx: cannot make a ring io_uring needs more locked memory Add --ulimit memlock=-1:-1 to docker run
Container exits immediately Missing --ulimit memlock or wrong flag Check logs with docker logs <container>
400 invalid_request_error Prompt exceeds --max-model-len Shorten prompt or increase --max-model-len

Resources

Footnote for the curious: the same guide validates on the ASRock 8x R9600D box (TP2 on 2 of the 8 cards, 4 sets, with model router) produced identical KV sizing and byte-identical benchmark output, within normal run-to-run variance on speed. Two cards are the sweet spot for this model; the rest of a bigger box is better spent on other instances, other models, or scaling for many MANY users. Video on that soon.

3 Likes

Since this is compatible with the Open AI connections, can I drop Turnstone on top of this for a model mixture? Assuming I have enough hardware headroom?

Yep this plus a small fast judge model is bananas

1 Like

A 1x R9700 follow-up for folks without the second card could be interesting; It could be a killer combination for cheap for serving Qwen3.8-Flash-Next, and hopefully GLM-5.3-Flash eventually.

You can put one R9700 into a lot of existing setups without too much hassle, but 2x R9700 would probably have to be a dedicated rebuild regarding PCIe lanes, slot layout, and PSU headroom.

Hey there, long time fan of the channel and the author of the Radiance engine. Super cool to see you talking about this :+1:

4 Likes

More exciting things are also in the works for Radiance, we just added EmbeddingGemma2 support. Longer term plans are full gguf compatibility, and support for more hardware. Currently thinking of targeting strix halo next. Feel free to hang out with us at the Launch80 discord, where we talk about all this stuff.

3 Likes

I run Qwen3.8-Flash-Next on Radiance and it’s so damn fast and gets even faster every day!

Asrock Rack ROMED8-2T
Epyc 7F52
2xR9700 Pro @240W
256GB 8x32GB 3200 DDR4

Just as a PSA, the Swift models are up for download now:

1 Like

nice guide. a few things that bit me on a 2x R9700 box (AM4, X570, 5700X) that might save others time:

check that p2p is actually on. radiance falls back to staging through host memory when it isn’t, so it still works but slower. the log tells you (“reach each other’s memory directly (peer to peer)”). some distro kernels ship without CONFIG_PCI_P2PDMA (unraid does, debian’s default did for someone on discord), then no bios setting will help. if your board has no rebar toggle, pci=realloc,nocrs plus resizing BAR0 before amdgpu binds works, but don’t unbind/rebind amdgpu to do it, that locks kfd until a reboot.

–ngram-placement disk is only fine on nvme. with the .rad on a hard drive I was at ~110 ms per decode step instead of ~15.

on 1.2.x clients that send reasoning_effort “high” get a 400, –reasoning-effort-map high=xhigh fixes that.

and for comparing runs, step time from /stats is a lot more stable than tok/s, since tok/s moves with the draft acceptance rate. flash-next here: ~15 ms/step, ~140 t/s single stream at a 210 W power cap.

One handy tip…don’t forget to set up your prefix cache:

      --prefix-cache-host-mib 8192 --prefix-cache-dir /data/kvcache/swiftnext
      --prefix-cache-disk-mib 131072

It won’t create the data/kvcache directory, so the subdirectory creation will fail (there will be a log entry if that happens).

The 128GB disk cache is a default, I haven’t used it enough to work out whether that’s actually necessary or not. I changed the system RAM value from 4GB to 8GB, just because I’ve got a bit of spare RAM after the model’s loaded.

1 Like

Bananas doesn’t even begin to cover it. I have never seen OpenWebUI stutter and struggle like it does here. I have a dedicated RTX A4000 running Gemma E4B as a judge model (which funny enough is mostly for performance reasons) and it can’t even keep up with the output from this setup. :distorted_face:

I always knew something like this was possible given the hardware, and it sat in the back of my mind as a, “Well, maybe someday…”

Today is that day.

1 Like

hello and welcome! I have the unpublished guide almost ready. I have an 8x R9600D system we can use for testing if need be (things dont slice evenly beyond tp=4 but tp=4 is loads of fun let me tell you)

Any color you want to add to the how-to on the vllm > radiance transition, lol? I saw some.. spicy pull requests vllm rejected, adjacent to the work you were doing, haha

also bumped your trust level if you wanna set icons, link to project, whathave you.

1 Like

Sure, perhaps a little bit of background on the whole vllm to standalone story.

I got some R9700s earlier this year and at the beginning just installed vLLM. Performance was pretty meh on Qwen3.6 27B FP8 at the time. Long context attention was completely crushing to the point where even going beyond 128K context was basically impossible as things slowed to a crawl. Cold starts took >13min and even warm starts took a few minutes.

Then i saw a reddit post that spread the word about the now pretty widely known AITER patch. Someone realized that the AITER attention kernels originally meant for CDNA actually run on RDNA4 and outperform the standard kernels vLLM selected at the time by a wide margin. That was when i started vllm-radiance. My goal was to make vLLM not suck for RDNA4, and so i embarked on a long journey of patches. Fast forwarding a bit i had ~20 patches and one of the fastest vLLM forks around. It used a sideproject i had started called libr4d which was a collection library of gfx1201 optimized HIP kernels for various operations (attention, GDN etc), i also had a significantly quantized drafter, adaptive depth drafting, and a fix for that awfull long cold start (there was a fallback in one place where amd would fall back to an F32 operation which then caused a huge amount of LLVM JIT compilations during vLLMs autotuning stage, lowering it to BF16 did the job, there was also no reason to have that op in F32 but vLLM is primarily optimized for nvidia so it is what it is).

Then i started working on a 4bit version of Qwen3.6 27B and this is what finally brought me over the edge of starting my own project. I found myself with many many patches and the realization that while this was working, if vLLM ever updates i will have to reconsider my life choices. Also working in the vLLM codebase was generally hell, everything is way too dynamic, there is no centralized anything and even just trying to find out where something gets dispatched from is an actual adventure. Additionally they chose python to write a high performance multithreaded inference engine, which in itself is already a bad choice in my oppinion. It was not rare to see vLLM pin multiple cpu cores to 100%.

Thats when i started radiance. The core idea was to use the general concept of vLLM (concurrency, batching, paged attention etc) in the tech stack of llama.cpp (C++/HIP). Additionally i wanted to ensure that it will not be neccesary to crate lots of forks when people inevitably want to extend the engine so i chose a modular core that provides all the base features like the api server and all the plumbing and then a plugin system for model architectures and gpu kernels so people can retrofit those into the engine without having to fork. Around a thousand commits (to my internal repo) and three months of work later i released radiance publicly at StillDeadcode/radiance: modular inference engine for people that want to get the most out of their hardware - Codeberg.org .

My goal with this project is also not to support every model on every possible hardware, i want to just run a few select models as fast as possible on the hardware i use and i want to empower others to do the same without having to reinvent the wheel.

So far the open source developement has been pretty good, but since the project is still so fresh off the press, we are still dealing with a lot of initial bugs like client incompatibilities because of missing endpoint features etc.

Hope this helps :+1:

7 Likes

Well, I for one very much appreciate your efforts, and I’m pretty damn glad you gave up on vLLM :wink:

As an aside…having spent a lot of time dealing with llama.cpp, I really like the fact that radiance will happily peg the disks at near full-speed. Loading a model at 1.5GB/s when you know your SSD can do 6GB/s+ is one of the most frustrating irritations possible.

I do have one question, though: do you think you’ll ever incorporate an equivalent of llama.cpp’s router mode, where models can be swapped in and out at runtime on request? The use case is pretty simple…I find that Swift/Qwen Flash Next is very creative and good for planning, but not as reliable as 27B when it comes to writing solid code, so being able to assign them to separate plan/build agents (without resorting to a separate management layer) would be a godsend.

EDIT: And now my brainworm for the night is…wondering if it would be possible to vibe-code a plugin for that other family of cards that are famously under-performing relative to their hardware capabilities: Intel Arc.