Converting a data center GPU for desktop use

The first time I used Vulcan was in 2013 for the game Battlefield 4. Then it was called Mantle API and gave my then Radeon 7970 ghz a big fps boost.

It is just one way for a software to talk to the GPU hardware. It was mostly about gaming performance in the beginning. Then it became general GPU api with support for most models. I belive that is why Copilot is off. Are you pay for Copilot? It seems to be confidently wrong.

MI25 has the same memory bandwith of a Pro V620. Other than that the Pro V620 is faster, but more expansive.

Same total (16 gb) ram, double the compute actually. Which is a win for me @ $40 each. It’s essentially 2x mi25 with half memory each. I wanted to be able to run big models, and 64gb of hbm2 for $160 was by far the cheapest option I could find.

There is also the (hard to find) v340 (non-L), which is 32gb total, and is just actually 2x mi25, but those are $400-700 on ebay when you can find them from what I’ve seen.

Maybe I am reading this wrong, but it appears the V620 is in essence two cards with 8GB each in one package. Pretty cool for video, but awkward for compute. Can we “shard” a model across the two compute units and get the same throughput?

Wow that’s a crazy similar build lol, x99, open air testbench, the red gskill kit, and everything lol. A bit jelly you have four cards though. I’ll definitely check out your perf testing, may do some comps against the same models.

One thing I may not have made clear in my earlier replies, I flashed the cards because I had issues booting with both cards (likely due to the crappy b350 board I was running on at the time), not because of any issues I had with llama.cpp and the mi25 vbios.

You should be able offload different layers to different GPUs (Intel has Battlematrix for this), and while it does scale better than some tasks like the infamous example of real-time graphics, it’s still not perfect.

I don’t have multiple GPUs in a single system to test, but it’s possible that as the model needs to run sequentially, this may not scale compute that well in that configuration. This was my understanding of how these networks work, but my knowledge on the internals is getting old.

It really depends what you want to do. I wouldn’t use these v340l for any actual graphics tasks unless I was going to be doing at most 1 physical gpu per virtual gpu. Dividing down is good, trying to combine gets messy quickly. To use for actual video out and graphics, the single gpu is always going to be a better choice.

For compute/LLM however… using llama.cpp with layer split strategy, it simply divides the model layers out among the cards, and does compute with a single card at a time on the card where the layer in question sits. In practice this means it behaves like a single gfx900 but with 64gb of hbm2 vram. It also makes heat basically a non-issue, because each card has a compute duty cycle of 1/# of cards. Power use is base system + 30w*# of cards + 80w max for the card running compute at a given moment.

You can also do tensor parallelism to multiply the compute available by dividing a given layer among multiple cards and doing parallel processing, but you need the pcie bandwidth to do it. This means you need both enough pcie bandwidth in the first place, and also the bandwidth requirement scales linearly with the number of gpus, so more powerful fewer units are better. So far with llama.cpp I can combine 4 gfx900 without any bandwidth limitations on x99. I’m now playing with vllm as it’s supposed to be much more efficient using the bandwidth, so i’m hoping I can reach a higher ceiling. In my case I’m trying to see what the limits are with a low enough budget that this is accessible to almost anyone, my entire system build is < $400.

TL;DR: the v340l is unmatched for vram and compute per $, but by the same token it’s a disadvantageous way to assemble the compute unless it can still fit within your use case without compromise. so far that limit is 4x gfx900 for llama.cpp tensor parallelism.


My 24gb Tesla M40s just have dell blower server fans held on with copper tape, i had plans and designs to 3d print and mount but its only for me to play with so tape is good enough. rest of the pc is 2x 2699v3 18 core xeons, cheap dual socket dual channel x99 board, 1000w power supply and 128gb ram. the fans are controlled by the esp32 c3 board, the plan was to use a pair of i2c temperature sensors on the back side of the cards but because i didn’t have the sensors in at the time i just set a decent fan speed and left it and it ran fine, it wasnt quiet at idle but that didnt really matter and i still havent got around to fitting the sensors as it was good enough.

the case is a vintage Antec 900 Nine Hundred

1 Like

I was about to say, look at that steely!

Also, exactly how an ESP32 should work, hanging in the air with wires

2 Likes

Well, I am at the end with Microsoft Copilot for this exercise. (To be fair, this is a fairly exotic topic.) With the last iteration, tried using PyTorch directly, no LLM ran, and Copilot’s suggestions started to loop (again).

It did suggest a simple test to load the GPU, which got me temperatures under load. This MI25 BIOS got the card up to 220W at 100% GPU. Temperatures stayed just under 90C, and fan only got up to about 85%. I tweaked the fan-control script to be slightly more aggressive.

At full load, VRM temperatures were higher than the hotspot, but not by much (<10°C). I have not done anything extra to cool the VRM. Did change the script to use the higher of hotspot or VRM temperatures (rather than an override for VRMs).

(Example log from run under load is below.)

Started with Google’s Gemini, and the first thing it suggested was to use Vulcan. (Where Copilot was quite certain that Vulcan would not work.)

Gemini also suggests that some MI25 BIOS might allow up to 300W. Do not know if that is true, and not sure if the GPU fan would be adequate. Might go looking for a higher-power BIOS, but want to get LLMs running first.

So … next, Vulcan. :slight_smile:

Jun 20 10:32:33 beast mi25-fan-control.sh[439123]: edge: 42°C hotspot: 43°C vrm: 43°C pwm have:  18 last:  30 next:  30 want:  57
Jun 20 10:32:34 beast mi25-fan-control.sh[439123]: edge: 42°C hotspot: 43°C vrm: 43°C pwm have:  28 last:  40 next:  40 want:  57
Jun 20 10:32:35 beast mi25-fan-control.sh[439123]: edge: 42°C hotspot: 43°C vrm: 43°C pwm have:  39 last:  50 next:  50 want:  57
Jun 20 10:32:36 beast mi25-fan-control.sh[439123]: edge: 43°C hotspot: 43°C vrm: 43°C pwm have:  49 last:  57 next:  57 want:  57
Jun 20 10:32:51 beast mi25-fan-control.sh[439123]: edge: 47°C hotspot: 60°C vrm: 61°C pwm have:  56 last:  67 next:  67 want: 105
Jun 20 10:32:52 beast mi25-fan-control.sh[439123]: edge: 47°C hotspot: 66°C vrm: 64°C pwm have:  66 last:  77 next:  77 want: 130
Jun 20 10:32:53 beast mi25-fan-control.sh[439123]: edge: 50°C hotspot: 67°C vrm: 66°C pwm have:  75 last:  87 next:  87 want: 135
Jun 20 10:32:54 beast mi25-fan-control.sh[439123]: edge: 52°C hotspot: 68°C vrm: 66°C pwm have:  86 last:  97 next:  97 want: 140
Jun 20 10:32:55 beast mi25-fan-control.sh[439123]: edge: 52°C hotspot: 67°C vrm: 68°C pwm have:  96 last: 107 next: 107 want: 140
Jun 20 10:32:56 beast mi25-fan-control.sh[439123]: edge: 53°C hotspot: 69°C vrm: 68°C pwm have: 105 last: 117 next: 117 want: 145
Jun 20 10:32:57 beast mi25-fan-control.sh[439123]: edge: 52°C hotspot: 70°C vrm: 70°C pwm have: 115 last: 127 next: 127 want: 150
Jun 20 10:32:58 beast mi25-fan-control.sh[439123]: edge: 55°C hotspot: 70°C vrm: 70°C pwm have: 126 last: 137 next: 137 want: 150
Jun 20 10:32:59 beast mi25-fan-control.sh[439123]: edge: 56°C hotspot: 71°C vrm: 71°C pwm have: 136 last: 147 next: 147 want: 155
Jun 20 10:33:00 beast mi25-fan-control.sh[439123]: edge: 56°C hotspot: 72°C vrm: 71°C pwm have: 145 last: 157 next: 157 want: 160
Jun 20 10:33:01 beast mi25-fan-control.sh[439123]: edge: 60°C hotspot: 73°C vrm: 72°C pwm have: 156 last: 165 next: 165 want: 165
Jun 20 10:33:04 beast mi25-fan-control.sh[439123]: edge: 62°C hotspot: 75°C vrm: 74°C pwm have: 164 last: 175 next: 175 want: 175
Jun 20 10:33:09 beast mi25-fan-control.sh[439123]: edge: 64°C hotspot: 77°C vrm: 77°C pwm have: 173 last: 185 next: 185 want: 185
Jun 20 10:33:14 beast mi25-fan-control.sh[439123]: edge: 65°C hotspot: 78°C vrm: 79°C pwm have: 183 last: 195 next: 195 want: 195
Jun 20 10:33:23 beast mi25-fan-control.sh[439123]: edge: 66°C hotspot: 79°C vrm: 81°C pwm have: 194 last: 202 next: 202 want: 202
Jun 20 10:33:32 beast mi25-fan-control.sh[439123]: edge: 66°C hotspot: 81°C vrm: 83°C pwm have: 200 last: 208 next: 208 want: 208
Jun 20 10:33:51 beast mi25-fan-control.sh[439123]: edge: 68°C hotspot: 80°C vrm: 85°C pwm have: 207 last: 213 next: 213 want: 213
Jun 20 10:34:28 beast mi25-fan-control.sh[439123]: edge: 68°C hotspot: 82°C vrm: 87°C pwm have: 211 last: 219 next: 219 want: 219
Jun 20 10:35:36 beast mi25-fan-control.sh[439123]: edge: 68°C hotspot: 69°C vrm: 71°C pwm have: 217 last: 213 next: 213 want: 155
Jun 20 10:35:37 beast mi25-fan-control.sh[439123]: edge: 67°C hotspot: 69°C vrm: 67°C pwm have: 211 last: 207 next: 207 want: 145
Jun 20 10:35:38 beast mi25-fan-control.sh[439123]: edge: 66°C hotspot: 68°C vrm: 66°C pwm have: 205 last: 201 next: 201 want: 140
Jun 20 10:35:39 beast mi25-fan-control.sh[439123]: edge: 66°C hotspot: 67°C vrm: 65°C pwm have: 200 last: 195 next: 195 want: 135
Jun 20 10:35:40 beast mi25-fan-control.sh[439123]: edge: 63°C hotspot: 66°C vrm: 63°C pwm have: 194 last: 189 next: 189 want: 130
Jun 20 10:35:41 beast mi25-fan-control.sh[439123]: edge: 65°C hotspot: 65°C vrm: 62°C pwm have: 188 last: 183 next: 183 want: 125
Jun 20 10:35:42 beast mi25-fan-control.sh[439123]: edge: 61°C hotspot: 65°C vrm: 61°C pwm have: 181 last: 177 next: 177 want: 125
Jun 20 10:35:43 beast mi25-fan-control.sh[439123]: edge: 62°C hotspot: 64°C vrm: 61°C pwm have: 175 last: 171 next: 171 want: 120
Jun 20 10:35:44 beast mi25-fan-control.sh[439123]: edge: 60°C hotspot: 64°C vrm: 60°C pwm have: 170 last: 165 next: 165 want: 120
Jun 20 10:35:45 beast mi25-fan-control.sh[439123]: edge: 56°C hotspot: 62°C vrm: 59°C pwm have: 164 last: 159 next: 159 want: 110
Jun 20 10:35:46 beast mi25-fan-control.sh[439123]: edge: 58°C hotspot: 62°C vrm: 59°C pwm have: 158 last: 153 next: 153 want: 110
Jun 20 10:35:47 beast mi25-fan-control.sh[439123]: edge: 55°C hotspot: 61°C vrm: 58°C pwm have: 153 last: 147 next: 147 want: 105
Jun 20 10:35:48 beast mi25-fan-control.sh[439123]: edge: 57°C hotspot: 60°C vrm: 57°C pwm have: 145 last: 141 next: 141 want: 100
Jun 20 10:35:49 beast mi25-fan-control.sh[439123]: edge: 56°C hotspot: 60°C vrm: 57°C pwm have: 139 last: 135 next: 135 want: 100
Jun 20 10:35:50 beast mi25-fan-control.sh[439123]: edge: 55°C hotspot: 58°C vrm: 56°C pwm have: 134 last: 129 next: 129 want:  95
Jun 20 10:35:51 beast mi25-fan-control.sh[439123]: edge: 52°C hotspot: 57°C vrm: 56°C pwm have: 128 last: 123 next: 123 want:  92
Jun 20 10:35:52 beast mi25-fan-control.sh[439123]: edge: 54°C hotspot: 56°C vrm: 56°C pwm have: 122 last: 117 next: 117 want:  90
Jun 20 10:35:53 beast mi25-fan-control.sh[439123]: edge: 54°C hotspot: 56°C vrm: 55°C pwm have: 115 last: 111 next: 111 want:  90
Jun 20 10:35:54 beast mi25-fan-control.sh[439123]: edge: 54°C hotspot: 55°C vrm: 55°C pwm have: 109 last: 105 next: 105 want:  87
Jun 20 10:35:55 beast mi25-fan-control.sh[439123]: edge: 53°C hotspot: 54°C vrm: 54°C pwm have: 103 last:  99 next:  99 want:  85
Jun 20 10:35:56 beast mi25-fan-control.sh[439123]: edge: 53°C hotspot: 54°C vrm: 54°C pwm have:  98 last:  93 next:  93 want:  85
Jun 20 10:35:57 beast mi25-fan-control.sh[439123]: edge: 53°C hotspot: 54°C vrm: 54°C pwm have:  92 last:  87 next:  87 want:  85
Jun 20 10:35:58 beast mi25-fan-control.sh[439123]: edge: 53°C hotspot: 54°C vrm: 54°C pwm have:  86 last:  85 next:  85 want:  85
Jun 20 10:36:15 beast mi25-fan-control.sh[439123]: edge: 50°C hotspot: 51°C vrm: 51°C pwm have:  85 last:  79 next:  79 want:  77
Jun 20 10:36:16 beast mi25-fan-control.sh[439123]: edge: 50°C hotspot: 51°C vrm: 51°C pwm have:  77 last:  77 next:  77 want:  77
Jun 20 10:36:44 beast mi25-fan-control.sh[439123]: edge: 45°C hotspot: 48°C vrm: 48°C pwm have:  75 last:  71 next:  71 want:  70
Jun 20 10:36:45 beast mi25-fan-control.sh[439123]: edge: 48°C hotspot: 49°C vrm: 48°C pwm have:  69 last:  72 next:  72 want:  72
Jun 20 10:37:17 beast mi25-fan-control.sh[439123]: edge: 43°C hotspot: 46°C vrm: 46°C pwm have:  71 last:  66 next:  66 want:  65
Jun 20 10:37:18 beast mi25-fan-control.sh[439123]: edge: 46°C hotspot: 47°C vrm: 46°C pwm have:  64 last:  67 next:  67 want:  67
Jun 20 10:38:30 beast mi25-fan-control.sh[439123]: edge: 44°C hotspot: 44°C vrm: 44°C pwm have:  66 last:  61 next:  61 want:  60
Jun 20 10:38:31 beast mi25-fan-control.sh[439123]: edge: 42°C hotspot: 44°C vrm: 45°C pwm have:  60 last:  62 next:  62 want:  62

Turns out, getting an LLM running using the Vulcan backend is easy. So the excursion trying to use ROCm to run LLMs seems a waste.

Rebuilt llama.cpp to use Vulcan, and can run LLMs on the GPU.
No notion if these numbers are any good. :slight_smile:

During the tests, card temperatures stayed below 80°C.

Belatedly realized that Vulcan is using both the MI25, and the NVIDIA GeForce GTX 1050 (meant only for display).

GPT-OSS is odd, as it runs the 1050 flat out, while the MI25 is loafing.

----------------------------------------
Benchmark : Mistral-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Mistral-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/mistral-7b/mistral-7b-instruct-v0.2.Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp512 297.38 ± 0.44
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp2048 437.29 ± 0.90
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp4096 398.06 ± 0.65
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 tg128 39.13 ± 0.23
----------------------------------------
Benchmark : Llama-3-8B-Q4_K_M
Save to   : /home/preston/models/benchmark_Llama-3-8B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/llama3-8b/Meta-Llama-3-8B-Instruct-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp512 294.69 ± 0.44
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp2048 432.83 ± 0.36
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp4096 396.20 ± 0.57
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 tg128 37.35 ± 0.33
----------------------------------------
Benchmark : Qwen2.5-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Qwen2.5-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-7b/qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp512 301.47 ± 0.04
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp2048 449.88 ± 0.18
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp4096 425.19 ± 0.20
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 tg128 37.96 ± 0.30
----------------------------------------
Benchmark : Qwen2.5-Coder-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Qwen2.5-Coder-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-coder-7b/qwen2.5-coder-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp512 300.77 ± 0.02
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp2048 448.56 ± 0.16
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp4096 424.41 ± 0.29
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 tg128 37.79 ± 0.36
----------------------------------------
Benchmark : Gemma-4-12B-it-Q4_K_M
Save to   : /home/preston/models/benchmark_Gemma-4-12B-it-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gemma-4-12b/gemma-4-12b-it-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp512 181.89 ± 0.06
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp2048 259.72 ± 0.04
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp4096 233.31 ± 0.14
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 tg128 21.09 ± 0.20
----------------------------------------
Benchmark : GPT-OSS-20B-Q4_K_M
Save to   : /home/preston/models/benchmark_GPT-OSS-20B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gpt-oss-20b/gpt-oss-20b-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp512 193.57 ± 2.55
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp2048 225.67 ± 0.67
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp4096 223.70 ± 0.55
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 tg128 47.82 ± 1.50

Turns out, I can force Vulcan to use only the MI25, and the numbers improve.

$ GGML_VK_VISIBLE_DEVICES=1 bash download-mi25-models.sh 
----------------------------------------
Benchmark : Mistral-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Mistral-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/mistral-7b/mistral-7b-instruct-v0.2.Q4_K_M.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp512 480.81 ± 0.49
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp2048 454.78 ± 1.63
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 pp4096 412.46 ± 0.74
llama 7B Q4_K - Medium 4.07 GiB 7.24 B Vulkan 99 tg128 53.47 ± 1.03
----------------------------------------
Benchmark : Llama-3-8B-Q4_K_M
Save to   : /home/preston/models/benchmark_Llama-3-8B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/llama3-8b/Meta-Llama-3-8B-Instruct-Q4_K_M.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp512 471.10 ± 0.19
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp2048 448.63 ± 0.38
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 pp4096 409.78 ± 0.52
llama 8B Q4_K - Medium 4.58 GiB 8.03 B Vulkan 99 tg128 49.48 ± 0.41
----------------------------------------
Benchmark : Qwen2.5-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Qwen2.5-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-7b/qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp512 521.02 ± 0.49
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp2048 495.96 ± 0.31
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp4096 462.32 ± 0.33
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 tg128 50.60 ± 0.52
----------------------------------------
Benchmark : Qwen2.5-Coder-7B-Q4_K_M
Save to   : /home/preston/models/benchmark_Qwen2.5-Coder-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-coder-7b/qwen2.5-coder-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp512 519.39 ± 0.41
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp2048 495.68 ± 0.58
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 pp4096 462.14 ± 0.17
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B Vulkan 99 tg128 50.54 ± 0.19
----------------------------------------
Benchmark : Gemma-4-12B-it-Q4_K_M
Save to   : /home/preston/models/benchmark_Gemma-4-12B-it-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gemma-4-12b/gemma-4-12b-it-Q4_K_M.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp512 290.11 ± 0.55
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp2048 272.17 ± 0.36
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 pp4096 256.51 ± 0.03
gemma4 ?B Q4_K - Medium 6.62 GiB 11.91 B Vulkan 99 tg128 28.75 ± 0.01
----------------------------------------
Benchmark : GPT-OSS-20B-Q4_K_M
Save to   : /home/preston/models/benchmark_GPT-OSS-20B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gpt-oss-20b/gpt-oss-20b-Q4_K_M.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
model size params backend ngl test t/s
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp512 761.52 ± 8.28
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp2048 716.74 ± 0.21
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp4096 664.73 ± 0.91
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 tg128 69.39 ± 0.54

Setup llama.cpp in “router” mode (user can choose between models) to serve my local subnet. Performance is surprisingly good.

This is where I was trying to go at the start of the exercise.

Documented setup: local-LLM-service

(Yes, I am easily amused.)

Updated: local-LLM-service

Updated the list of models downloaded and benchmarked. I am not remotely expert in this area, so picked models that looked interesting or significant.

Download step runs one prompt, with the aim of getting some comparative notion of startup and response times, and of the quality of the answers.

Benchmark command was recommended, no notion as to how useful. At least it offers a comparison.

Suggestions welcome! :slight_smile:

Have an AWK script to summarize the log from the exercise.

[EDIT: Omitted log as obviated by later work.]

Should note, posted on llama.cpp thread, and got some interesting response. To summarize:

  • For LLMs, ROCm should give better performance than Vulcan.
  • The ArchLinux build of ROCm kept gfx900 support.
  • ROCm 7.13 re-adds gfx900 support(!!).
  • The “turboquant” builds of llama.cpp can improve performance.

Seems there is another episode to this exercise…

Yeah the re-added support has been great. I use 7.1.1 with the gfx906. AWQ builds are promising for me at least on the gfx906 llama.cpp build in llama-swap I’m using

1 Like

FWIW TQ does not provide much benefit over the existing KV cache quants in llama.cpp. Anbeeld (owner of beellama) has done some great writeup with extensive tests comparing it:

The existing quants also have the advantage of having better kernel support for AMD in both Vulkan and ROCm when compared to TQ.

1 Like

The v620 is a single GPU v710. The v320[L] is the dual GPU variant, just like the v520 and v540. They are all intended for VDi but have shown to have some good use cases for LLMs and Machine Learning/Vision. You may want to check out Country Boy Computers as he is doing the lord’s work of trying these niche cards.

1 Like

Debian also compiles all pre-existing gfx versions in their maintained version of ROCm. They just note in the wiki what things may not work or need manual tweaks if there is an incompatibility between new and old libraries for any particular architecture.

I use Debian GNU/Linux and ArchLinux for my setups.

I have a v340, BC-160, Mi25, v540, W5700 Pro, a RX 5700, and a RX 9070 XT.

1 Like

I did some similar testing on the mi25s with the wx9100 vbios on, same launch commands.
Main card results (also driving gui)

model size params backend ngl test t/s
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp512 301.19 ± 12.66
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp2048 369.39 ± 1.41
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp4096 363.08 ± 5.12
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 tg128 36.32 ± 0.85

Other card

model size params backend ngl test t/s
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp512 300.15 ± 14.97
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp2048 371.28 ± 3.00
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp4096 369.90 ± 0.22
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 tg128 41.13 ± 0.14

Both cards:

model size params backend ngl test t/s
qwen35 9B Q4_K - Medium 5.23 GiB 8.95 B Vulkan 99 pp512 300.16 ± 10.58
qwen35 9B Q4_K - Medium 5.23 GiB 8.95 B Vulkan 99 pp2048 346.19 ± 5.38
qwen35 9B Q4_K - Medium 5.23 GiB 8.95 B Vulkan 99 pp4096 329.79 ± 1.36
qwen35 9B Q4_K - Medium 5.23 GiB 8.95 B Vulkan 99 tg128 42.57 ± 0.05

Couple takeaways from my testing:

1: Vulkan version is important, while not shown in this testing different versions of vulkan gave me 30% lower t/s. Might be an interesting point for more testing.
2: There’s a bit of overhead for the main card since it’s running the display, that’s why I included both cards separately.
3: There doesn’t seem to be a big jump in the score for both cards together, both cards seem to be getting utilized, so this could just be layer split being inefficient.
4: In all the tests the cards seem to be hitting the vbios power limit of 140w, so there may be additional headroom by reverting back.
5. Prompt processing seems to be lower across the board

I may try testing out a rocm build then going back the mi25 vbios and see how that performs

ROCm is pretty much always slightly faster than Vulkan at prompt processing, but Vulkan usually has a major advantage in MoE token generation (not helpful here, since you’re using the 9B dense model). You’ll probably find that Qwen 3.6 35B shows a bigger difference between them.

1 Like