Turns out, getting an LLM running using the Vulcan backend is easy. So the excursion trying to use ROCm to run LLMs seems a waste.
Rebuilt llama.cpp to use Vulcan, and can run LLMs on the GPU.
No notion if these numbers are any good. 
During the tests, card temperatures stayed below 80°C.
Belatedly realized that Vulcan is using both the MI25, and the NVIDIA GeForce GTX 1050 (meant only for display).
GPT-OSS is odd, as it runs the 1050 flat out, while the MI25 is loafing.
----------------------------------------
Benchmark : Mistral-7B-Q4_K_M
Save to : /home/preston/models/benchmark_Mistral-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/mistral-7b/mistral-7b-instruct-v0.2.Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| llama 7B Q4_K - Medium |
4.07 GiB |
7.24 B |
Vulkan |
99 |
pp512 |
297.38 ± 0.44 |
| llama 7B Q4_K - Medium |
4.07 GiB |
7.24 B |
Vulkan |
99 |
pp2048 |
437.29 ± 0.90 |
| llama 7B Q4_K - Medium |
4.07 GiB |
7.24 B |
Vulkan |
99 |
pp4096 |
398.06 ± 0.65 |
| llama 7B Q4_K - Medium |
4.07 GiB |
7.24 B |
Vulkan |
99 |
tg128 |
39.13 ± 0.23 |
----------------------------------------
Benchmark : Llama-3-8B-Q4_K_M
Save to : /home/preston/models/benchmark_Llama-3-8B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/llama3-8b/Meta-Llama-3-8B-Instruct-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| llama 8B Q4_K - Medium |
4.58 GiB |
8.03 B |
Vulkan |
99 |
pp512 |
294.69 ± 0.44 |
| llama 8B Q4_K - Medium |
4.58 GiB |
8.03 B |
Vulkan |
99 |
pp2048 |
432.83 ± 0.36 |
| llama 8B Q4_K - Medium |
4.58 GiB |
8.03 B |
Vulkan |
99 |
pp4096 |
396.20 ± 0.57 |
| llama 8B Q4_K - Medium |
4.58 GiB |
8.03 B |
Vulkan |
99 |
tg128 |
37.35 ± 0.33 |
----------------------------------------
Benchmark : Qwen2.5-7B-Q4_K_M
Save to : /home/preston/models/benchmark_Qwen2.5-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-7b/qwen2.5-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp512 |
301.47 ± 0.04 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp2048 |
449.88 ± 0.18 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp4096 |
425.19 ± 0.20 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
tg128 |
37.96 ± 0.30 |
----------------------------------------
Benchmark : Qwen2.5-Coder-7B-Q4_K_M
Save to : /home/preston/models/benchmark_Qwen2.5-Coder-7B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/qwen2.5-coder-7b/qwen2.5-coder-7b-instruct-q4_k_m-00001-of-00002.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp512 |
300.77 ± 0.02 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp2048 |
448.56 ± 0.16 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
pp4096 |
424.41 ± 0.29 |
| qwen2 7B Q4_K - Medium |
4.36 GiB |
7.62 B |
Vulkan |
99 |
tg128 |
37.79 ± 0.36 |
----------------------------------------
Benchmark : Gemma-4-12B-it-Q4_K_M
Save to : /home/preston/models/benchmark_Gemma-4-12B-it-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gemma-4-12b/gemma-4-12b-it-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| gemma4 ?B Q4_K - Medium |
6.62 GiB |
11.91 B |
Vulkan |
99 |
pp512 |
181.89 ± 0.06 |
| gemma4 ?B Q4_K - Medium |
6.62 GiB |
11.91 B |
Vulkan |
99 |
pp2048 |
259.72 ± 0.04 |
| gemma4 ?B Q4_K - Medium |
6.62 GiB |
11.91 B |
Vulkan |
99 |
pp4096 |
233.31 ± 0.14 |
| gemma4 ?B Q4_K - Medium |
6.62 GiB |
11.91 B |
Vulkan |
99 |
tg128 |
21.09 ± 0.20 |
----------------------------------------
Benchmark : GPT-OSS-20B-Q4_K_M
Save to : /home/preston/models/benchmark_GPT-OSS-20B-Q4_K_M.txt
----------------------------------------
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -m /home/preston/models/gpt-oss-20b/gpt-oss-20b-Q4_K_M.gguf
ggml_vulkan: Found 2 Vulkan devices:
ggml_vulkan: 0 = NVIDIA GeForce GTX 1050 (NVIDIA) | uma: 0 | fp16: 0 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: none
ggml_vulkan: 1 = AMD Radeon Instinct MI25 (RADV VEGA10) (radv) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 64 | shared memory: 65536 | int dot: 0 | matrix cores: none
| model |
size |
params |
backend |
ngl |
test |
t/s |
| gpt-oss 20B Q4_K - Medium |
10.81 GiB |
20.91 B |
Vulkan |
99 |
pp512 |
193.57 ± 2.55 |
| gpt-oss 20B Q4_K - Medium |
10.81 GiB |
20.91 B |
Vulkan |
99 |
pp2048 |
225.67 ± 0.67 |
| gpt-oss 20B Q4_K - Medium |
10.81 GiB |
20.91 B |
Vulkan |
99 |
pp4096 |
223.70 ± 0.55 |
| gpt-oss 20B Q4_K - Medium |
10.81 GiB |
20.91 B |
Vulkan |
99 |
tg128 |
47.82 ± 1.50 |