GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP)

TL;DR: This fork promises up to to double prompt processing speeds for deep infill on gfx906 based setups.

The title pretty much says it all. Since my setup is based on a wild arrangement of two Radeon VII and a laptop with an internal 8gb rtx3080 maxq, i have since worked with GLM to optimize said setup to try and get the most out of it.

Disclaimer: I did not code this myself but only steered GLM 5.3 towards the result but i have been using it in production ever since. I do not like talking about things i barely understand and this is really at the edge of what is comprehensible for me as a lowly IoT engineer.

Setup: Linux, Qwen 3.8 27B at Q6_K_XL, llama.cpp and opencode with “oh-my-openagent” but customized to be more token efficient.

Summary of changes:

I really really hope that this can also be useful for the many other users out there that use these bottom-of-the-barrel cheap but available Vega20 GPUs. At the same time though, i feel uneasy about releasing something that i barely comprehend so i am looking for feedback if i made any logical errors in this or messed up in some other way.

All these changes hugely improved prompt processing speeds for me, especially with quantized or half-quantized KV cache.

Since this happened over the past few weeks i have had Qwen 3.8 sync the upstream repo to incroporate most of the latest changes. I promise to sync it every now and again.

Maybe someday some of this could be useful enough to be incorporated into upstream llama.cpp but i think it’s still far from it. Let me know what you think.

Anyway, heres the repo:

I also had the AI write up some stuff for people like me to at least coarsely understand what the changes are about:

I would be happy if whoever has similar cards in use could try it out and report back if it was of any use to them.

This was built and tested on top of rocm 6.1 which has repeatedly proven to me as the fastest rocm version to use with these cards.

4 Likes

Nice you have two RVII’s?

For reference:

Summary

llama.cpp gfx906 Forks

Repository Description Key Features Status
iacopPBK/llama.cpp-gfx906 Primary dedicated MI50/MI60 fork Custom gfx906 kernels, optimized flash attention, MMVQ, fused operations, Q8/MXFP4 paths Active; builds from 2026 era with ROCm 7.1.x
arte-fact/llamacpp-gfx-906-turbo iacopPBK-based fork with additional tuning TurboQuant KV-cache compression, HIP fixes, gfx906-specific tuning Active
ollo12-prog/llama.cpp-gfx906-solve-tri-fix Qwen3.5/MI50 fork addressing rocBLAS issues Resolves SOLVE_TRI/rocBLAS failure on gfx906 Archived April 5, 2026
Mainline llama.cpp Upstream gfx906 tuning efforts MMQ tuning specifically for gfx906; performance varies by quantization (Q8↑, Q4_K↓) Ongoing

Legacy References

  • mixa3607/ML-gfx906 — now maintained as part of the larger gfx906 software stack; most recent update 2026
  • Lemonade’s gfx90X-dcgpu — closest “gfx90X” label for backend/ROCm device mapping (includes gfx908, gfx90a)

vLLM gfx906 Variants

Repository Type Status Last Activity
ai-infos/vllm-gfx906-mobydick Newer vLLM fork (recent upstream) Active — current development line Ongoing
mixa3607/vllm-gfx906 Prebuilt Docker image/build artifact Available as part of gfx906 stack Active stack
mixa3607/ML-gfx906/vllm vLLM source/build integration in stack Active stack; vLLM marked paused Paused
nlzy/vllm-gfx906 Original dedicated gfx906 fork Archived February 20, 2026
ai-infos/vllm-gfx906-deepseek DeepSeek-focused fork of nlzy Archived February 22, 2026
PowerfulGhost/vllm-mi50 MI50-targeted fork Archived; redirects to newer fork March 2, 2026
nalanzeyu/vllm-gfx906 Docker image (older nlzy base) Listed; legacy lineage Not maintained

I too have finetuned a llama.cpp or two, but for one card, we mashed flashattention into a variant in jan/feb then one of these beat me to it, ended up just using it.

I actively use knguyen298/llama-swap-gfx906 - Docker Image

They have built router features into newer llama.cpp, but my harness has the extra model field in API calls and that is easier for me at least. I use ai-infos vLLM for that container.

I’ll spin up this variant later on, the RVII is crunching tokens right now

1 Like

I, in fact, saw your post on Reddit and then noticed the URL and came here. I will surely try it out and report back with numbers. I do have several MI50 that currently run the default llama.cpp.

1 Like

Awesome, thank you!

Yeah i bought two RVIIs. They were more expensive than MI50 but i didn’t want blower style cards on my desk.

Our own production config has moved from using MTP and adaptive MTP to using Dflash 2 in order to keep the speed-gain but minimize the vram footprint at runtime.

This along with some bugs has shown to invalidate many of our optimizations and actually made us slower than mainline. Among others, we regressed badly in TG speeds.

Our previous opts were either workload specific to our old MTP based workloads, no longer relevant or interacted in a bad way with the new configuation.

Thats when we shifted to isolating and measuring all the optimizations we had added since.

Well, we (that is Qwen 3.8 27B, GLM 5.3 and I) went back to work the past days and have something new to show for it.

We ran controlled A/B lanes and actually isolated regressions.

It turned out our native gfx906 FATTN path which was supposed to be optimized was actually performing poorly for our new production workload.

With dflash2 the actual decode/verify shapes being sent through FATTN are different from what some of the earlier optimization work had targeted. In our initially profiled workload we were seeing extremely small Q shapes. Specifically the hot path was effectively Q=5.

It turned out the existing native FATTN geometry was poorly suited to Q=5. This lead the kernel to operate on a much larger tile than the useful query size.

That meant a lot of the work performed was effectively padding/wasted. It was estimated that around 85% of the QK vector work was wasted in that situation.

The supposedly less elegant conversion path was substantially faster than native FATTN. Our numbers were roughly:

native: 9.2t/s TG

convert: 12.6t/s TG

Instead of immediately hardcoding convert everywhere, we implemented selectable native/convert paths for validation. Then we added an AUTO selector: It is not just choosing which kernel to launch but rather sizing/allocation and dispatch all agree on the same choice.

Our first such implementation had a bypass bug, claiming convert but actually running native. We caught it in our A/B tests and fixed it eventually.

Afterwards we investigated why the native path hated small Q. This led us to introduce a special narrow geometry for small-Q/ Q <= 16 instead of changing the entire kernel. Larger Q workloads retain existing geometry. This targets the actual hot shape instead of broadly retuning FATTN.

The small-Q change took native from about 9.2 to 12.3t/s which is a roughly 34% improvement. Fill performance remained intact.

Why AUTO is useful:

- Convert requires additional temporary VRAM.

- Native is much cheaper in memory.

- At lower ctx or with enough vram, convert may be the better perf choice.

- At very high ctx or under vram pressure, native can become the better overall choice because the conversion buffer may no longer be practical.

-> hence the context/mem aware path selection.

We also took a look at all existing gfx906 related llama cpp forks and optimizations we could find, evaluated them all very carefully but found none of them to be useful to our actual workloads.

Now for the numbers:

- Prompt processing for the first 16384 token batch:

fork: 379.2t/s

mainline: 332.3t/s

+14.1% for the fork

- 120k long ctx fill:

fork: 252.6 t/s

mainline: 231.1t/s

+9.3% for the fork

- deep TG at 120k ctx, 512 tokens, temp 0:

fork: 13.6t/s

mainline: 13.5t/s

parity

Dflash acceptance stays the same.

The comparison used the same configs for both runs:

Qwen 3.8 27B Q6_K + Q4_K_M DFlash2 drafter, n_max=4, mmproj, f16 K / q8_0 V, tensor split 35/20/45, ~200k context

The fork has been synced with mainline llama.cpp on 31st of August.

Other Vega 20 users: what speeds do you get with what setup? I’m now at around 400tps for PP16385 and 10-15tps TG

Time for another update! We have been busy and managed to improve the gains substantially (mostly from exploring existing llama cpp PRs and adopting relevant things).

Among other things the README.md was also appended to provide a better overall picture of what’s in the fork, why and from whom.

metric upstream t/s fork t/s gain
prefill PP16384 332.5 ~410 +23%
120k deep fill 231.4 ~264 +14%
TG @120k depth 13.6 ~15.1 +11% (parity pre-mirror)
context cannot fit 250k on 40 GB tight-fit machinery
outputs - - bit-identical (sha + token-for-token)
1 Like

Does this work for you with Qwen3.8 Flash Next? I got it working with Qwen 3.6 35b a3b without issue. But i’m struggling with Qwen 3.8 Flash Next. (Got a nice speed boost on Qwen 3.6 35b from 50tps to 70-tps! thx!)

Sorry i have not tried it as my GPUs only have a total of 40gb of vram. What error do you encounter exactly?

rpcribari@2990MI50:~/builds-llama.cpp/milpster-gfx906-llama-cpp/build/bin$ ./llama-server -m ~/models/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -t 16 -ngl 999 --split-mode layer -ctk q8_0 -ctv q8_0 -c 8192 --host 0.0.0.0 --port 8080
0.01.880.020 I cmn common_param: common_params_print_info: verbosity = 3 (adjust with the -lv N CLI arg)
0.01.880.518 W srv llama_server: -----------------
0.01.880.522 W srv llama_server: CORS is set to allow all origins (‘*’) and no API key is set
0.01.880.523 W srv llama_server: this can be a security risk (cross-origin attacks)
0.01.880.524 W srv llama_server: more info: server: add --cors-* options by ngxson · Pull Request #25655 · ggml-org/llama.cpp · GitHub
0.01.880.524 W srv llama_server: -----------------
0.01.881.885 I srv load_model: loading model ‘/home/rpcribari/models/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf’
0.28.959.850 I cmn init: llama threadpool init, n_threads = 16
/home/rpcribari/gfx906-llama-cpp/ggml/src/ggml-cuda/ggml-cuda.cu:108: ROCm error
0.28.977.966 E ROCm error: invalid device function
0.28.977.973 E current device: 0, in function ggml_cuda_kernel_launch at /home/rpcribari/gfx906-llama-cpp/ggml/src/ggml-cuda/common.cuh:1723
0.28.977.974 E hipGetLastError()

New LWP 137736
New LWP 137737
New LWP 137741
New LWP 137743
New LWP 137744
New LWP 137745
New LWP 137746
New LWP 137747
New LWP 137748
New LWP 137749
New LWP 137750
New LWP 137751
New LWP 137752
New LWP 137753
New LWP 137754
New LWP 137755
New LWP 137756
New LWP 137757
New LWP 137758
New LWP 137759
New LWP 137760
New LWP 137761
New LWP 137762
New LWP 137763
New LWP 137764
New LWP 137765
New LWP 137766
New LWP 137767
New LWP 137768
New LWP 137769
New LWP 137770
New LWP 137771
New LWP 137772
New LWP 137773
New LWP 137774
New LWP 137775
New LWP 137776
New LWP 137777
New LWP 137778
New LWP 137779
New LWP 137780
New LWP 137781
New LWP 137782
New LWP 137783
New LWP 137784
New LWP 137785
New LWP 137786
New LWP 137787
New LWP 137788
New LWP 137789
New LWP 137790
New LWP 137791
New LWP 137792
New LWP 137793
New LWP 137794
New LWP 137795
New LWP 137796
New LWP 137797
New LWP 137798
New LWP 137799
New LWP 137800
New LWP 137801
New LWP 137802
New LWP 137803
New LWP 137804
New LWP 137805
New LWP 137806
New LWP 137808
Thread debugging using libthread_db enabled

Using host libthread_db library “/lib/x86_64-linux-gnu/libthread_db.so.1”.
0x00007731214ea3ef in __GI___wait4 (pid=137851, stat_loc=0x0, options=0, usage=0x0) at ../sysdeps/unix/sysv/linux/wait4.c:30
30 ../sysdeps/unix/sysv/linux/wait4.c: No such file or directory.*
#0 0x00007731214ea3ef in __GI___wait4 (pid=137851, stat_loc=0x0, options=0, usage=0x0) at ../sysdeps/unix/sysv/linux/wait4.c:30
30 in ../sysdeps/unix/sysv/linux/wait4.c

#1 0x0000773122259e3b in ggml_print_backtrace () from libggml-base.so.0
#2 0x0000773122259fd2 in ggml_abort () from libggml-base.so.0
#3 0x000077311d0c8e42 in ggml_cuda_error(char const, char const*, char const*, int, char const*) () from libggml-hip.so.0
#4 0x000077311cfe5468 in void ggml_cuda_op_bin_bcast<bin_bcast_cuda<&(op_repeat(float, float)), 0> >(ggml_tensor const*, ggml_tensor const*, ggml_tensor*, void const*, void const*, void*, ihipStream_t*) () from libggml-hip.so.0
#5 0x000077311d0d4fc7 in ggml_cuda_graph_evaluate_and_capture(ggml_backend_cuda_context*, ggml_cgraph*, bool, bool, void const*) () from libggml-hip.so.0
#6 0x000077311d0cf0ac in ggml_backend_cuda_graph_compute(ggml_backend*, ggml_cgraph*) () from libggml-hip.so.0
#7 0x0000773122279c27 in ggml_backend_sched_graph_compute_async () from libggml-base.so.0
#8 0x0000773120afea01 in llama_context::graph_compute(ggml_cgraph*, bool) () from libllama.so.0
#9 0x0000773120b01a6a in llama_context::process_ubatch(llama_ubatch const&, llm_graph_type, llama_memory_context_i*, ggml_status&) () from libllama.so.0
#10 0x0000773120b09bff in llama_context::decode(llama_batch const&) () from libllama.so.0
#11 0x0000773120b0af10 in llama_decode () from libllama.so.0
#12 0x000077312100fbd0 in common_init_from_params(common_params&, bool) () from libllama-common.so.0
#13 0x0000773121d7476b in server_context_impl::load_model(common_params&) () from libllama-server-impl.so
#14 0x0000773121cc5c29 in llama_server(common_params&, int, char**) () from libllama-server-impl.so
#15 0x0000773121cc755f in llama_server(int, char**) () from libllama-server-impl.so
#16 0x0000773121429d90 in __libc_start_call_main (main=main@entry=0x59759fe23260 , argc=argc@entry=19, argv=argv@entry=0x7ffef35e7548) at ../sysdeps/nptl/libc_start_call_main.h:58
58 ../sysdeps/nptl/libc_start_call_main.h: No such file or directory.
#17 0x0000773121429e40 in __libc_start_main_impl (main=0x59759fe23260 , argc=19, argv=0x7ffef35e7548, init=, fini=, rtld_fini=, stack_end=0x7ffef35e7538) at ../csu/libc-start.c:392
392 ../csu/libc-start.c: No such file or directory.
#18 0x000059759fe23295 in _start ()

Inferior 1 (process 137734) detached

Aborted (core dumped)

1 Like

I’m getting a nice speed boost on Qwen 3.8 27b too. Went from 18tps using Vulkan to 30tps using ROCm.

1 Like

Turns out the problem is my 4th MI50 is running the Apple Vega Pro II firmware. I needed it for displayport out. But it turns off ECC apparently and ROCm wasn’t happy with 3 using ECC and 1 not. I disabled it and the model loads without issue.

1 Like