TL;DR: This fork promises up to to double prompt processing speeds for deep infill on gfx906 based setups.
The title pretty much says it all. Since my setup is based on a wild arrangement of two Radeon VII and a laptop with an internal 8gb rtx3080 maxq, i have since worked with GLM to optimize said setup to try and get the most out of it.
Disclaimer: I did not code this myself but only steered GLM 5.3 towards the result but i have been using it in production ever since. I do not like talking about things i barely understand and this is really at the edge of what is comprehensible for me as a lowly IoT engineer.
Setup: Linux, Qwen 3.8 27B at Q6_K_XL, llama.cpp and opencode with “oh-my-openagent” but customized to be more token efficient.
Summary of changes:
-
First we incorporated another patch that added pipeline parallelism and cleaned up some of the mtp memory allocations for more vram efficiency. All credit for that to this guy: Index of /lm/bug
-
From this dude’s repo we also imported a skill that will help your own agent to properly balance and optimize your model with your llama cpp backend.
-
Then GLM came up with a new cost-based split mode which can sometimes help if you also have a weird setup like mine, with two rocm cards and a vulkan based card in the middle. Sometimes it helps, sometimes it doesn’t - you need to bench it yourself: llama : add LLAMA_SPLIT_MODE_COST for hybrid models · milpster/gfx906-llama-cpp@604fca5 · GitHub
-
Next up GLM created a customized tile config for Vega 20: mmq : add vega20-specific tile config · milpster/gfx906-llama-cpp@468c164 · GitHub
-
Then we went on to quantize our KV cache which required another performance boost, thus GLM created a custom q8_0-native tile kernel for GCN HIP: fattn : q8_0-native tile kernel for GCN HIP · milpster/gfx906-llama-cpp@d14628d · GitHub
-
From there i went to only quantizing the V cache to q8 where GLM created a mixed F16-K q8_0-V tile kernel for GCN HIP: fattn : mixed F16-K q8_0-V tile kernel for GCN HIP · milpster/gfx906-llama-cpp@eaca6d4 · GitHub
I really really hope that this can also be useful for the many other users out there that use these bottom-of-the-barrel cheap but available Vega20 GPUs. At the same time though, i feel uneasy about releasing something that i barely comprehend so i am looking for feedback if i made any logical errors in this or messed up in some other way.
All these changes hugely improved prompt processing speeds for me, especially with quantized or half-quantized KV cache.
Since this happened over the past few weeks i have had Qwen 3.8 sync the upstream repo to incroporate most of the latest changes. I promise to sync it every now and again.
Maybe someday some of this could be useful enough to be incorporated into upstream llama.cpp but i think it’s still far from it. Let me know what you think.
Anyway, heres the repo:
I also had the AI write up some stuff for people like me to at least coarsely understand what the changes are about:
I would be happy if whoever has similar cards in use could try it out and report back if it was of any use to them.
This was built and tested on top of rocm 6.1 which has repeatedly proven to me as the fastest rocm version to use with these cards.
