Converting a data center GPU for desktop use

Have you tried using the tensor split mode? This is the mode that actually spreads compute across GPUs, but also requires way more data passing across them.

Took a side trip to upgrade the CPU (2x Xeon at $70 for the pair). This went sideways. Not sure if it was a bent pin in the CPU socket, or a bit of gunk, but took a couple of tries. This is pretty much the limit for this old box - maxxed out for CPU, RAM, have practically infinite storage and on 10GbE.

(Clean rebuild of Linux with 32-threads takes 16-odd minutes. Not bad for near 9-hours CPU-time.)

Had the GPU overheat once, as seems the fan-control script came up before the card’s sysfs interface was ready. Updated the fan-control script to adapt.

Learned a bit more…

Seems Vulcan is improving steadily, and current Vulcan and ROCm are pretty much equivalent (but not identical). Also good news as AMD re-added support for gfx900 in ROCm 7.13.

Also should note there is a lot of churn in this space. Seems there is continuous improvement all through the stack, so you really want to be on bleeding-edge/current versions. Building from someone’s last-night Git-commit is not crazy.

Updated my list of models to test - hello “Gemma_4_QAT”.

So … at present, have a working setup. Will learn to make best use of local LLMs for development - which was the original purpose.

Still have a “to-do” list, from what I learned:

  1. Want to install ROCm 7.13 (or later).
    Get performance or feature gain (that might matter).
  2. Want to re-install Linux (again).
    Not locked to ROCm 5.7 means not locked to Ubuntu 22.
    Would slightly prefer Debian-latest (or ArchLinux?).
  3. Build custom Linux kernel.
    Some gain for CPU-specific build, and disabled mitigations.
  4. Update the MI25 BIOS to 260 (or 300) watt limit.
    Current BIOS has 220 watt limit. Cooling seems adequate.

Might be a while. Also have two Voron Trident kits (3D-printers) to build, update to R2, and add Bondtech INDX. (Another topic altogether.)

1 Like

Oh that’s good to hear. I have one machine with an RX580 and a W6800 in it and the newer ROCm would just crash on the unsupported card (gfx803) instead of ignoring it and using the W6800 (I only used the RX580 for extra video outputs).

I got past it by exposing the supported AMD card to my ComfyUI Docker container:

[Issue]: rocm initialization fails when incompatible GPU present (W6800 and RX580 in same system) · Issue #6883 · ROCm/rocm-systems · GitHub

I would recommend using the wx9100 vbios + edited pptable(using the upp tool) over the mi25 vbios with the higher PPT limit, the performance is a bit better in general.

An important note about loading a custom pptable, is that it will disable avfs even if you only edit the power limit, thus it is important to renable AVFS via the ppfeature mask , otherwise you will get excessive voltage and higher than normal power draw, I have sucessfully taken my mi25 above 400w using the wx9100 vbios using some powerful cooling on a cold day.( don’t recommend doing that unless you are certain you have a variant with the full vrm, rather than the cut down one)

Is there a good way to tell the boards apart without taking apart the heatsinks? Wasn’t sure if the boards would show differently in GPU-z or something like that.

Good hardwork.

My understanding is that if the board has 8+8pin pcie power connectors , its more likely to be a full vrm card, vs if its equiped with a 6+8pin, though generally you will know if its the weak vrm version, because using a 200w+ power limit will cause the vrm to get really hot very quickly, which visible in gpu-z sensors, or just by touching it no matter how much air flow you have.

My MI25 has dual power connectors. Should also note that I simply do not have any trouble with VRM temperatures, as they are very little different from hotspot temperatures.

As a guess, perhaps this is due to cooling the card as-designed (using a GPU fan). Whatever the reason, the GPU fan alone is more than able to cool the VRMs.

2 Likes

Got a bit further along. To be clear, the hardware part of the equation seems to be working well. The MI25 temperatures are well-controlled with the MI25 220W BIOS and the fan-control service. So the card seems well settled.

Also to be clear, I spent the last several years working in a cave (for a DoD contractor), so am very much catching up. There is an impressive amount of work going on around LLMs. I expect there are a lot of folk in the same space. But also I am of the habit of jumping into a new space, asking a lot of dumb questions, then finding the bleeding edge in a few months.

There is something about llama.cpp, contexts, and loading more than one model. In operating systems, there is long-settled technique around hotspots and LRU, and intelligent caching. Best I can tell, in the LLM case, this is still very primitive.

Seems llama.cpp is fairly stupid when you load more than one model, and one is used more than the other.

Of the many choices, loaded the llama.vscode extension into VSCode.

Well … guess what … llama.vscode offers you the choice to use different LLMs for “chat” and “tools” and … yeh, other things. What I would choose and why is unclear.

I have the (old) “beast” server, with dual (8-core each) Xeons, 256GB RAM, and absurd amounts of storage. Have a Zen 3 desktop (12-cores) with 128GB RAM and also absurd storage. The “beast” also has the AMD MI25, an (incidental) NVidia 1050 with 2GB VRAM. This desktop has an AMD RX 5500 with 8GB of VRAM.

So I could load different models into different spaces.

Tokens per second I am seeing with this old datacenter GPU seem acceptable.

MIght have learned a bit, so updated my list of models for test:

model_add "DeepSeek-R1-Distill-Qwen-1.5B"   ":UD-Q4_K_XL"   "unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF"
model_add "DeepSeek-R1-Distill-Qwen-14B"    ":Q4_K_M"       "unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF"

# Gemma 4 non-QAT GGUF crashes on MI25 (segfault). The QAT version below may work.
model_add "Gemma-4-E2B-QAT"                 ":UD-Q4_K_XL"   "unsloth/gemma-4-E2B-it-qat-GGUF"                 "--jinja"
model_add "Gemma-4-E4B-QAT"                 ":UD-Q4_K_XL"   "unsloth/gemma-4-E4B-it-qat-GGUF"                 "--jinja"
model_add "Gemma-4-12B-QAT"                 ":UD-Q4_K_XL"   "unsloth/gemma-4-12B-it-qat-GGUF"                 "--jinja"
model_add "Gemma-4-26B-A4B-QAT"             ":UD-Q4_K_XL"   "unsloth/gemma-4-26B-A4B-it-qat-GGUF"             "--jinja"
model_add "Gemma-4-31B-QAT"                 ":UD-Q4_K_XL"   "unsloth/gemma-4-31B-it-qat-GGUF"                 "--jinja"

model_add "GPT-OSS-20B"                     ":Q4_K_M"       "unsloth/gpt-oss-20b-GGUF"   

model_add "LLama-3.2-1B"                    ":Q4_K_M"       "unsloth/Llama-3.2-1B-Instruct-GGUF"              
model_add "LLama-3.2-3B"                    ":Q4_K_M"       "unsloth/Llama-3.2-3B-Instruct-GGUF"              

# model_add "Microsoft-Phi-4"                 ":Q4_K_M"       "unsloth/Phi-4-mini-instruct-GGUF"                  

model_add "Devstral-Small-2-24B"            ":Q4_K_M"       "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF"       
model_add "Mistral-Small-3.2-24B"           ":Q4_K_S"       "unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"       

model_add "Qwen-3.5-2B"                     ":Q4_K_M"       "unsloth/Qwen3.5-2B-GGUF"                           
model_add "Qwen-3.5-4B"                     ":Q4_K_M"       "unsloth/Qwen3.5-4B-GGUF"                           
model_add "Qwen-3.5-9B"                     ":Q4_K_M"       "unsloth/Qwen3.5-9B-GGUF"                           
model_add "Qwen-3.5-27B"                    ":Q4_K_S"       "unsloth/Qwen3.5-27B-GGUF"                           

In theory, models that do not fit into GPU VRAM could be partitioned into CPU RAM. With a proper caching hierarchy, this could work well. Best I can tell, this does not work at all well. Models that end up partly in CPU memory suffer badly.

Also to be clear, I am looking for the “knee” in the curve. There is usually a point where you end up spending a lot for small improvement. In my testing, the “small” models seem to perform surprisingly well. As a first rough guess, the “Q4_K_M” models that fit in 16GB VRAM seem to be largely good enough.

So … might end up loading small models into the local GPU, and larger models into the datacenter GPU … maybe?

[Edit: Was missing Gemma-4 benchmark results.]

Download - DeepSeek-R1-Distill-Qwen-1.5B

Model Family Model Name
DeepSeek-R1-Distill-Qwen-1.5B unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.08.341.702 I common_perf_print:    sampling time =     574.71 ms
0.08.341.705 I common_perf_print:    samplers time =     195.84 ms /   821 tokens
0.08.341.713 I common_perf_print:        load time =     963.31 ms
0.08.341.717 I common_perf_print: prompt eval time =      88.29 ms /    35 tokens (    2.52 ms per token,   396.43 tokens per second)
0.08.341.721 I common_perf_print:        eval time =    5668.42 ms /   785 runs   (    7.22 ms per token,   138.49 tokens per second)
0.08.341.723 I common_perf_print:       total time =    6364.00 ms /   820 tokens
0.08.341.726 I common_perf_print: unaccounted time =      32.58 ms /   0.5 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m8.457s 0m4.732s 0m1.022s

Download - DeepSeek-R1-Distill-Qwen-14B

Model Family Model Name
DeepSeek-R1-Distill-Qwen-14B unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
+ llama-completion -hf unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.41.403.661 I common_perf_print:    sampling time =     883.69 ms
0.41.403.665 I common_perf_print:    samplers time =     295.72 ms /   997 tokens
0.41.403.676 I common_perf_print:        load time =    6132.24 ms
0.41.403.683 I common_perf_print: prompt eval time =     409.00 ms /    35 tokens (   11.69 ms per token,    85.57 tokens per second)
0.41.403.689 I common_perf_print:        eval time =   32508.98 ms /   961 runs   (   33.83 ms per token,    29.56 tokens per second)
0.41.403.692 I common_perf_print:       total time =   33853.14 ms /   996 tokens
0.41.403.695 I common_perf_print: unaccounted time =      51.46 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m41.560s 0m17.038s 0m4.310s

Download - Gemma-4-E2B-QAT

Model Family Model Name
Gemma-4-E2B-QAT unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja

0.17.396.149 I common_perf_print:    sampling time =     740.27 ms
0.17.396.152 I common_perf_print:    samplers time =     254.77 ms /  1255 tokens
0.17.396.157 I common_perf_print:        load time =    1324.85 ms
0.17.396.159 I common_perf_print: prompt eval time =     128.93 ms /    47 tokens (    2.74 ms per token,   364.54 tokens per second)
0.17.396.162 I common_perf_print:        eval time =   13295.65 ms /  1207 runs   (   11.02 ms per token,    90.78 tokens per second)
0.17.396.163 I common_perf_print:       total time =   14197.52 ms /  1254 tokens
0.17.396.164 I common_perf_print: unaccounted time =      32.67 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m17.629s 0m8.842s 0m1.626s

Download - Gemma-4-E4B-QAT

Model Family Model Name
Gemma-4-E4B-QAT unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja

0.27.116.205 I common_perf_print:    sampling time =    1352.30 ms
0.27.116.211 I common_perf_print:    samplers time =     467.16 ms /  1297 tokens
0.27.116.218 I common_perf_print:        load time =    2182.16 ms
0.27.116.221 I common_perf_print: prompt eval time =     184.00 ms /    47 tokens (    3.91 ms per token,   255.43 tokens per second)
0.27.116.228 I common_perf_print:        eval time =   21471.61 ms /  1249 runs   (   17.19 ms per token,    58.17 tokens per second)
0.27.116.236 I common_perf_print:       total time =   23058.86 ms /  1296 tokens
0.27.116.239 I common_perf_print: unaccounted time =      50.96 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m27.369s 0m14.144s 0m2.741s

Download - Gemma-4-12B-QAT

Model Family Model Name
Gemma-4-12B-QAT unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja

0.47.226.558 I common_perf_print:    sampling time =    1769.61 ms
0.47.226.563 I common_perf_print:    samplers time =     610.27 ms /  1156 tokens
0.47.226.573 I common_perf_print:        load time =    4824.19 ms
0.47.226.579 I common_perf_print: prompt eval time =     297.95 ms /    47 tokens (    6.34 ms per token,   157.75 tokens per second)
0.47.226.584 I common_perf_print:        eval time =   38375.83 ms /  1108 runs   (   34.64 ms per token,    28.87 tokens per second)
0.47.226.587 I common_perf_print:       total time =   40509.26 ms /  1155 tokens
0.47.226.590 I common_perf_print: unaccounted time =      65.87 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m47.499s 0m20.459s 0m4.336s

Download - Gemma-4-26B-A4B-QAT

Model Family Model Name
Gemma-4-26B-A4B-QAT unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja

0.39.427.510 I common_perf_print:    sampling time =    1098.37 ms
0.39.427.512 I common_perf_print:    samplers time =     378.43 ms /  1547 tokens
0.39.427.516 I common_perf_print:        load time =    9611.67 ms
0.39.427.519 I common_perf_print: prompt eval time =     328.79 ms /    47 tokens (    7.00 ms per token,   142.95 tokens per second)
0.39.427.521 I common_perf_print:        eval time =   25684.05 ms /  1499 runs   (   17.13 ms per token,    58.36 tokens per second)
0.39.427.522 I common_perf_print:       total time =   27156.52 ms /  1546 tokens
0.39.427.524 I common_perf_print: unaccounted time =      45.32 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m39.661s 0m19.980s 0m5.176s

Download - Gemma-4-31B-QAT

Model Family Model Name
Gemma-4-31B-QAT unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
+ llama-completion -hf unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja

2.57.607.770 I common_perf_print:    sampling time =     591.32 ms
2.57.607.774 I common_perf_print:    samplers time =     157.10 ms /   990 tokens
2.57.607.781 I common_perf_print:        load time =    9971.08 ms
2.57.607.783 I common_perf_print: prompt eval time =    2204.15 ms /    47 tokens (   46.90 ms per token,    21.32 tokens per second)
2.57.607.785 I common_perf_print:        eval time =  158801.82 ms /   942 runs   (  168.58 ms per token,     5.93 tokens per second)
2.57.607.787 I common_perf_print:       total time =  161672.46 ms /   989 tokens
2.57.607.789 I common_perf_print: unaccounted time =      75.17 ms /   0.0 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
2m57.861s 0m19.980s 0m9.568s

Download - GPT-OSS-20B

Model Family Model Name
GPT-OSS-20B unsloth/gpt-oss-20b-GGUF:Q4_K_M
+ llama-completion -hf unsloth/gpt-oss-20b-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.20.527.213 I common_perf_print:    sampling time =     579.61 ms
0.20.527.215 I common_perf_print:    samplers time =     192.78 ms /   799 tokens
0.20.527.222 I common_perf_print:        load time =    7784.85 ms
0.20.527.227 I common_perf_print: prompt eval time =     282.51 ms /    38 tokens (    7.43 ms per token,   134.51 tokens per second)
0.20.527.231 I common_perf_print:        eval time =   10211.50 ms /   760 runs   (   13.44 ms per token,    74.43 tokens per second)
0.20.527.232 I common_perf_print:       total time =   11103.83 ms /   798 tokens
0.20.527.237 I common_perf_print: unaccounted time =      30.21 ms /   0.3 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m20.762s 0m10.969s 0m3.722s

Download - LLama-3.2-1B

Model Family Model Name
LLama-3.2-1B unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.04.781.760 I common_perf_print:    sampling time =     156.94 ms
0.04.781.765 I common_perf_print:    samplers time =      52.71 ms /   560 tokens
0.04.781.772 I common_perf_print:        load time =     859.66 ms
0.04.781.775 I common_perf_print: prompt eval time =      51.36 ms /    42 tokens (    1.22 ms per token,   817.76 tokens per second)
0.04.781.780 I common_perf_print:        eval time =    2506.05 ms /   517 runs   (    4.85 ms per token,   206.30 tokens per second)
0.04.781.781 I common_perf_print:       total time =    2726.30 ms /   559 tokens
0.04.781.791 I common_perf_print: unaccounted time =      11.94 ms /   0.4 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m4.925s 0m2.565s 0m0.679s

Download - LLama-3.2-3B

Model Family Model Name
LLama-3.2-3B unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.10.207.955 I common_perf_print:    sampling time =     282.44 ms
0.10.207.959 I common_perf_print:    samplers time =     117.30 ms /   680 tokens
0.10.207.966 I common_perf_print:        load time =    1623.02 ms
0.10.207.971 I common_perf_print: prompt eval time =      99.35 ms /    42 tokens (    2.37 ms per token,   422.73 tokens per second)
0.10.207.975 I common_perf_print:        eval time =    6484.59 ms /   637 runs   (   10.18 ms per token,    98.23 tokens per second)
0.10.207.977 I common_perf_print:       total time =    6891.03 ms /   679 tokens
0.10.207.979 I common_perf_print: unaccounted time =      24.65 ms /   0.4 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m10.367s 0m5.117s 0m1.139s

Download - Devstral-Small-2-24B

Model Family Model Name
Devstral-Small-2-24B unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.57.924.327 I common_perf_print:    sampling time =     657.25 ms
0.57.924.331 I common_perf_print:    samplers time =     236.95 ms /   838 tokens
0.57.924.340 I common_perf_print:        load time =    9599.16 ms
0.57.924.345 I common_perf_print: prompt eval time =     555.55 ms /    35 tokens (   15.87 ms per token,    63.00 tokens per second)
0.57.924.350 I common_perf_print:        eval time =   45340.26 ms /   802 runs   (   56.53 ms per token,    17.69 tokens per second)
0.57.924.354 I common_perf_print:       total time =   46596.88 ms /   837 tokens
0.57.924.357 I common_perf_print: unaccounted time =      43.83 ms /   0.1 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m58.111s 0m22.715s 0m5.470s

Download - Mistral-Small-3.2-24B

Model Family Model Name
Mistral-Small-3.2-24B unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
+ llama-completion -hf unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.29.461.436 I common_perf_print:    sampling time =     289.50 ms
0.29.461.440 I common_perf_print:    samplers time =     104.16 ms /   389 tokens
0.29.461.449 I common_perf_print:        load time =    9206.67 ms
0.29.461.455 I common_perf_print: prompt eval time =     560.87 ms /    35 tokens (   16.02 ms per token,    62.40 tokens per second)
0.29.461.461 I common_perf_print:        eval time =   17639.85 ms /   353 runs   (   49.97 ms per token,    20.01 tokens per second)
0.29.461.463 I common_perf_print:       total time =   18510.58 ms /   388 tokens
0.29.461.467 I common_perf_print: unaccounted time =      20.36 ms /   0.1 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m29.650s 0m11.903s 0m4.503s

Download - Qwen-3.5-2B

Model Family Model Name
Qwen-3.5-2B unsloth/Qwen3.5-2B-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Qwen3.5-2B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.22.867.546 I common_perf_print:    sampling time =    1308.94 ms
0.22.867.549 I common_perf_print:    samplers time =     432.88 ms /  2178 tokens
0.22.867.553 I common_perf_print:        load time =    1358.44 ms
0.22.867.555 I common_perf_print: prompt eval time =      88.98 ms /    40 tokens (    2.22 ms per token,   449.54 tokens per second)
0.22.867.556 I common_perf_print:        eval time =   18511.10 ms /  2137 runs   (    8.66 ms per token,   115.44 tokens per second)
0.22.867.557 I common_perf_print:       total time =   19962.84 ms /  2177 tokens
0.22.867.558 I common_perf_print: unaccounted time =      53.82 ms /   0.3 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m23.038s 0m11.927s 0m2.457s

Download - Qwen-3.5-4B

Model Family Model Name
Qwen-3.5-4B unsloth/Qwen3.5-4B-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.16.010.620 I common_perf_print:    sampling time =     347.32 ms
0.16.010.621 I common_perf_print:    samplers time =     109.96 ms /   786 tokens
0.16.010.625 I common_perf_print:        load time =    2335.15 ms
0.16.010.626 I common_perf_print: prompt eval time =     140.85 ms /    40 tokens (    3.52 ms per token,   284.00 tokens per second)
0.16.010.628 I common_perf_print:        eval time =   11521.64 ms /   745 runs   (   15.47 ms per token,    64.66 tokens per second)
0.16.010.629 I common_perf_print:       total time =   12031.35 ms /   785 tokens
0.16.010.632 I common_perf_print: unaccounted time =      21.55 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m16.183s 0m7.228s 0m1.775s

Download - Qwen-3.5-9B

Model Family Model Name
Qwen-3.5-9B unsloth/Qwen3.5-9B-GGUF:Q4_K_M
+ llama-completion -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

0.57.681.435 I common_perf_print:    sampling time =    2695.21 ms
0.57.681.439 I common_perf_print:    samplers time =     886.45 ms /  2041 tokens
0.57.681.450 I common_perf_print:        load time =    3879.13 ms
0.57.681.456 I common_perf_print: prompt eval time =     201.42 ms /    40 tokens (    5.04 ms per token,   198.59 tokens per second)
0.57.681.462 I common_perf_print:        eval time =   49290.19 ms /  2000 runs   (   24.65 ms per token,    40.58 tokens per second)
0.57.681.464 I common_perf_print:       total time =   52281.45 ms /  2040 tokens
0.57.681.468 I common_perf_print: unaccounted time =      94.62 ms /   0.2 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
0m57.904s 0m26.875s 0m5.596s

Download - Qwen-3.5-27B

Model Family Model Name
Qwen-3.5-27B unsloth/Qwen3.5-27B-GGUF:Q4_K_S
+ llama-completion -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_S --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

2.17.918.517 I common_perf_print:    sampling time =    2112.98 ms
2.17.918.524 I common_perf_print:    samplers time =     558.55 ms /  1689 tokens
2.17.918.533 I common_perf_print:        load time =   10450.55 ms
2.17.918.538 I common_perf_print: prompt eval time =     878.25 ms /    40 tokens (   21.96 ms per token,    45.55 tokens per second)
2.17.918.542 I common_perf_print:        eval time =  119977.82 ms /  1648 runs   (   72.80 ms per token,    13.74 tokens per second)
2.17.918.545 I common_perf_print:       total time =  123125.12 ms /  1688 tokens
2.17.918.551 I common_perf_print: unaccounted time =     156.07 ms /   0.1 %      (total - sampling - prompt eval - eval) / (total)
real user sys time
2m18.154s 6m26.572s 0m10.839s

Benchmark - DeepSeek-R1-Distill-Qwen-1.5B

Model Family Model Name
DeepSeek-R1-Distill-Qwen-1.5B unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
qwen2 1.5B Q4_K - Medium 1.10 GiB 1.78 B Vulkan 99 pp512 2198.49 ± 4.97
qwen2 1.5B Q4_K - Medium 1.10 GiB 1.78 B Vulkan 99 pp2048 1903.26 ± 0.99
qwen2 1.5B Q4_K - Medium 1.10 GiB 1.78 B Vulkan 99 pp4096 1638.22 ± 1.45
qwen2 1.5B Q4_K - Medium 1.10 GiB 1.78 B Vulkan 99 tg128 146.36 ± 0.16
real user sys time
0m19.729s 0m4.565s 0m1.223s

Benchmark - DeepSeek-R1-Distill-Qwen-14B

Model Family Model Name
DeepSeek-R1-Distill-Qwen-14B unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
model size params backend ngl test t/s
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 pp512 250.11 ± 0.50
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 pp2048 227.32 ± 0.28
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 pp4096 203.33 ± 0.06
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B Vulkan 99 tg128 27.35 ± 0.19
real user sys time
2m25.928s 0m29.068s 0m7.502s

Benchmark - Gemma-4-E2B-QAT

Model Family Model Name
Gemma-4-E2B-QAT unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
gemma4 E2B Q4_0 2.43 GiB 4.63 B Vulkan 99 pp512 1462.95 ± 3.57
gemma4 E2B Q4_0 2.43 GiB 4.63 B Vulkan 99 pp2048 1308.52 ± 1.91
gemma4 E2B Q4_0 2.43 GiB 4.63 B Vulkan 99 pp4096 1177.14 ± 0.68
gemma4 E2B Q4_0 2.43 GiB 4.63 B Vulkan 99 tg128 94.76 ± 0.24
real user sys time
0m28.518s 0m9.309s 0m2.212s

Benchmark - Gemma-4-E4B-QAT

Model Family Model Name
Gemma-4-E4B-QAT unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
gemma4 E4B Q4_0 3.91 GiB 7.46 B Vulkan 99 pp512 790.98 ± 1.24
gemma4 E4B Q4_0 3.91 GiB 7.46 B Vulkan 99 pp2048 737.58 ± 0.46
gemma4 E4B Q4_0 3.91 GiB 7.46 B Vulkan 99 pp4096 692.21 ± 0.55
gemma4 E4B Q4_0 3.91 GiB 7.46 B Vulkan 99 tg128 58.50 ± 0.22
real user sys time
0m47.646s 0m13.713s 0m3.355s

Benchmark - Gemma-4-12B-QAT

Model Family Model Name
Gemma-4-12B-QAT unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
gemma4 ?B Q4_0 6.24 GiB 11.91 B Vulkan 99 pp512 314.82 ± 0.09
gemma4 ?B Q4_0 6.24 GiB 11.91 B Vulkan 99 pp2048 287.90 ± 0.44
gemma4 ?B Q4_0 6.24 GiB 11.91 B Vulkan 99 pp4096 269.61 ± 0.10
gemma4 ?B Q4_0 6.24 GiB 11.91 B Vulkan 99 tg128 27.70 ± 0.02
real user sys time
1m55.918s 0m27.792s 0m6.576s

Benchmark - Gemma-4-26B-A4B-QAT

Model Family Model Name
Gemma-4-26B-A4B-QAT unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B Vulkan 99 pp512 644.79 ± 4.22
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B Vulkan 99 pp2048 579.16 ± 0.35
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B Vulkan 99 pp4096 533.38 ± 1.71
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B Vulkan 99 tg128 61.67 ± 0.07
real user sys time
1m5.368s 0m19.110s 0m6.268s

Benchmark - Gemma-4-31B-QAT

Model Family Model Name
Gemma-4-31B-QAT unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
model size params backend ngl test t/s
gemma4 31B Q4_0 16.09 GiB 30.70 B Vulkan 99 pp512 13.00 ± 0.09
gemma4 31B Q4_0 16.09 GiB 30.70 B Vulkan 99 pp2048 12.77 ± 0.01
gemma4 31B Q4_0 16.09 GiB 30.70 B Vulkan 99 pp4096 12.59 ± 0.02
gemma4 31B Q4_0 16.09 GiB 30.70 B Vulkan 99 tg128 4.11 ± 0.00
real user sys time
1m5.368s 8m12.121s 2m0.144s

Benchmark - GPT-OSS-20B

Model Family Model Name
GPT-OSS-20B unsloth/gpt-oss-20b-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gpt-oss-20b-GGUF:Q4_K_M
model size params backend ngl test t/s
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp512 760.14 ± 10.00
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp2048 695.92 ± 0.90
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 pp4096 628.11 ± 0.98
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B Vulkan 99 tg128 70.54 ± 1.43
real user sys time
0m55.004s 0m13.544s 0m4.647s

Benchmark - LLama-3.2-1B

Model Family Model Name
LLama-3.2-1B unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
model size params backend ngl test t/s
llama 1B Q4_K - Medium 762.81 MiB 1.24 B Vulkan 99 pp512 3006.24 ± 8.26
llama 1B Q4_K - Medium 762.81 MiB 1.24 B Vulkan 99 pp2048 2388.89 ± 2.90
llama 1B Q4_K - Medium 762.81 MiB 1.24 B Vulkan 99 pp4096 1905.26 ± 0.99
llama 1B Q4_K - Medium 762.81 MiB 1.24 B Vulkan 99 tg128 210.30 ± 0.31
real user sys time
0m16.341s 0m3.748s 0m1.003s

Benchmark - LLama-3.2-3B

Model Family Model Name
LLama-3.2-3B unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
model size params backend ngl test t/s
llama 3B Q4_K - Medium 1.87 GiB 3.21 B Vulkan 99 pp512 1092.27 ± 1.55
llama 3B Q4_K - Medium 1.87 GiB 3.21 B Vulkan 99 pp2048 951.35 ± 0.63
llama 3B Q4_K - Medium 1.87 GiB 3.21 B Vulkan 99 pp4096 811.01 ± 0.32
llama 3B Q4_K - Medium 1.87 GiB 3.21 B Vulkan 99 tg128 94.83 ± 1.40
real user sys time
0m37.230s 0m7.795s 0m2.195s

Benchmark - Devstral-Small-2-24B

Model Family Model Name
Devstral-Small-2-24B unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
model size params backend ngl test t/s
mistral3 14B Q4_K - Medium 13.34 GiB 23.57 B Vulkan 99 pp512 149.21 ± 0.28
mistral3 14B Q4_K - Medium 13.34 GiB 23.57 B Vulkan 99 pp2048 143.30 ± 0.02
mistral3 14B Q4_K - Medium 13.34 GiB 23.57 B Vulkan 99 pp4096 136.49 ± 0.03
mistral3 14B Q4_K - Medium 13.34 GiB 23.57 B Vulkan 99 tg128 16.70 ± 0.04
real user sys time
3m44.573s 0m41.935s 0m11.212s

Benchmark - Mistral-Small-3.2-24B

Model Family Model Name
Mistral-Small-3.2-24B unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
model size params backend ngl test t/s
llama 13B Q4_K - Small 12.61 GiB 23.57 B Vulkan 99 pp512 144.61 ± 0.19
llama 13B Q4_K - Small 12.61 GiB 23.57 B Vulkan 99 pp2048 139.35 ± 0.10
llama 13B Q4_K - Small 12.61 GiB 23.57 B Vulkan 99 pp4096 132.68 ± 0.05
llama 13B Q4_K - Small 12.61 GiB 23.57 B Vulkan 99 tg128 18.03 ± 0.42
real user sys time
3m47.993s 0m40.593s 0m11.294s

Benchmark - Qwen-3.5-2B

Model Family Model Name
Qwen-3.5-2B unsloth/Qwen3.5-2B-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-2B-GGUF:Q4_K_M
model size params backend ngl test t/s
qwen35 2B Q4_K - Medium 1.18 GiB 1.88 B Vulkan 99 pp512 1784.73 ± 5.41
qwen35 2B Q4_K - Medium 1.18 GiB 1.88 B Vulkan 99 pp2048 1752.94 ± 2.29
qwen35 2B Q4_K - Medium 1.18 GiB 1.88 B Vulkan 99 pp4096 1707.30 ± 0.60
qwen35 2B Q4_K - Medium 1.18 GiB 1.88 B Vulkan 99 tg128 112.94 ± 1.68
real user sys time
0m21.157s 0m5.875s 0m1.513s

Benchmark - Qwen-3.5-4B

Model Family Model Name
Qwen-3.5-4B unsloth/Qwen3.5-4B-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M
model size params backend ngl test t/s
qwen35 4B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 99 pp512 775.41 ± 1.01
qwen35 4B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 99 pp2048 756.50 ± 0.58
qwen35 4B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 99 pp4096 730.54 ± 0.17
qwen35 4B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 99 tg128 61.02 ± 0.05
real user sys time
0m45.529s 0m10.969s 0m2.767s

Benchmark - Qwen-3.5-9B

Model Family Model Name
Qwen-3.5-9B unsloth/Qwen3.5-9B-GGUF:Q4_K_M
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M
model size params backend ngl test t/s
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp512 436.24 ± 0.32
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp2048 429.95 ± 0.28
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 pp4096 420.92 ± 0.27
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B Vulkan 99 tg128 39.25 ± 0.24
real user sys time
1m17.347s 0m18.605s 0m4.788s

Benchmark - Qwen-3.5-27B

Model Family Model Name
Qwen-3.5-27B unsloth/Qwen3.5-27B-GGUF:Q4_K_S
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_S
model size params backend ngl test t/s
qwen35 27B Q4_K - Small 14.68 GiB 26.90 B Vulkan 99 pp512 132.59 ± 0.12
qwen35 27B Q4_K - Small 14.68 GiB 26.90 B Vulkan 99 pp2048 130.23 ± 0.10
qwen35 27B Q4_K - Small 14.68 GiB 26.90 B Vulkan 99 pp4096 127.90 ± 0.02
qwen35 27B Q4_K - Small 14.68 GiB 26.90 B Vulkan 99 tg128 14.93 ± 0.01

Models

Model Family Model Name
DeepSeek-R1-Distill-Qwen-1.5B unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
DeepSeek-R1-Distill-Qwen-14B unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
Gemma-4-E2B-QAT unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
Gemma-4-E4B-QAT unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
Gemma-4-12B-QAT unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
Gemma-4-26B-A4B-QAT unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
Gemma-4-31B-QAT unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
GPT-OSS-20B unsloth/gpt-oss-20b-GGUF:Q4_K_M
LLama-3.2-1B unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
LLama-3.2-3B unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
Devstral-Small-2-24B unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
Mistral-Small-3.2-24B unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
Qwen-3.5-2B unsloth/Qwen3.5-2B-GGUF:Q4_K_M
Qwen-3.5-4B unsloth/Qwen3.5-4B-GGUF:Q4_K_M
Qwen-3.5-9B unsloth/Qwen3.5-9B-GGUF:Q4_K_M
Qwen-3.5-27B unsloth/Qwen3.5-27B-GGUF:Q4_K_S

Compared to your results for the WX9100 BIOS, I seem to get uniformly higher performance with the 220W MI25 BIOS. The difference is roughly comparable to the difference in power limit (but not exact).

BIOS power pp512 pp2048 pp4096 tg128
WX9100 170W 301 369 363 36
MI25 220W 436 430 421 39
ratio 1.29 1.45 1.17 1.14 1.08

This is of course a not a properly controlled experiment, as our systems are likely different. I am running the latest llama.cpp (tag: b10069 from July 19) .

I am picking up some performance improvement … somehow.

From the prior exercise, saw that when model size was too large for GPU memory, part of the model is on CPU, and processing slowed dramatically. Would expect slower processing, but did not expect such a dramatic difference.

Changed test script to only exercise models below a fixed size. Added the Qwen 2.5 Coder models as likely relevant for use in code-completion.

Saw a couple instances were the model got caught in some sort of processing loop, and had to be forcibly terminated. (Have seen where a public model returned nothing … perhaps the same base problem?) Wonder if this is a common problem with LLMs, and how to detect and handle?

Added a summary of benchmark results, first by model name, and then by model size.

Model Size Params pp512 pp2048 pp4096 tg128
gemma4 12B Q4_0 6.24 GiB 11.91 B 314.55 ± 0.13 287.84 ± 0.36 269.63 ± 0.04 27.72 ± 0.03
gemma4 26B.A4B Q4_0 13.26 GiB 25.23 B 644.30 ± 4.19 579.66 ± 0.37 533.55 ± 1.53 61.24 ± 0.34
gemma4 E2B Q4_0 2.43 GiB 4.63 B 1468.65 ± 2.55 1307.93 ± 0.72 1177.84 ± 0.72 94.71 ± 0.18
gemma4 E4B Q4_0 3.91 GiB 7.46 B 790.26 ± 0.52 737.20 ± 0.45 691.74 ± 0.41 58.37 ± 0.27
gpt-oss 20B Q4_K - Medium 10.81 GiB 20.91 B 758.65 ± 9.95 695.07 ± 1.29 627.88 ± 0.80 70.96 ± 0.38
llama 13B Q4_K - Small 12.61 GiB 23.57 B 144.65 ± 0.16 139.38 ± 0.07 132.82 ± 0.03 18.54 ± 0.54
llama 1B Q4_K - Medium 762.81 MiB 1.24 B 3010.95 ± 3.21 2387.34 ± 2.55 1904.03 ± 0.52 204.15 ± 5.05
llama 3B Q4_K - Medium 1.87 GiB 3.21 B 1092.33 ± 0.81 950.45 ± 0.97 810.43 ± 0.39 95.36 ± 0.28
mistral3 14B Q4_K - Medium 13.34 GiB 23.57 B 149.17 ± 0.12 143.31 ± 0.10 136.44 ± 0.04 16.58 ± 0.16
qwen2 14B Q4_K - Medium 8.37 GiB 14.77 B 250.68 ± 0.71 227.47 ± 0.28 203.32 ± 0.01 27.13 ± 0.15
qwen2 1.5B Q4_K - Medium 1.10 GiB 1.78 B 2211.66 ± 9.61 1903.15 ± 3.48 1639.08 ± 0.98 145.98 ± 2.55
qwen2 3B Q4_K - Medium 1.95 GiB 3.40 B 1104.07 ± 1.19 970.28 ± 0.83 846.95 ± 0.22 91.91 ± 0.17
qwen2 7B Q4_K - Medium 4.36 GiB 7.62 B 522.45 ± 0.05 481.33 ± 0.27 438.19 ± 0.04 48.92 ± 0.58
qwen35 2B Q4_K - Medium 1.18 GiB 1.88 B 1778.62 ± 6.07 1753.56 ± 3.03 1705.87 ± 0.75 112.37 ± 1.22
qwen35 4B Q4_K - Medium 2.54 GiB 4.21 B 774.11 ± 1.28 756.31 ± 0.36 730.66 ± 0.12 61.01 ± 0.07
qwen35 9B Q4_K - Medium 5.28 GiB 8.95 B 436.51 ± 0.30 429.72 ± 0.29 421.09 ± 0.18 39.19 ± 0.17

Performance relative to model size

From the second table we can see that Gemma-4 has some interesting outliers. Of course, this is only measure for speed of processing, not quality of results.

params size pp512 pp2048 pp4096 tg128 family / model / spec
1.24 B 762.81 MiB 3010.95 ± 3.21 2387.34 ± 2.55 1904.03 ± 0.52 204.15 ± 5.05 LLama-3.2-1B
llama 1B Q4_K - Medium
unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
1.78 B 1.10 GiB 2211.66 ± 9.61 1903.15 ± 3.48 1639.08 ± 0.98 145.98 ± 2.55 DeepSeek-R1-Distill-Qwen-1.5B
qwen2 1.5B Q4_K - Medium
unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
1.88 B 1.18 GiB 1778.62 ± 6.07 1753.56 ± 3.03 1705.87 ± 0.75 112.37 ± 1.22 Qwen-3.5-2B
qwen35 2B Q4_K - Medium
unsloth/Qwen3.5-2B-GGUF:Q4_K_M
3.21 B 1.87 GiB 1092.33 ± 0.81 950.45 ± 0.97 810.43 ± 0.39 95.36 ± 0.28 LLama-3.2-3B
llama 3B Q4_K - Medium
unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
3.40 B 1.95 GiB 1104.07 ± 1.19 970.28 ± 0.83 846.95 ± 0.22 91.91 ± 0.17 Qwen-2.5-Coder-3B
qwen2 3B Q4_K - Medium
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M
4.21 B 2.54 GiB 774.11 ± 1.28 756.31 ± 0.36 730.66 ± 0.12 61.01 ± 0.07 Qwen-3.5-4B
qwen35 4B Q4_K - Medium
unsloth/Qwen3.5-4B-GGUF:Q4_K_M
4.63 B 2.43 GiB 1468.65 ± 2.55 1307.93 ± 0.72 1177.84 ± 0.72 94.71 ± 0.18 Gemma-4-E2B-QAT
gemma4 E2B Q4_0
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
7.46 B 3.91 GiB 790.26 ± 0.52 737.20 ± 0.45 691.74 ± 0.41 58.37 ± 0.27 Gemma-4-E4B-QAT
gemma4 E4B Q4_0
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
7.62 B 4.36 GiB 522.45 ± 0.05 481.33 ± 0.27 438.19 ± 0.04 48.92 ± 0.58 Qwen-2.5-Coder-7B
qwen2 7B Q4_K - Medium
Qwen/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M
8.95 B 5.28 GiB 436.51 ± 0.30 429.72 ± 0.29 421.09 ± 0.18 39.19 ± 0.17 Qwen-3.5-9B
qwen35 9B Q4_K - Medium
unsloth/Qwen3.5-9B-GGUF:Q4_K_M
11.91 B 6.24 GiB 314.55 ± 0.13 287.84 ± 0.36 269.63 ± 0.04 27.72 ± 0.03 Gemma-4-12B-QAT
gemma4 12B Q4_0
unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
14.77 B 8.37 GiB 250.68 ± 0.71 227.47 ± 0.28 203.32 ± 0.01 27.13 ± 0.15 DeepSeek-R1-Distill-Qwen-14B
qwen2 14B Q4_K - Medium
unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
20.91 B 10.81 GiB 758.65 ± 9.95 695.07 ± 1.29 627.88 ± 0.80 70.96 ± 0.38 GPT-OSS-20B
gpt-oss 20B Q4_K - Medium
unsloth/gpt-oss-20b-GGUF:Q4_K_M
23.57 B 12.61 GiB 144.65 ± 0.16 139.38 ± 0.07 132.82 ± 0.03 18.54 ± 0.54 Mistral-Small-3.2-24B
llama 13B Q4_K - Small
unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
23.57 B 13.34 GiB 149.17 ± 0.12 143.31 ± 0.10 136.44 ± 0.04 16.58 ± 0.16 Devstral-Small-2-24B
mistral3 14B Q4_K - Medium
unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
25.23 B 13.26 GiB 644.30 ± 4.19 579.66 ± 0.37 533.55 ± 1.53 61.24 ± 0.34 Gemma-4-26B-A4B-QAT
gemma4 26B.A4B Q4_0
unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
1 Like

Added a few interesting(?) models. The table at the end measures throughput, but not quality.

Adjusted the prompt used for the one-shot smoke test, as models were failing or returning wildly variable response.

# This prompt is for testing the model's ability to summarize a book. It is not related to coding.
# Note that some models (Qwen 3.5 in particular) get stupid without specifying the year of publication.
# Note that Qwen 2.5 Coder gets stuck in a loop on this prompt.
#PROMPT='Please summarize the book from Adam Smith published in 1776 - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'

# Coding related prompt.
# Note that some models - sometimes! - get stuck in a loop on this prompt.
#PROMPT='Generate a Javascript program to compute Pi to 100 decimal places.'

# Hopefully this prompt is less likely to get stuck in a loop.
PROMPT='Generate a Javascript program to present a rotating cube in a web browser.'

Found that I could run the same prompt against the same model, with varying result:

  1. Could return in similar time.
  2. Could take quite a lot longer.
  3. Could simply stop in the middle of a response.
  4. Could get stuck in a loop.

Not sure how to deal with the last two.

These results are for models running on the 16GB AMD Instinct MI25 GPU. The Gemma-4 models tend to be outliers.

params size pp512 pp2048 pp4096 tg128 family / model / spec
1.24 B 762.81 MiB 3074.77 ± 1.89 2426.16 ± 1.22 1925.44 ± 1.74 216.27 ± 1.00 LLama-3.2-1B
llama 1B Q4_K - Medium
unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
1.78 B 1.10 GiB 2210.50 ± 5.96 1904.94 ± 1.62 1639.38 ± 1.59 147.71 ± 1.43 DeepSeek-R1-Distill-Qwen-1.5B
qwen2 1.5B Q4_K - Medium
unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
1.88 B 1.18 GiB 1782.24 ± 2.23 1754.50 ± 0.75 1708.67 ± 2.05 113.85 ± 1.28 Qwen-3.5-2B
qwen35 2B Q4_K - Medium
unsloth/Qwen3.5-2B-GGUF:Q4_K_M
3.21 B 1.87 GiB 1104.34 ± 1.35 959.58 ± 0.29 815.63 ± 1.36 99.65 ± 2.13 LLama-3.2-3B
llama 3B Q4_K - Medium
unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
3.40 B 1.95 GiB 1104.94 ± 1.05 972.45 ± 1.16 848.26 ± 0.50 93.76 ± 0.14 Qwen-2.5-Coder-3B
qwen2 3B Q4_K - Medium
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M
4.21 B 2.54 GiB 775.56 ± 0.89 757.00 ± 0.20 731.12 ± 0.31 60.97 ± 0.46 Qwen-3.5-4B
qwen35 4B Q4_K - Medium
unsloth/Qwen3.5-4B-GGUF:Q4_K_M
4.63 B 2.43 GiB 1467.58 ± 0.56 1307.88 ± 0.82 1176.73 ± 1.33 95.26 ± 0.30 Gemma-4-E2B-QAT
gemma4 E2B Q4_0
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
7.46 B 3.91 GiB 790.20 ± 0.06 738.13 ± 0.78 692.48 ± 0.12 58.62 ± 0.49 Gemma-4-E4B-QAT
gemma4 E4B Q4_0
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
7.52 B 5.35 GiB 720.25 ± 1.07 680.63 ± 1.13 641.91 ± 0.43 46.17 ± 0.16 Gemma-4-E4B-Coder
gemma4 E4B Q5_K - Medium
josephmayo/gemma-4-E4B-it-Coder-GGUF:Q5_K_M
7.62 B 4.36 GiB 522.83 ± 0.49 481.94 ± 0.32 438.49 ± 0.10 49.22 ± 0.47 Qwen-2.5-Coder-7B
qwen2 7B Q4_K - Medium
Qwen/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M
8.95 B 5.28 GiB 434.38 ± 0.34 427.68 ± 0.43 418.80 ± 0.17 39.77 ± 0.26 Qwen-3.5-9B
qwen35 9B Q4_K - Medium
unsloth/Qwen3.5-9B-GGUF:Q4_K_M
11.91 B 6.24 GiB 314.85 ± 0.37 288.12 ± 0.46 269.69 ± 0.02 27.81 ± 0.01 Gemma-4-12B-QAT
gemma4 12B Q4_0
unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
11.91 B 6.86 GiB 292.33 ± 0.19 268.73 ± 0.28 250.93 ± 0.12 27.53 ± 0.03 Gemma-4-12B-agentic
gemma4 ?B Q4_K - Medium
yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF:Q4_K_M
14.77 B 8.37 GiB 250.49 ± 0.50 227.34 ± 0.22 203.34 ± 0.03 27.25 ± 0.18 DeepSeek-R1-Distill-Qwen-14B
qwen2 14B Q4_K - Medium
unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
20.91 B 10.81 GiB 758.80 ± 10.36 695.44 ± 1.10 628.04 ± 0.92 70.91 ± 0.84 GPT-OSS-20B
gpt-oss 20B Q4_K - Medium
unsloth/gpt-oss-20b-GGUF:Q4_K_M
23.57 B 12.61 GiB 145.15 ± 0.32 139.47 ± 0.07 132.98 ± 0.02 18.83 ± 0.15 Mistral-Small-3.2-24B
llama 13B Q4_K - Small
unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
23.57 B 13.34 GiB 149.28 ± 0.12 143.46 ± 0.06 136.53 ± 0.03 16.62 ± 0.17 Devstral-Small-2-24B
mistral3 14B Q4_K - Medium
unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
25.23 B 13.26 GiB 644.46 ± 4.56 580.23 ± 0.51 534.53 ± 1.30 62.40 ± 0.11 Gemma-4-26B-A4B-QAT
gemma4 26B.A4B Q4_0
unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL

On the chance that I might need to run a second smaller model in addition to the larger model, collected measures for those models that might run on my desktop GPU.

The graphics card on my desktop has 8GB of VRAM, but is otherwise unexceptional (an AMD Radeon RX 5500 XT). Turns out this older model is not too far from the performance of the datacenter GPU (for those models that fit).

params size pp512 pp2048 pp4096 tg128 family / model / spec
1.24 B 762.81 MiB 723.82 ± 2.95 713.91 ± 1.27 636.17 ± 0.12 175.39 ± 0.18 LLama-3.2-1B
llama 1B Q4_K - Medium
unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
1.78 B 1.10 GiB 1313.24 ± 36.53 1141.64 ± 9.41 979.34 ± 6.52 132.96 ± 0.51 DeepSeek-R1-Distill-Qwen-1.5B
qwen2 1.5B Q4_K - Medium
unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
1.88 B 1.18 GiB 588.49 ± 7.66 582.02 ± 1.44 572.60 ± 0.50 33.72 ± 0.03 Qwen-3.5-2B
qwen35 2B Q4_K - Medium
unsloth/Qwen3.5-2B-GGUF:Q4_K_M
3.21 B 1.87 GiB 607.78 ± 8.15 337.52 ± 5.12 410.61 ± 1.21 72.38 ± 0.41 LLama-3.2-3B
llama 3B Q4_K - Medium
unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
3.40 B 1.95 GiB 611.93 ± 8.14 555.13 ± 4.22 488.62 ± 1.63 74.70 ± 0.25 Qwen-2.5-Coder-3B
qwen2 3B Q4_K - Medium
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M
4.21 B 2.54 GiB 438.83 ± 11.38 398.53 ± 1.43 224.74 ± 8.77 31.75 ± 0.11 Qwen-3.5-4B
qwen35 4B Q4_K - Medium
unsloth/Qwen3.5-4B-GGUF:Q4_K_M
4.63 B 2.43 GiB 761.96 ± 17.99 720.33 ± 6.12 691.04 ± 3.26 34.35 ± 0.13 Gemma-4-E2B-QAT
gemma4 E2B Q4_0
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
7.46 B 3.91 GiB 499.95 ± 12.30 452.93 ± 5.08 264.84 ± 3.09 47.44 ± 0.74 Gemma-4-E4B-QAT
gemma4 E4B Q4_0
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
7.52 B 5.35 GiB 282.33 ± 1.25 273.06 ± 0.38 183.96 ± 2.14 23.40 ± 0.07 Gemma-4-E4B-Coder
gemma4 E4B Q5_K - Medium
josephmayo/gemma-4-E4B-it-Coder-GGUF:Q5_K_M
1 Like

Again, had a village-idiot moment.

Had a list of questions accumulated that I expected to work through. Realized that of habit I was scoping-down questions presented to the online LLM (Google Gemini) as I would to a coworker. Instead presented my full questions direct … and Gemini did very well. (Not so good news for coworkers.)

Now using (possibly) optimal parameters to llama.cpp for each model. Asked for suggested models. Added large models split between GPU+CPU, and CPU only.

As before, there are some apparent outliers - measuring throughput, not quality.

This is from the server hosting an AMD Instinct MI25 16GB GPU (also 2x 8-core older Xeons with 256GB memory).

params
size
pp512
pp2048
pp4096
tg128 family / model / spec
1.24 B
1.22 GiB
3200.34 ± 2.08
2504.46 ± 4.35
1981.35 ± 1.37
169.30 ± 0.84 Llama-3.2-1B (llama 1B Q8_0)
unsloth/Llama-3.2-1B-Instruct-GGUF:Q8_0
1.78 B
1.76 GiB
2357.82 ± 7.98
2001.67 ± 2.73
1712.70 ± 0.26
131.00 ± 0.24 Qwen-2.5-Coder-1.5B (qwen2 1.5B Q8_0)
Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:Q8_0
3.21 B
3.18 GiB
1156.29 ± 1.02
998.85 ± 0.20
845.68 ± 0.71
67.16 ± 0.45 Llama-3.2-3B (llama 3B Q8_0)
unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0
3.40 B
3.36 GiB
1170.97 ± 2.70
1022.90 ± 0.46
884.34 ± 0.77
71.71 ± 1.56 Qwen-2.5-Coder-3B (qwen2 3B Q8_0)
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q8_0
4.63 B
2.43 GiB
1490.31 ± 2.64
1327.49 ± 0.96
1194.94 ± 0.86
96.23 ± 0.04 Gemma-4-E2B-QAT (gemma4 E2B Q4_0)
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
4.63 B
2.43 GiB
721.53 ± 43.34
647.34 ± 27.10
593.12 ± 10.22
20.76 ± 0.02 Gemma-4-E2B-QAT:CPU (gemma4 E2B Q4_0)
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
7.46 B
3.91 GiB
364.12 ± 20.81
340.91 ± 1.13
324.84 ± 4.50
10.21 ± 0.01 Gemma-4-E4B-QAT:CPU (gemma4 E4B Q4_0)
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
7.46 B
3.91 GiB
796.89 ± 1.05
746.93 ± 0.53
700.66 ± 0.76
61.41 ± 0.31 Gemma-4-E4B-QAT (gemma4 E4B Q4_0)
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
8.95 B
5.28 GiB
436.30 ± 0.10
430.60 ± 0.10
422.56 ± 0.08
40.12 ± 0.30 Qwen-3.5-9B (qwen35 9B Q4_K - Medium)
unsloth/Qwen3.5-9B-GGUF:Q4_K_M
11.91 B
6.24 GiB
315.02 ± 0.28
288.13 ± 0.25
270.29 ± 0.04
27.91 ± 0.06 Gemma-4-12B-QAT (gemma4 12B Q4_0)
unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
12.25 B
6.96 GiB
310.26 ± 0.37
286.61 ± 0.13
261.10 ± 0.06
33.67 ± 0.03 Mistral-Nemo-12B (llama 13B Q4_K - Medium)
bartowski/Mistral-Nemo-Instruct-2407-GGUF:Q4_K_M
14.77 B
8.37 GiB
249.30 ± 0.49
227.43 ± 0.02
203.79 ± 0.03
27.49 ± 0.26 DeepSeek-R1-Distill-Qwen-14B:GPU+CPU (qwen2 14B Q4_K - Medium)
unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
20.91 B
11.04 GiB
760.98 ± 6.42
697.44 ± 2.05
628.17 ± 1.11
70.17 ± 2.00 GPT-OSS-20B (gpt-oss 20B Q4_K - Medium)
unsloth/gpt-oss-20b-GGUF:UD-Q4_K_XL
23.57 B
13.34 GiB
149.34 ± 0.20
143.56 ± 0.07
136.83 ± 0.05
16.94 ± 0.02 Devstral-Small-2-24B (mistral3 14B Q4_K - Medium)
unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
25.23 B
13.26 GiB
645.56 ± 4.23
581.94 ± 1.23
535.61 ± 1.42
62.05 ± 0.15 Gemma-4-26B-A4B-QAT (gemma4 26B.A4B Q4_0)
unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
26.90 B
15.58 GiB
104.59 ± 4.04
98.95 ± 0.15
97.45 ± 0.06
6.17 ± 0.01 Qwen-3.5-27B:GPU+CPU (qwen35 27B Q4_K - Medium)
unsloth/Qwen3.5-27B-GGUF:Q4_K_M
32.76 B
18.48 GiB
86.91 ± 0.34
83.78 ± 0.06
79.04 ± 0.04
5.25 ± 0.01 DeepSeek-R1-Distill-32B (qwen2 32B Q4_K - Medium)
unsloth/DeepSeek-R1-Distill-Qwen-32B-GGUF:Q4_K_M
32.76 B
18.48 GiB
87.71 ± 0.36
83.81 ± 0.05
78.97 ± 0.11
5.33 ± 0.01 Qwen-2.5-Coder-32B:GPU+CPU (qwen2 32B Q4_K - Medium)
Qwen/Qwen2.5-Coder-32B-Instruct-GGUF:Q4_K_M
70.55 B
39.73 GiB
28.16 ± 0.63
26.69 ± 0.14
26.04 ± 0.11
1.05 ± 0.00 Llama-3.3-70B:CPU (llama 70B Q4_K - Medium)
unsloth/Llama-3.3-70B-Instruct-GGUF:UD-Q4_K_XL
70.55 B
39.73 GiB
29.12 ± 0.73
27.27 ± 0.13
26.26 ± 0.18
0.97 ± 0.00 DeepSeek-R1-Distill-70B:CPU (llama 70B Q4_K - Medium)
unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF:UD-Q4_K_XL
116.83 B
58.68 GiB
17.83 ± 1.92
34.91 ± 1.15
36.31 ± 0.30
8.24 ± 0.06 GPT-OSS-120B:CPU (gpt-oss 120B Q4_K - Medium)
unsloth/gpt-oss-120b-GGUF:UD-Q4_K_XL

Same updated list of models as prior (that will fit), from my desktop with an AMD Radeon RX 5500 XL 8GB GPU, and 12-core AMD Zen-3 CPU with 128GB memory.

Again, highlighted the outliers.

params
size
pp512
pp2048
pp4096
tg128 family / model / spec
1.24 B
1.22 GiB
702.28 ± 2.35
685.91 ± 2.55
627.05 ± 0.76
40.04 ± 0.02 Llama-3.2-1B (llama 1B Q8_0)
unsloth/Llama-3.2-1B-Instruct-GGUF:Q8_0
1.78 B
1.76 GiB
1796.77 ± 0.66
1633.89 ± 0.15
1281.08 ± 0.47
98.03 ± 0.17 Qwen-2.5-Coder-1.5B (qwen2 1.5B Q8_0)
Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:Q8_0
3.21 B
3.18 GiB
597.02 ± 19.65
352.87 ± 4.89
399.31 ± 3.04
29.80 ± 0.01 Llama-3.2-3B (llama 3B Q8_0)
unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0
3.40 B
3.36 GiB
874.08 ± 0.45
760.75 ± 1.58
591.51 ± 20.78
53.69 ± 0.07 Qwen-2.5-Coder-3B (qwen2 3B Q8_0)
Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q8_0
4.63 B
2.43 GiB
716.55 ± 12.39
674.87 ± 4.08
644.17 ± 0.45
37.41 ± 0.01 Gemma-4-E2B-QAT (gemma4 E2B Q4_0)
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
4.63 B
2.43 GiB
779.72 ± 44.94
814.93 ± 3.49
774.38 ± 3.09
21.71 ± 0.08 Gemma-4-E2B-QAT:CPU (gemma4 E2B Q4_0)
unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
7.46 B
3.91 GiB
416.55 ± 3.99
410.64 ± 0.75
355.50 ± 7.41
12.05 ± 0.04 Gemma-4-E4B-QAT:CPU (gemma4 E4B Q4_0)
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
7.46 B
3.91 GiB
429.02 ± 10.56
408.08 ± 4.79
245.75 ± 7.60
45.67 ± 0.84 Gemma-4-E4B-QAT (gemma4 E4B Q4_0)
unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
70.55 B
39.73 GiB
19.56 ± 0.47
18.07 ± 0.02
15.33 ± 0.04
0.91 ± 0.00 Llama-3.3-70B:CPU (llama 70B Q4_K - Medium)
unsloth/Llama-3.3-70B-Instruct-GGUF:UD-Q4_K_XL
70.55 B
39.73 GiB
19.90 ± 0.28
18.12 ± 0.04
15.34 ± 0.03
0.91 ± 0.00 DeepSeek-R1-Distill-70B:CPU (llama 70B Q4_K - Medium)
unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF:UD-Q4_K_XL
116.83 B
58.68 GiB
62.73 ± 1.89
65.86 ± 0.61
64.23 ± 1.06
8.42 ± 0.02 GPT-OSS-120B:CPU (gpt-oss 120B Q4_K - Medium)
unsloth/gpt-oss-120b-GGUF:UD-Q4_K_XL

Reached at least the initial goal.

  • Have a datacenter GPU (AMD Instinct MI25) working in a desktop computer.
  • Found LLM models suitable for desktop use in development.
  • End goal was to use the datacenter GPU to support local development.
  • Generated an example project to prove use of the local LLMs in development.

The example project (at present a work in progress) is in:

This takes advantage of carrier-grade external LLMs, as well as local LLMs.

I am still very much on the learning curve with incorporating LLMs into development, but at least now have working local LLMs for much.

This depends on a service to control the MI25 fan:

The llama.cpp setup for my rig is captured in:

At this point, the datacenter GPU is working, and I need get further on learning to make best use of LLMs. :confused:

1 Like

Should note, thought to ask an LLM about this sort of solution:

This seems to have changed with the update that happened in June because I’m seeing 86-110 tokens per second now in (Generation) Qwen3.6-35b-A3B UD_Q6-K-XL MTP.

From the January drivers to now I see a 30% speedup in MOE using llama.cpp based softwares.

1 Like

What are you getting with Vulkan?

Without MTP, I see around 110t/s TG on that same model. With MTP, it’s usually around 150-160t/s for code, 130-140t/s otherwise.

I will have to test with the Vulcan and LM Studio again, maybe tomorrow. I have not been using LM Studio and Unsloth Studio does not use Vulcan.

I can tell you that people have switched to ROCM that Vulcan did definitely not have as high performance numbers as your claims or what I mentioned above. It was worse than what I have now.

It will be interesting to see if the updates to the runtimes bring my token generation performance up drastically on Vulcan in LM Studio.