Have you tried using the tensor split mode? This is the mode that actually spreads compute across GPUs, but also requires way more data passing across them.
Took a side trip to upgrade the CPU (2x Xeon at $70 for the pair). This went sideways. Not sure if it was a bent pin in the CPU socket, or a bit of gunk, but took a couple of tries. This is pretty much the limit for this old box - maxxed out for CPU, RAM, have practically infinite storage and on 10GbE.
(Clean rebuild of Linux with 32-threads takes 16-odd minutes. Not bad for near 9-hours CPU-time.)
Had the GPU overheat once, as seems the fan-control script came up before the card’s sysfs interface was ready. Updated the fan-control script to adapt.
Learned a bit more…
Seems Vulcan is improving steadily, and current Vulcan and ROCm are pretty much equivalent (but not identical). Also good news as AMD re-added support for gfx900 in ROCm 7.13.
Also should note there is a lot of churn in this space. Seems there is continuous improvement all through the stack, so you really want to be on bleeding-edge/current versions. Building from someone’s last-night Git-commit is not crazy.
Updated my list of models to test - hello “Gemma_4_QAT”.
So … at present, have a working setup. Will learn to make best use of local LLMs for development - which was the original purpose.
Still have a “to-do” list, from what I learned:
- Want to install ROCm 7.13 (or later).
Get performance or feature gain (that might matter). - Want to re-install Linux (again).
Not locked to ROCm 5.7 means not locked to Ubuntu 22.
Would slightly prefer Debian-latest (or ArchLinux?). - Build custom Linux kernel.
Some gain for CPU-specific build, and disabled mitigations. - Update the MI25 BIOS to 260 (or 300) watt limit.
Current BIOS has 220 watt limit. Cooling seems adequate.
Might be a while. Also have two Voron Trident kits (3D-printers) to build, update to R2, and add Bondtech INDX. (Another topic altogether.)
Oh that’s good to hear. I have one machine with an RX580 and a W6800 in it and the newer ROCm would just crash on the unsupported card (gfx803) instead of ignoring it and using the W6800 (I only used the RX580 for extra video outputs).
I got past it by exposing the supported AMD card to my ComfyUI Docker container:
I would recommend using the wx9100 vbios + edited pptable(using the upp tool) over the mi25 vbios with the higher PPT limit, the performance is a bit better in general.
An important note about loading a custom pptable, is that it will disable avfs even if you only edit the power limit, thus it is important to renable AVFS via the ppfeature mask , otherwise you will get excessive voltage and higher than normal power draw, I have sucessfully taken my mi25 above 400w using the wx9100 vbios using some powerful cooling on a cold day.( don’t recommend doing that unless you are certain you have a variant with the full vrm, rather than the cut down one)
Is there a good way to tell the boards apart without taking apart the heatsinks? Wasn’t sure if the boards would show differently in GPU-z or something like that.
Good hardwork.
My understanding is that if the board has 8+8pin pcie power connectors , its more likely to be a full vrm card, vs if its equiped with a 6+8pin, though generally you will know if its the weak vrm version, because using a 200w+ power limit will cause the vrm to get really hot very quickly, which visible in gpu-z sensors, or just by touching it no matter how much air flow you have.
My MI25 has dual power connectors. Should also note that I simply do not have any trouble with VRM temperatures, as they are very little different from hotspot temperatures.
As a guess, perhaps this is due to cooling the card as-designed (using a GPU fan). Whatever the reason, the GPU fan alone is more than able to cool the VRMs.
Got a bit further along. To be clear, the hardware part of the equation seems to be working well. The MI25 temperatures are well-controlled with the MI25 220W BIOS and the fan-control service. So the card seems well settled.
Also to be clear, I spent the last several years working in a cave (for a DoD contractor), so am very much catching up. There is an impressive amount of work going on around LLMs. I expect there are a lot of folk in the same space. But also I am of the habit of jumping into a new space, asking a lot of dumb questions, then finding the bleeding edge in a few months.
There is something about llama.cpp, contexts, and loading more than one model. In operating systems, there is long-settled technique around hotspots and LRU, and intelligent caching. Best I can tell, in the LLM case, this is still very primitive.
Seems llama.cpp is fairly stupid when you load more than one model, and one is used more than the other.
Of the many choices, loaded the llama.vscode extension into VSCode.
Well … guess what … llama.vscode offers you the choice to use different LLMs for “chat” and “tools” and … yeh, other things. What I would choose and why is unclear.
I have the (old) “beast” server, with dual (8-core each) Xeons, 256GB RAM, and absurd amounts of storage. Have a Zen 3 desktop (12-cores) with 128GB RAM and also absurd storage. The “beast” also has the AMD MI25, an (incidental) NVidia 1050 with 2GB VRAM. This desktop has an AMD RX 5500 with 8GB of VRAM.
So I could load different models into different spaces.
Tokens per second I am seeing with this old datacenter GPU seem acceptable.
MIght have learned a bit, so updated my list of models for test:
model_add "DeepSeek-R1-Distill-Qwen-1.5B" ":UD-Q4_K_XL" "unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF"
model_add "DeepSeek-R1-Distill-Qwen-14B" ":Q4_K_M" "unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF"
# Gemma 4 non-QAT GGUF crashes on MI25 (segfault). The QAT version below may work.
model_add "Gemma-4-E2B-QAT" ":UD-Q4_K_XL" "unsloth/gemma-4-E2B-it-qat-GGUF" "--jinja"
model_add "Gemma-4-E4B-QAT" ":UD-Q4_K_XL" "unsloth/gemma-4-E4B-it-qat-GGUF" "--jinja"
model_add "Gemma-4-12B-QAT" ":UD-Q4_K_XL" "unsloth/gemma-4-12B-it-qat-GGUF" "--jinja"
model_add "Gemma-4-26B-A4B-QAT" ":UD-Q4_K_XL" "unsloth/gemma-4-26B-A4B-it-qat-GGUF" "--jinja"
model_add "Gemma-4-31B-QAT" ":UD-Q4_K_XL" "unsloth/gemma-4-31B-it-qat-GGUF" "--jinja"
model_add "GPT-OSS-20B" ":Q4_K_M" "unsloth/gpt-oss-20b-GGUF"
model_add "LLama-3.2-1B" ":Q4_K_M" "unsloth/Llama-3.2-1B-Instruct-GGUF"
model_add "LLama-3.2-3B" ":Q4_K_M" "unsloth/Llama-3.2-3B-Instruct-GGUF"
# model_add "Microsoft-Phi-4" ":Q4_K_M" "unsloth/Phi-4-mini-instruct-GGUF"
model_add "Devstral-Small-2-24B" ":Q4_K_M" "unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF"
model_add "Mistral-Small-3.2-24B" ":Q4_K_S" "unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"
model_add "Qwen-3.5-2B" ":Q4_K_M" "unsloth/Qwen3.5-2B-GGUF"
model_add "Qwen-3.5-4B" ":Q4_K_M" "unsloth/Qwen3.5-4B-GGUF"
model_add "Qwen-3.5-9B" ":Q4_K_M" "unsloth/Qwen3.5-9B-GGUF"
model_add "Qwen-3.5-27B" ":Q4_K_S" "unsloth/Qwen3.5-27B-GGUF"
In theory, models that do not fit into GPU VRAM could be partitioned into CPU RAM. With a proper caching hierarchy, this could work well. Best I can tell, this does not work at all well. Models that end up partly in CPU memory suffer badly.
Also to be clear, I am looking for the “knee” in the curve. There is usually a point where you end up spending a lot for small improvement. In my testing, the “small” models seem to perform surprisingly well. As a first rough guess, the “Q4_K_M” models that fit in 16GB VRAM seem to be largely good enough.
So … might end up loading small models into the local GPU, and larger models into the datacenter GPU … maybe?
[Edit: Was missing Gemma-4 benchmark results.]
Download - DeepSeek-R1-Distill-Qwen-1.5B
| Model Family | Model Name |
|---|---|
| DeepSeek-R1-Distill-Qwen-1.5B | unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.08.341.702 I common_perf_print: sampling time = 574.71 ms
0.08.341.705 I common_perf_print: samplers time = 195.84 ms / 821 tokens
0.08.341.713 I common_perf_print: load time = 963.31 ms
0.08.341.717 I common_perf_print: prompt eval time = 88.29 ms / 35 tokens ( 2.52 ms per token, 396.43 tokens per second)
0.08.341.721 I common_perf_print: eval time = 5668.42 ms / 785 runs ( 7.22 ms per token, 138.49 tokens per second)
0.08.341.723 I common_perf_print: total time = 6364.00 ms / 820 tokens
0.08.341.726 I common_perf_print: unaccounted time = 32.58 ms / 0.5 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m8.457s | 0m4.732s | 0m1.022s |
Download - DeepSeek-R1-Distill-Qwen-14B
| Model Family | Model Name |
|---|---|
| DeepSeek-R1-Distill-Qwen-14B | unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.41.403.661 I common_perf_print: sampling time = 883.69 ms
0.41.403.665 I common_perf_print: samplers time = 295.72 ms / 997 tokens
0.41.403.676 I common_perf_print: load time = 6132.24 ms
0.41.403.683 I common_perf_print: prompt eval time = 409.00 ms / 35 tokens ( 11.69 ms per token, 85.57 tokens per second)
0.41.403.689 I common_perf_print: eval time = 32508.98 ms / 961 runs ( 33.83 ms per token, 29.56 tokens per second)
0.41.403.692 I common_perf_print: total time = 33853.14 ms / 996 tokens
0.41.403.695 I common_perf_print: unaccounted time = 51.46 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m41.560s | 0m17.038s | 0m4.310s |
Download - Gemma-4-E2B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-E2B-QAT | unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja
0.17.396.149 I common_perf_print: sampling time = 740.27 ms
0.17.396.152 I common_perf_print: samplers time = 254.77 ms / 1255 tokens
0.17.396.157 I common_perf_print: load time = 1324.85 ms
0.17.396.159 I common_perf_print: prompt eval time = 128.93 ms / 47 tokens ( 2.74 ms per token, 364.54 tokens per second)
0.17.396.162 I common_perf_print: eval time = 13295.65 ms / 1207 runs ( 11.02 ms per token, 90.78 tokens per second)
0.17.396.163 I common_perf_print: total time = 14197.52 ms / 1254 tokens
0.17.396.164 I common_perf_print: unaccounted time = 32.67 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m17.629s | 0m8.842s | 0m1.626s |
Download - Gemma-4-E4B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-E4B-QAT | unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja
0.27.116.205 I common_perf_print: sampling time = 1352.30 ms
0.27.116.211 I common_perf_print: samplers time = 467.16 ms / 1297 tokens
0.27.116.218 I common_perf_print: load time = 2182.16 ms
0.27.116.221 I common_perf_print: prompt eval time = 184.00 ms / 47 tokens ( 3.91 ms per token, 255.43 tokens per second)
0.27.116.228 I common_perf_print: eval time = 21471.61 ms / 1249 runs ( 17.19 ms per token, 58.17 tokens per second)
0.27.116.236 I common_perf_print: total time = 23058.86 ms / 1296 tokens
0.27.116.239 I common_perf_print: unaccounted time = 50.96 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m27.369s | 0m14.144s | 0m2.741s |
Download - Gemma-4-12B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-12B-QAT | unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja
0.47.226.558 I common_perf_print: sampling time = 1769.61 ms
0.47.226.563 I common_perf_print: samplers time = 610.27 ms / 1156 tokens
0.47.226.573 I common_perf_print: load time = 4824.19 ms
0.47.226.579 I common_perf_print: prompt eval time = 297.95 ms / 47 tokens ( 6.34 ms per token, 157.75 tokens per second)
0.47.226.584 I common_perf_print: eval time = 38375.83 ms / 1108 runs ( 34.64 ms per token, 28.87 tokens per second)
0.47.226.587 I common_perf_print: total time = 40509.26 ms / 1155 tokens
0.47.226.590 I common_perf_print: unaccounted time = 65.87 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m47.499s | 0m20.459s | 0m4.336s |
Download - Gemma-4-26B-A4B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-26B-A4B-QAT | unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja
0.39.427.510 I common_perf_print: sampling time = 1098.37 ms
0.39.427.512 I common_perf_print: samplers time = 378.43 ms / 1547 tokens
0.39.427.516 I common_perf_print: load time = 9611.67 ms
0.39.427.519 I common_perf_print: prompt eval time = 328.79 ms / 47 tokens ( 7.00 ms per token, 142.95 tokens per second)
0.39.427.521 I common_perf_print: eval time = 25684.05 ms / 1499 runs ( 17.13 ms per token, 58.36 tokens per second)
0.39.427.522 I common_perf_print: total time = 27156.52 ms / 1546 tokens
0.39.427.524 I common_perf_print: unaccounted time = 45.32 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m39.661s | 0m19.980s | 0m5.176s |
Download - Gemma-4-31B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-31B-QAT | unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-completion -hf unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.' --jinja
2.57.607.770 I common_perf_print: sampling time = 591.32 ms
2.57.607.774 I common_perf_print: samplers time = 157.10 ms / 990 tokens
2.57.607.781 I common_perf_print: load time = 9971.08 ms
2.57.607.783 I common_perf_print: prompt eval time = 2204.15 ms / 47 tokens ( 46.90 ms per token, 21.32 tokens per second)
2.57.607.785 I common_perf_print: eval time = 158801.82 ms / 942 runs ( 168.58 ms per token, 5.93 tokens per second)
2.57.607.787 I common_perf_print: total time = 161672.46 ms / 989 tokens
2.57.607.789 I common_perf_print: unaccounted time = 75.17 ms / 0.0 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 2m57.861s | 0m19.980s | 0m9.568s |
Download - GPT-OSS-20B
| Model Family | Model Name |
|---|---|
| GPT-OSS-20B | unsloth/gpt-oss-20b-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/gpt-oss-20b-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.20.527.213 I common_perf_print: sampling time = 579.61 ms
0.20.527.215 I common_perf_print: samplers time = 192.78 ms / 799 tokens
0.20.527.222 I common_perf_print: load time = 7784.85 ms
0.20.527.227 I common_perf_print: prompt eval time = 282.51 ms / 38 tokens ( 7.43 ms per token, 134.51 tokens per second)
0.20.527.231 I common_perf_print: eval time = 10211.50 ms / 760 runs ( 13.44 ms per token, 74.43 tokens per second)
0.20.527.232 I common_perf_print: total time = 11103.83 ms / 798 tokens
0.20.527.237 I common_perf_print: unaccounted time = 30.21 ms / 0.3 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m20.762s | 0m10.969s | 0m3.722s |
Download - LLama-3.2-1B
| Model Family | Model Name |
|---|---|
| LLama-3.2-1B | unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.04.781.760 I common_perf_print: sampling time = 156.94 ms
0.04.781.765 I common_perf_print: samplers time = 52.71 ms / 560 tokens
0.04.781.772 I common_perf_print: load time = 859.66 ms
0.04.781.775 I common_perf_print: prompt eval time = 51.36 ms / 42 tokens ( 1.22 ms per token, 817.76 tokens per second)
0.04.781.780 I common_perf_print: eval time = 2506.05 ms / 517 runs ( 4.85 ms per token, 206.30 tokens per second)
0.04.781.781 I common_perf_print: total time = 2726.30 ms / 559 tokens
0.04.781.791 I common_perf_print: unaccounted time = 11.94 ms / 0.4 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m4.925s | 0m2.565s | 0m0.679s |
Download - LLama-3.2-3B
| Model Family | Model Name |
|---|---|
| LLama-3.2-3B | unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.10.207.955 I common_perf_print: sampling time = 282.44 ms
0.10.207.959 I common_perf_print: samplers time = 117.30 ms / 680 tokens
0.10.207.966 I common_perf_print: load time = 1623.02 ms
0.10.207.971 I common_perf_print: prompt eval time = 99.35 ms / 42 tokens ( 2.37 ms per token, 422.73 tokens per second)
0.10.207.975 I common_perf_print: eval time = 6484.59 ms / 637 runs ( 10.18 ms per token, 98.23 tokens per second)
0.10.207.977 I common_perf_print: total time = 6891.03 ms / 679 tokens
0.10.207.979 I common_perf_print: unaccounted time = 24.65 ms / 0.4 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m10.367s | 0m5.117s | 0m1.139s |
Download - Devstral-Small-2-24B
| Model Family | Model Name |
|---|---|
| Devstral-Small-2-24B | unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.57.924.327 I common_perf_print: sampling time = 657.25 ms
0.57.924.331 I common_perf_print: samplers time = 236.95 ms / 838 tokens
0.57.924.340 I common_perf_print: load time = 9599.16 ms
0.57.924.345 I common_perf_print: prompt eval time = 555.55 ms / 35 tokens ( 15.87 ms per token, 63.00 tokens per second)
0.57.924.350 I common_perf_print: eval time = 45340.26 ms / 802 runs ( 56.53 ms per token, 17.69 tokens per second)
0.57.924.354 I common_perf_print: total time = 46596.88 ms / 837 tokens
0.57.924.357 I common_perf_print: unaccounted time = 43.83 ms / 0.1 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m58.111s | 0m22.715s | 0m5.470s |
Download - Mistral-Small-3.2-24B
| Model Family | Model Name |
|---|---|
| Mistral-Small-3.2-24B | unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S |
+ llama-completion -hf unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.29.461.436 I common_perf_print: sampling time = 289.50 ms
0.29.461.440 I common_perf_print: samplers time = 104.16 ms / 389 tokens
0.29.461.449 I common_perf_print: load time = 9206.67 ms
0.29.461.455 I common_perf_print: prompt eval time = 560.87 ms / 35 tokens ( 16.02 ms per token, 62.40 tokens per second)
0.29.461.461 I common_perf_print: eval time = 17639.85 ms / 353 runs ( 49.97 ms per token, 20.01 tokens per second)
0.29.461.463 I common_perf_print: total time = 18510.58 ms / 388 tokens
0.29.461.467 I common_perf_print: unaccounted time = 20.36 ms / 0.1 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m29.650s | 0m11.903s | 0m4.503s |
Download - Qwen-3.5-2B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-2B | unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Qwen3.5-2B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.22.867.546 I common_perf_print: sampling time = 1308.94 ms
0.22.867.549 I common_perf_print: samplers time = 432.88 ms / 2178 tokens
0.22.867.553 I common_perf_print: load time = 1358.44 ms
0.22.867.555 I common_perf_print: prompt eval time = 88.98 ms / 40 tokens ( 2.22 ms per token, 449.54 tokens per second)
0.22.867.556 I common_perf_print: eval time = 18511.10 ms / 2137 runs ( 8.66 ms per token, 115.44 tokens per second)
0.22.867.557 I common_perf_print: total time = 19962.84 ms / 2177 tokens
0.22.867.558 I common_perf_print: unaccounted time = 53.82 ms / 0.3 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m23.038s | 0m11.927s | 0m2.457s |
Download - Qwen-3.5-4B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-4B | unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.16.010.620 I common_perf_print: sampling time = 347.32 ms
0.16.010.621 I common_perf_print: samplers time = 109.96 ms / 786 tokens
0.16.010.625 I common_perf_print: load time = 2335.15 ms
0.16.010.626 I common_perf_print: prompt eval time = 140.85 ms / 40 tokens ( 3.52 ms per token, 284.00 tokens per second)
0.16.010.628 I common_perf_print: eval time = 11521.64 ms / 745 runs ( 15.47 ms per token, 64.66 tokens per second)
0.16.010.629 I common_perf_print: total time = 12031.35 ms / 785 tokens
0.16.010.632 I common_perf_print: unaccounted time = 21.55 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m16.183s | 0m7.228s | 0m1.775s |
Download - Qwen-3.5-9B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-9B | unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
+ llama-completion -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
0.57.681.435 I common_perf_print: sampling time = 2695.21 ms
0.57.681.439 I common_perf_print: samplers time = 886.45 ms / 2041 tokens
0.57.681.450 I common_perf_print: load time = 3879.13 ms
0.57.681.456 I common_perf_print: prompt eval time = 201.42 ms / 40 tokens ( 5.04 ms per token, 198.59 tokens per second)
0.57.681.462 I common_perf_print: eval time = 49290.19 ms / 2000 runs ( 24.65 ms per token, 40.58 tokens per second)
0.57.681.464 I common_perf_print: total time = 52281.45 ms / 2040 tokens
0.57.681.468 I common_perf_print: unaccounted time = 94.62 ms / 0.2 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 0m57.904s | 0m26.875s | 0m5.596s |
Download - Qwen-3.5-27B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-27B | unsloth/Qwen3.5-27B-GGUF:Q4_K_S |
+ llama-completion -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_S --single-turn --prompt 'Please summarize the book from Adam Smith - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
2.17.918.517 I common_perf_print: sampling time = 2112.98 ms
2.17.918.524 I common_perf_print: samplers time = 558.55 ms / 1689 tokens
2.17.918.533 I common_perf_print: load time = 10450.55 ms
2.17.918.538 I common_perf_print: prompt eval time = 878.25 ms / 40 tokens ( 21.96 ms per token, 45.55 tokens per second)
2.17.918.542 I common_perf_print: eval time = 119977.82 ms / 1648 runs ( 72.80 ms per token, 13.74 tokens per second)
2.17.918.545 I common_perf_print: total time = 123125.12 ms / 1688 tokens
2.17.918.551 I common_perf_print: unaccounted time = 156.07 ms / 0.1 % (total - sampling - prompt eval - eval) / (total)
| real | user | sys | time |
|---|---|---|---|
| 2m18.154s | 6m26.572s | 0m10.839s |
Benchmark - DeepSeek-R1-Distill-Qwen-1.5B
| Model Family | Model Name |
|---|---|
| DeepSeek-R1-Distill-Qwen-1.5B | unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen2 1.5B Q4_K - Medium | 1.10 GiB | 1.78 B | Vulkan | 99 | pp512 | 2198.49 ± 4.97 |
| qwen2 1.5B Q4_K - Medium | 1.10 GiB | 1.78 B | Vulkan | 99 | pp2048 | 1903.26 ± 0.99 |
| qwen2 1.5B Q4_K - Medium | 1.10 GiB | 1.78 B | Vulkan | 99 | pp4096 | 1638.22 ± 1.45 |
| qwen2 1.5B Q4_K - Medium | 1.10 GiB | 1.78 B | Vulkan | 99 | tg128 | 146.36 ± 0.16 |
| real | user | sys | time |
|---|---|---|---|
| 0m19.729s | 0m4.565s | 0m1.223s |
Benchmark - DeepSeek-R1-Distill-Qwen-14B
| Model Family | Model Name |
|---|---|
| DeepSeek-R1-Distill-Qwen-14B | unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | pp512 | 250.11 ± 0.50 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | pp2048 | 227.32 ± 0.28 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | pp4096 | 203.33 ± 0.06 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | Vulkan | 99 | tg128 | 27.35 ± 0.19 |
| real | user | sys | time |
|---|---|---|---|
| 2m25.928s | 0m29.068s | 0m7.502s |
Benchmark - Gemma-4-E2B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-E2B-QAT | unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gemma4 E2B Q4_0 | 2.43 GiB | 4.63 B | Vulkan | 99 | pp512 | 1462.95 ± 3.57 |
| gemma4 E2B Q4_0 | 2.43 GiB | 4.63 B | Vulkan | 99 | pp2048 | 1308.52 ± 1.91 |
| gemma4 E2B Q4_0 | 2.43 GiB | 4.63 B | Vulkan | 99 | pp4096 | 1177.14 ± 0.68 |
| gemma4 E2B Q4_0 | 2.43 GiB | 4.63 B | Vulkan | 99 | tg128 | 94.76 ± 0.24 |
| real | user | sys | time |
|---|---|---|---|
| 0m28.518s | 0m9.309s | 0m2.212s |
Benchmark - Gemma-4-E4B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-E4B-QAT | unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gemma4 E4B Q4_0 | 3.91 GiB | 7.46 B | Vulkan | 99 | pp512 | 790.98 ± 1.24 |
| gemma4 E4B Q4_0 | 3.91 GiB | 7.46 B | Vulkan | 99 | pp2048 | 737.58 ± 0.46 |
| gemma4 E4B Q4_0 | 3.91 GiB | 7.46 B | Vulkan | 99 | pp4096 | 692.21 ± 0.55 |
| gemma4 E4B Q4_0 | 3.91 GiB | 7.46 B | Vulkan | 99 | tg128 | 58.50 ± 0.22 |
| real | user | sys | time |
|---|---|---|---|
| 0m47.646s | 0m13.713s | 0m3.355s |
Benchmark - Gemma-4-12B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-12B-QAT | unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gemma4 ?B Q4_0 | 6.24 GiB | 11.91 B | Vulkan | 99 | pp512 | 314.82 ± 0.09 |
| gemma4 ?B Q4_0 | 6.24 GiB | 11.91 B | Vulkan | 99 | pp2048 | 287.90 ± 0.44 |
| gemma4 ?B Q4_0 | 6.24 GiB | 11.91 B | Vulkan | 99 | pp4096 | 269.61 ± 0.10 |
| gemma4 ?B Q4_0 | 6.24 GiB | 11.91 B | Vulkan | 99 | tg128 | 27.70 ± 0.02 |
| real | user | sys | time |
|---|---|---|---|
| 1m55.918s | 0m27.792s | 0m6.576s |
Benchmark - Gemma-4-26B-A4B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-26B-A4B-QAT | unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | 99 | pp512 | 644.79 ± 4.22 |
| gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | 99 | pp2048 | 579.16 ± 0.35 |
| gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | 99 | pp4096 | 533.38 ± 1.71 |
| gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | 99 | tg128 | 61.67 ± 0.07 |
| real | user | sys | time |
|---|---|---|---|
| 1m5.368s | 0m19.110s | 0m6.268s |
Benchmark - Gemma-4-31B-QAT
| Model Family | Model Name |
|---|---|
| Gemma-4-31B-QAT | unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gemma4 31B Q4_0 | 16.09 GiB | 30.70 B | Vulkan | 99 | pp512 | 13.00 ± 0.09 |
| gemma4 31B Q4_0 | 16.09 GiB | 30.70 B | Vulkan | 99 | pp2048 | 12.77 ± 0.01 |
| gemma4 31B Q4_0 | 16.09 GiB | 30.70 B | Vulkan | 99 | pp4096 | 12.59 ± 0.02 |
| gemma4 31B Q4_0 | 16.09 GiB | 30.70 B | Vulkan | 99 | tg128 | 4.11 ± 0.00 |
| real | user | sys | time |
|---|---|---|---|
| 1m5.368s | 8m12.121s | 2m0.144s |
Benchmark - GPT-OSS-20B
| Model Family | Model Name |
|---|---|
| GPT-OSS-20B | unsloth/gpt-oss-20b-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/gpt-oss-20b-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | pp512 | 760.14 ± 10.00 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | pp2048 | 695.92 ± 0.90 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | pp4096 | 628.11 ± 0.98 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | Vulkan | 99 | tg128 | 70.54 ± 1.43 |
| real | user | sys | time |
|---|---|---|---|
| 0m55.004s | 0m13.544s | 0m4.647s |
Benchmark - LLama-3.2-1B
| Model Family | Model Name |
|---|---|
| LLama-3.2-1B | unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| llama 1B Q4_K - Medium | 762.81 MiB | 1.24 B | Vulkan | 99 | pp512 | 3006.24 ± 8.26 |
| llama 1B Q4_K - Medium | 762.81 MiB | 1.24 B | Vulkan | 99 | pp2048 | 2388.89 ± 2.90 |
| llama 1B Q4_K - Medium | 762.81 MiB | 1.24 B | Vulkan | 99 | pp4096 | 1905.26 ± 0.99 |
| llama 1B Q4_K - Medium | 762.81 MiB | 1.24 B | Vulkan | 99 | tg128 | 210.30 ± 0.31 |
| real | user | sys | time |
|---|---|---|---|
| 0m16.341s | 0m3.748s | 0m1.003s |
Benchmark - LLama-3.2-3B
| Model Family | Model Name |
|---|---|
| LLama-3.2-3B | unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| llama 3B Q4_K - Medium | 1.87 GiB | 3.21 B | Vulkan | 99 | pp512 | 1092.27 ± 1.55 |
| llama 3B Q4_K - Medium | 1.87 GiB | 3.21 B | Vulkan | 99 | pp2048 | 951.35 ± 0.63 |
| llama 3B Q4_K - Medium | 1.87 GiB | 3.21 B | Vulkan | 99 | pp4096 | 811.01 ± 0.32 |
| llama 3B Q4_K - Medium | 1.87 GiB | 3.21 B | Vulkan | 99 | tg128 | 94.83 ± 1.40 |
| real | user | sys | time |
|---|---|---|---|
| 0m37.230s | 0m7.795s | 0m2.195s |
Benchmark - Devstral-Small-2-24B
| Model Family | Model Name |
|---|---|
| Devstral-Small-2-24B | unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| mistral3 14B Q4_K - Medium | 13.34 GiB | 23.57 B | Vulkan | 99 | pp512 | 149.21 ± 0.28 |
| mistral3 14B Q4_K - Medium | 13.34 GiB | 23.57 B | Vulkan | 99 | pp2048 | 143.30 ± 0.02 |
| mistral3 14B Q4_K - Medium | 13.34 GiB | 23.57 B | Vulkan | 99 | pp4096 | 136.49 ± 0.03 |
| mistral3 14B Q4_K - Medium | 13.34 GiB | 23.57 B | Vulkan | 99 | tg128 | 16.70 ± 0.04 |
| real | user | sys | time |
|---|---|---|---|
| 3m44.573s | 0m41.935s | 0m11.212s |
Benchmark - Mistral-Small-3.2-24B
| Model Family | Model Name |
|---|---|
| Mistral-Small-3.2-24B | unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| llama 13B Q4_K - Small | 12.61 GiB | 23.57 B | Vulkan | 99 | pp512 | 144.61 ± 0.19 |
| llama 13B Q4_K - Small | 12.61 GiB | 23.57 B | Vulkan | 99 | pp2048 | 139.35 ± 0.10 |
| llama 13B Q4_K - Small | 12.61 GiB | 23.57 B | Vulkan | 99 | pp4096 | 132.68 ± 0.05 |
| llama 13B Q4_K - Small | 12.61 GiB | 23.57 B | Vulkan | 99 | tg128 | 18.03 ± 0.42 |
| real | user | sys | time |
|---|---|---|---|
| 3m47.993s | 0m40.593s | 0m11.294s |
Benchmark - Qwen-3.5-2B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-2B | unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-2B-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 2B Q4_K - Medium | 1.18 GiB | 1.88 B | Vulkan | 99 | pp512 | 1784.73 ± 5.41 |
| qwen35 2B Q4_K - Medium | 1.18 GiB | 1.88 B | Vulkan | 99 | pp2048 | 1752.94 ± 2.29 |
| qwen35 2B Q4_K - Medium | 1.18 GiB | 1.88 B | Vulkan | 99 | pp4096 | 1707.30 ± 0.60 |
| qwen35 2B Q4_K - Medium | 1.18 GiB | 1.88 B | Vulkan | 99 | tg128 | 112.94 ± 1.68 |
| real | user | sys | time |
|---|---|---|---|
| 0m21.157s | 0m5.875s | 0m1.513s |
Benchmark - Qwen-3.5-4B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-4B | unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-4B-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 4B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 99 | pp512 | 775.41 ± 1.01 |
| qwen35 4B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 99 | pp2048 | 756.50 ± 0.58 |
| qwen35 4B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 99 | pp4096 | 730.54 ± 0.17 |
| qwen35 4B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 99 | tg128 | 61.02 ± 0.05 |
| real | user | sys | time |
|---|---|---|---|
| 0m45.529s | 0m10.969s | 0m2.767s |
Benchmark - Qwen-3.5-9B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-9B | unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-9B-GGUF:Q4_K_M
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 9B Q4_K - Medium | 5.28 GiB | 8.95 B | Vulkan | 99 | pp512 | 436.24 ± 0.32 |
| qwen35 9B Q4_K - Medium | 5.28 GiB | 8.95 B | Vulkan | 99 | pp2048 | 429.95 ± 0.28 |
| qwen35 9B Q4_K - Medium | 5.28 GiB | 8.95 B | Vulkan | 99 | pp4096 | 420.92 ± 0.27 |
| qwen35 9B Q4_K - Medium | 5.28 GiB | 8.95 B | Vulkan | 99 | tg128 | 39.25 ± 0.24 |
| real | user | sys | time |
|---|---|---|---|
| 1m17.347s | 0m18.605s | 0m4.788s |
Benchmark - Qwen-3.5-27B
| Model Family | Model Name |
|---|---|
| Qwen-3.5-27B | unsloth/Qwen3.5-27B-GGUF:Q4_K_S |
+ llama-bench -p 512,2048,4096 -n 128 -ngl 99 -r 3 -hf unsloth/Qwen3.5-27B-GGUF:Q4_K_S
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 27B Q4_K - Small | 14.68 GiB | 26.90 B | Vulkan | 99 | pp512 | 132.59 ± 0.12 |
| qwen35 27B Q4_K - Small | 14.68 GiB | 26.90 B | Vulkan | 99 | pp2048 | 130.23 ± 0.10 |
| qwen35 27B Q4_K - Small | 14.68 GiB | 26.90 B | Vulkan | 99 | pp4096 | 127.90 ± 0.02 |
| qwen35 27B Q4_K - Small | 14.68 GiB | 26.90 B | Vulkan | 99 | tg128 | 14.93 ± 0.01 |
Models
| Model Family | Model Name |
|---|---|
| DeepSeek-R1-Distill-Qwen-1.5B | unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
| DeepSeek-R1-Distill-Qwen-14B | unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
| Gemma-4-E2B-QAT | unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| Gemma-4-E4B-QAT | unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| Gemma-4-12B-QAT | unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
| Gemma-4-26B-A4B-QAT | unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
| Gemma-4-31B-QAT | unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL |
| GPT-OSS-20B | unsloth/gpt-oss-20b-GGUF:Q4_K_M |
| LLama-3.2-1B | unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
| LLama-3.2-3B | unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
| Devstral-Small-2-24B | unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
| Mistral-Small-3.2-24B | unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S |
| Qwen-3.5-2B | unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
| Qwen-3.5-4B | unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
| Qwen-3.5-9B | unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
| Qwen-3.5-27B | unsloth/Qwen3.5-27B-GGUF:Q4_K_S |
Compared to your results for the WX9100 BIOS, I seem to get uniformly higher performance with the 220W MI25 BIOS. The difference is roughly comparable to the difference in power limit (but not exact).
| BIOS | power | pp512 | pp2048 | pp4096 | tg128 |
|---|---|---|---|---|---|
| WX9100 | 170W | 301 | 369 | 363 | 36 |
| MI25 | 220W | 436 | 430 | 421 | 39 |
| ratio | 1.29 | 1.45 | 1.17 | 1.14 | 1.08 |
This is of course a not a properly controlled experiment, as our systems are likely different. I am running the latest llama.cpp (tag: b10069 from July 19) .
I am picking up some performance improvement … somehow.
From the prior exercise, saw that when model size was too large for GPU memory, part of the model is on CPU, and processing slowed dramatically. Would expect slower processing, but did not expect such a dramatic difference.
Changed test script to only exercise models below a fixed size. Added the Qwen 2.5 Coder models as likely relevant for use in code-completion.
Saw a couple instances were the model got caught in some sort of processing loop, and had to be forcibly terminated. (Have seen where a public model returned nothing … perhaps the same base problem?) Wonder if this is a common problem with LLMs, and how to detect and handle?
Added a summary of benchmark results, first by model name, and then by model size.
| Model | Size | Params | pp512 | pp2048 | pp4096 | tg128 |
|---|---|---|---|---|---|---|
| gemma4 12B Q4_0 | 6.24 GiB | 11.91 B | 314.55 ± 0.13 | 287.84 ± 0.36 | 269.63 ± 0.04 | 27.72 ± 0.03 |
| gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | 644.30 ± 4.19 | 579.66 ± 0.37 | 533.55 ± 1.53 | 61.24 ± 0.34 |
| gemma4 E2B Q4_0 | 2.43 GiB | 4.63 B | 1468.65 ± 2.55 | 1307.93 ± 0.72 | 1177.84 ± 0.72 | 94.71 ± 0.18 |
| gemma4 E4B Q4_0 | 3.91 GiB | 7.46 B | 790.26 ± 0.52 | 737.20 ± 0.45 | 691.74 ± 0.41 | 58.37 ± 0.27 |
| gpt-oss 20B Q4_K - Medium | 10.81 GiB | 20.91 B | 758.65 ± 9.95 | 695.07 ± 1.29 | 627.88 ± 0.80 | 70.96 ± 0.38 |
| llama 13B Q4_K - Small | 12.61 GiB | 23.57 B | 144.65 ± 0.16 | 139.38 ± 0.07 | 132.82 ± 0.03 | 18.54 ± 0.54 |
| llama 1B Q4_K - Medium | 762.81 MiB | 1.24 B | 3010.95 ± 3.21 | 2387.34 ± 2.55 | 1904.03 ± 0.52 | 204.15 ± 5.05 |
| llama 3B Q4_K - Medium | 1.87 GiB | 3.21 B | 1092.33 ± 0.81 | 950.45 ± 0.97 | 810.43 ± 0.39 | 95.36 ± 0.28 |
| mistral3 14B Q4_K - Medium | 13.34 GiB | 23.57 B | 149.17 ± 0.12 | 143.31 ± 0.10 | 136.44 ± 0.04 | 16.58 ± 0.16 |
| qwen2 14B Q4_K - Medium | 8.37 GiB | 14.77 B | 250.68 ± 0.71 | 227.47 ± 0.28 | 203.32 ± 0.01 | 27.13 ± 0.15 |
| qwen2 1.5B Q4_K - Medium | 1.10 GiB | 1.78 B | 2211.66 ± 9.61 | 1903.15 ± 3.48 | 1639.08 ± 0.98 | 145.98 ± 2.55 |
| qwen2 3B Q4_K - Medium | 1.95 GiB | 3.40 B | 1104.07 ± 1.19 | 970.28 ± 0.83 | 846.95 ± 0.22 | 91.91 ± 0.17 |
| qwen2 7B Q4_K - Medium | 4.36 GiB | 7.62 B | 522.45 ± 0.05 | 481.33 ± 0.27 | 438.19 ± 0.04 | 48.92 ± 0.58 |
| qwen35 2B Q4_K - Medium | 1.18 GiB | 1.88 B | 1778.62 ± 6.07 | 1753.56 ± 3.03 | 1705.87 ± 0.75 | 112.37 ± 1.22 |
| qwen35 4B Q4_K - Medium | 2.54 GiB | 4.21 B | 774.11 ± 1.28 | 756.31 ± 0.36 | 730.66 ± 0.12 | 61.01 ± 0.07 |
| qwen35 9B Q4_K - Medium | 5.28 GiB | 8.95 B | 436.51 ± 0.30 | 429.72 ± 0.29 | 421.09 ± 0.18 | 39.19 ± 0.17 |
Performance relative to model size
From the second table we can see that Gemma-4 has some interesting outliers. Of course, this is only measure for speed of processing, not quality of results.
| params | size | pp512 | pp2048 | pp4096 | tg128 | family / model / spec |
|---|---|---|---|---|---|---|
| 1.24 B | 762.81 MiB | 3010.95 ± 3.21 | 2387.34 ± 2.55 | 1904.03 ± 0.52 | 204.15 ± 5.05 | LLama-3.2-1B llama 1B Q4_K - Medium unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
| 1.78 B | 1.10 GiB | 2211.66 ± 9.61 | 1903.15 ± 3.48 | 1639.08 ± 0.98 | 145.98 ± 2.55 | DeepSeek-R1-Distill-Qwen-1.5B qwen2 1.5B Q4_K - Medium unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
| 1.88 B | 1.18 GiB | 1778.62 ± 6.07 | 1753.56 ± 3.03 | 1705.87 ± 0.75 | 112.37 ± 1.22 | Qwen-3.5-2B qwen35 2B Q4_K - Medium unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
| 3.21 B | 1.87 GiB | 1092.33 ± 0.81 | 950.45 ± 0.97 | 810.43 ± 0.39 | 95.36 ± 0.28 | LLama-3.2-3B llama 3B Q4_K - Medium unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
| 3.40 B | 1.95 GiB | 1104.07 ± 1.19 | 970.28 ± 0.83 | 846.95 ± 0.22 | 91.91 ± 0.17 | Qwen-2.5-Coder-3B qwen2 3B Q4_K - Medium Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M |
| 4.21 B | 2.54 GiB | 774.11 ± 1.28 | 756.31 ± 0.36 | 730.66 ± 0.12 | 61.01 ± 0.07 | Qwen-3.5-4B qwen35 4B Q4_K - Medium unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
| 4.63 B | 2.43 GiB | 1468.65 ± 2.55 | 1307.93 ± 0.72 | 1177.84 ± 0.72 | 94.71 ± 0.18 | Gemma-4-E2B-QAT gemma4 E2B Q4_0 unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B | 3.91 GiB | 790.26 ± 0.52 | 737.20 ± 0.45 | 691.74 ± 0.41 | 58.37 ± 0.27 | Gemma-4-E4B-QAT gemma4 E4B Q4_0 unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 7.62 B | 4.36 GiB | 522.45 ± 0.05 | 481.33 ± 0.27 | 438.19 ± 0.04 | 48.92 ± 0.58 | Qwen-2.5-Coder-7B qwen2 7B Q4_K - Medium Qwen/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M |
| 8.95 B | 5.28 GiB | 436.51 ± 0.30 | 429.72 ± 0.29 | 421.09 ± 0.18 | 39.19 ± 0.17 | Qwen-3.5-9B qwen35 9B Q4_K - Medium unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
| 11.91 B | 6.24 GiB | 314.55 ± 0.13 | 287.84 ± 0.36 | 269.63 ± 0.04 | 27.72 ± 0.03 | Gemma-4-12B-QAT gemma4 12B Q4_0 unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
| 14.77 B | 8.37 GiB | 250.68 ± 0.71 | 227.47 ± 0.28 | 203.32 ± 0.01 | 27.13 ± 0.15 | DeepSeek-R1-Distill-Qwen-14B qwen2 14B Q4_K - Medium unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
| 20.91 B | 10.81 GiB | 758.65 ± 9.95 | 695.07 ± 1.29 | 627.88 ± 0.80 | 70.96 ± 0.38 | GPT-OSS-20B gpt-oss 20B Q4_K - Medium unsloth/gpt-oss-20b-GGUF:Q4_K_M |
| 23.57 B | 12.61 GiB | 144.65 ± 0.16 | 139.38 ± 0.07 | 132.82 ± 0.03 | 18.54 ± 0.54 | Mistral-Small-3.2-24B llama 13B Q4_K - Small unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S |
| 23.57 B | 13.34 GiB | 149.17 ± 0.12 | 143.31 ± 0.10 | 136.44 ± 0.04 | 16.58 ± 0.16 | Devstral-Small-2-24B mistral3 14B Q4_K - Medium unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
| 25.23 B | 13.26 GiB | 644.30 ± 4.19 | 579.66 ± 0.37 | 533.55 ± 1.53 | 61.24 ± 0.34 | Gemma-4-26B-A4B-QAT gemma4 26B.A4B Q4_0 unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
Added a few interesting(?) models. The table at the end measures throughput, but not quality.
Adjusted the prompt used for the one-shot smoke test, as models were failing or returning wildly variable response.
# This prompt is for testing the model's ability to summarize a book. It is not related to coding.
# Note that some models (Qwen 3.5 in particular) get stupid without specifying the year of publication.
# Note that Qwen 2.5 Coder gets stuck in a loop on this prompt.
#PROMPT='Please summarize the book from Adam Smith published in 1776 - "Wealth of Nations" - in 3 paragraphs, and provide a list of the main points in bullet form.'
# Coding related prompt.
# Note that some models - sometimes! - get stuck in a loop on this prompt.
#PROMPT='Generate a Javascript program to compute Pi to 100 decimal places.'
# Hopefully this prompt is less likely to get stuck in a loop.
PROMPT='Generate a Javascript program to present a rotating cube in a web browser.'
Found that I could run the same prompt against the same model, with varying result:
- Could return in similar time.
- Could take quite a lot longer.
- Could simply stop in the middle of a response.
- Could get stuck in a loop.
Not sure how to deal with the last two.
These results are for models running on the 16GB AMD Instinct MI25 GPU. The Gemma-4 models tend to be outliers.
| params | size | pp512 | pp2048 | pp4096 | tg128 | family / model / spec |
|---|---|---|---|---|---|---|
| 1.24 B | 762.81 MiB | 3074.77 ± 1.89 | 2426.16 ± 1.22 | 1925.44 ± 1.74 | 216.27 ± 1.00 | LLama-3.2-1B llama 1B Q4_K - Medium unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
| 1.78 B | 1.10 GiB | 2210.50 ± 5.96 | 1904.94 ± 1.62 | 1639.38 ± 1.59 | 147.71 ± 1.43 | DeepSeek-R1-Distill-Qwen-1.5B qwen2 1.5B Q4_K - Medium unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
| 1.88 B | 1.18 GiB | 1782.24 ± 2.23 | 1754.50 ± 0.75 | 1708.67 ± 2.05 | 113.85 ± 1.28 | Qwen-3.5-2B qwen35 2B Q4_K - Medium unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
| 3.21 B | 1.87 GiB | 1104.34 ± 1.35 | 959.58 ± 0.29 | 815.63 ± 1.36 | 99.65 ± 2.13 | LLama-3.2-3B llama 3B Q4_K - Medium unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
| 3.40 B | 1.95 GiB | 1104.94 ± 1.05 | 972.45 ± 1.16 | 848.26 ± 0.50 | 93.76 ± 0.14 | Qwen-2.5-Coder-3B qwen2 3B Q4_K - Medium Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M |
| 4.21 B | 2.54 GiB | 775.56 ± 0.89 | 757.00 ± 0.20 | 731.12 ± 0.31 | 60.97 ± 0.46 | Qwen-3.5-4B qwen35 4B Q4_K - Medium unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
| 4.63 B | 2.43 GiB | 1467.58 ± 0.56 | 1307.88 ± 0.82 | 1176.73 ± 1.33 | 95.26 ± 0.30 | Gemma-4-E2B-QAT gemma4 E2B Q4_0 unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B | 3.91 GiB | 790.20 ± 0.06 | 738.13 ± 0.78 | 692.48 ± 0.12 | 58.62 ± 0.49 | Gemma-4-E4B-QAT gemma4 E4B Q4_0 unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 7.52 B | 5.35 GiB | 720.25 ± 1.07 | 680.63 ± 1.13 | 641.91 ± 0.43 | 46.17 ± 0.16 | Gemma-4-E4B-Coder gemma4 E4B Q5_K - Medium josephmayo/gemma-4-E4B-it-Coder-GGUF:Q5_K_M |
| 7.62 B | 4.36 GiB | 522.83 ± 0.49 | 481.94 ± 0.32 | 438.49 ± 0.10 | 49.22 ± 0.47 | Qwen-2.5-Coder-7B qwen2 7B Q4_K - Medium Qwen/Qwen2.5-Coder-7B-Instruct-GGUF:Q4_K_M |
| 8.95 B | 5.28 GiB | 434.38 ± 0.34 | 427.68 ± 0.43 | 418.80 ± 0.17 | 39.77 ± 0.26 | Qwen-3.5-9B qwen35 9B Q4_K - Medium unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
| 11.91 B | 6.24 GiB | 314.85 ± 0.37 | 288.12 ± 0.46 | 269.69 ± 0.02 | 27.81 ± 0.01 | Gemma-4-12B-QAT gemma4 12B Q4_0 unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
| 11.91 B | 6.86 GiB | 292.33 ± 0.19 | 268.73 ± 0.28 | 250.93 ± 0.12 | 27.53 ± 0.03 | Gemma-4-12B-agentic gemma4 ?B Q4_K - Medium yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF:Q4_K_M |
| 14.77 B | 8.37 GiB | 250.49 ± 0.50 | 227.34 ± 0.22 | 203.34 ± 0.03 | 27.25 ± 0.18 | DeepSeek-R1-Distill-Qwen-14B qwen2 14B Q4_K - Medium unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
| 20.91 B | 10.81 GiB | 758.80 ± 10.36 | 695.44 ± 1.10 | 628.04 ± 0.92 | 70.91 ± 0.84 | GPT-OSS-20B gpt-oss 20B Q4_K - Medium unsloth/gpt-oss-20b-GGUF:Q4_K_M |
| 23.57 B | 12.61 GiB | 145.15 ± 0.32 | 139.47 ± 0.07 | 132.98 ± 0.02 | 18.83 ± 0.15 | Mistral-Small-3.2-24B llama 13B Q4_K - Small unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_S |
| 23.57 B | 13.34 GiB | 149.28 ± 0.12 | 143.46 ± 0.06 | 136.53 ± 0.03 | 16.62 ± 0.17 | Devstral-Small-2-24B mistral3 14B Q4_K - Medium unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
| 25.23 B | 13.26 GiB | 644.46 ± 4.56 | 580.23 ± 0.51 | 534.53 ± 1.30 | 62.40 ± 0.11 | Gemma-4-26B-A4B-QAT gemma4 26B.A4B Q4_0 unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
On the chance that I might need to run a second smaller model in addition to the larger model, collected measures for those models that might run on my desktop GPU.
The graphics card on my desktop has 8GB of VRAM, but is otherwise unexceptional (an AMD Radeon RX 5500 XT). Turns out this older model is not too far from the performance of the datacenter GPU (for those models that fit).
| params | size | pp512 | pp2048 | pp4096 | tg128 | family / model / spec |
|---|---|---|---|---|---|---|
| 1.24 B | 762.81 MiB | 723.82 ± 2.95 | 713.91 ± 1.27 | 636.17 ± 0.12 | 175.39 ± 0.18 | LLama-3.2-1B llama 1B Q4_K - Medium unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M |
| 1.78 B | 1.10 GiB | 1313.24 ± 36.53 | 1141.64 ± 9.41 | 979.34 ± 6.52 | 132.96 ± 0.51 | DeepSeek-R1-Distill-Qwen-1.5B qwen2 1.5B Q4_K - Medium unsloth/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:UD-Q4_K_XL |
| 1.88 B | 1.18 GiB | 588.49 ± 7.66 | 582.02 ± 1.44 | 572.60 ± 0.50 | 33.72 ± 0.03 | Qwen-3.5-2B qwen35 2B Q4_K - Medium unsloth/Qwen3.5-2B-GGUF:Q4_K_M |
| 3.21 B | 1.87 GiB | 607.78 ± 8.15 | 337.52 ± 5.12 | 410.61 ± 1.21 | 72.38 ± 0.41 | LLama-3.2-3B llama 3B Q4_K - Medium unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_M |
| 3.40 B | 1.95 GiB | 611.93 ± 8.14 | 555.13 ± 4.22 | 488.62 ± 1.63 | 74.70 ± 0.25 | Qwen-2.5-Coder-3B qwen2 3B Q4_K - Medium Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q4_K_M |
| 4.21 B | 2.54 GiB | 438.83 ± 11.38 | 398.53 ± 1.43 | 224.74 ± 8.77 | 31.75 ± 0.11 | Qwen-3.5-4B qwen35 4B Q4_K - Medium unsloth/Qwen3.5-4B-GGUF:Q4_K_M |
| 4.63 B | 2.43 GiB | 761.96 ± 17.99 | 720.33 ± 6.12 | 691.04 ± 3.26 | 34.35 ± 0.13 | Gemma-4-E2B-QAT gemma4 E2B Q4_0 unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B | 3.91 GiB | 499.95 ± 12.30 | 452.93 ± 5.08 | 264.84 ± 3.09 | 47.44 ± 0.74 | Gemma-4-E4B-QAT gemma4 E4B Q4_0 unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 7.52 B | 5.35 GiB | 282.33 ± 1.25 | 273.06 ± 0.38 | 183.96 ± 2.14 | 23.40 ± 0.07 | Gemma-4-E4B-Coder gemma4 E4B Q5_K - Medium josephmayo/gemma-4-E4B-it-Coder-GGUF:Q5_K_M |
Again, had a village-idiot moment.
Had a list of questions accumulated that I expected to work through. Realized that of habit I was scoping-down questions presented to the online LLM (Google Gemini) as I would to a coworker. Instead presented my full questions direct … and Gemini did very well. (Not so good news for coworkers.)
Now using (possibly) optimal parameters to llama.cpp for each model. Asked for suggested models. Added large models split between GPU+CPU, and CPU only.
As before, there are some apparent outliers - measuring throughput, not quality.
This is from the server hosting an AMD Instinct MI25 16GB GPU (also 2x 8-core older Xeons with 256GB memory).
| params size |
pp512 pp2048 pp4096 |
tg128 | family / model / spec |
|---|---|---|---|
| 1.24 B 1.22 GiB |
3200.34 ± 2.08 2504.46 ± 4.35 1981.35 ± 1.37 |
169.30 ± 0.84 | Llama-3.2-1B (llama 1B Q8_0) unsloth/Llama-3.2-1B-Instruct-GGUF:Q8_0 |
| 1.78 B 1.76 GiB |
2357.82 ± 7.98 2001.67 ± 2.73 1712.70 ± 0.26 |
131.00 ± 0.24 | Qwen-2.5-Coder-1.5B (qwen2 1.5B Q8_0) Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:Q8_0 |
| 3.21 B 3.18 GiB |
1156.29 ± 1.02 998.85 ± 0.20 845.68 ± 0.71 |
67.16 ± 0.45 | Llama-3.2-3B (llama 3B Q8_0) unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0 |
| 3.40 B 3.36 GiB |
1170.97 ± 2.70 1022.90 ± 0.46 884.34 ± 0.77 |
71.71 ± 1.56 | Qwen-2.5-Coder-3B (qwen2 3B Q8_0) Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q8_0 |
| 4.63 B 2.43 GiB |
1490.31 ± 2.64 1327.49 ± 0.96 1194.94 ± 0.86 |
96.23 ± 0.04 | Gemma-4-E2B-QAT (gemma4 E2B Q4_0) unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 4.63 B 2.43 GiB |
721.53 ± 43.34 647.34 ± 27.10 593.12 ± 10.22 |
20.76 ± 0.02 | Gemma-4-E2B-QAT:CPU (gemma4 E2B Q4_0) unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B 3.91 GiB |
364.12 ± 20.81 340.91 ± 1.13 324.84 ± 4.50 |
10.21 ± 0.01 | Gemma-4-E4B-QAT:CPU (gemma4 E4B Q4_0) unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B 3.91 GiB |
796.89 ± 1.05 746.93 ± 0.53 700.66 ± 0.76 |
61.41 ± 0.31 | Gemma-4-E4B-QAT (gemma4 E4B Q4_0) unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 8.95 B 5.28 GiB |
436.30 ± 0.10 430.60 ± 0.10 422.56 ± 0.08 |
40.12 ± 0.30 | Qwen-3.5-9B (qwen35 9B Q4_K - Medium) unsloth/Qwen3.5-9B-GGUF:Q4_K_M |
| 11.91 B 6.24 GiB |
315.02 ± 0.28 288.13 ± 0.25 270.29 ± 0.04 |
27.91 ± 0.06 | Gemma-4-12B-QAT (gemma4 12B Q4_0) unsloth/gemma-4-12B-it-qat-GGUF:UD-Q4_K_XL |
| 12.25 B 6.96 GiB |
310.26 ± 0.37 286.61 ± 0.13 261.10 ± 0.06 |
33.67 ± 0.03 | Mistral-Nemo-12B (llama 13B Q4_K - Medium) bartowski/Mistral-Nemo-Instruct-2407-GGUF:Q4_K_M |
| 14.77 B 8.37 GiB |
249.30 ± 0.49 227.43 ± 0.02 203.79 ± 0.03 |
27.49 ± 0.26 | DeepSeek-R1-Distill-Qwen-14B:GPU+CPU (qwen2 14B Q4_K - Medium) unsloth/DeepSeek-R1-Distill-Qwen-14B-GGUF:Q4_K_M |
| 20.91 B 11.04 GiB |
760.98 ± 6.42 697.44 ± 2.05 628.17 ± 1.11 |
70.17 ± 2.00 | GPT-OSS-20B (gpt-oss 20B Q4_K - Medium) unsloth/gpt-oss-20b-GGUF:UD-Q4_K_XL |
| 23.57 B 13.34 GiB |
149.34 ± 0.20 143.56 ± 0.07 136.83 ± 0.05 |
16.94 ± 0.02 | Devstral-Small-2-24B (mistral3 14B Q4_K - Medium) unsloth/Devstral-Small-2-24B-Instruct-2512-GGUF:Q4_K_M |
| 25.23 B 13.26 GiB |
645.56 ± 4.23 581.94 ± 1.23 535.61 ± 1.42 |
62.05 ± 0.15 | Gemma-4-26B-A4B-QAT (gemma4 26B.A4B Q4_0) unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL |
| 26.90 B 15.58 GiB |
104.59 ± 4.04 98.95 ± 0.15 97.45 ± 0.06 |
6.17 ± 0.01 | Qwen-3.5-27B:GPU+CPU (qwen35 27B Q4_K - Medium) unsloth/Qwen3.5-27B-GGUF:Q4_K_M |
| 32.76 B 18.48 GiB |
86.91 ± 0.34 83.78 ± 0.06 79.04 ± 0.04 |
5.25 ± 0.01 | DeepSeek-R1-Distill-32B (qwen2 32B Q4_K - Medium) unsloth/DeepSeek-R1-Distill-Qwen-32B-GGUF:Q4_K_M |
| 32.76 B 18.48 GiB |
87.71 ± 0.36 83.81 ± 0.05 78.97 ± 0.11 |
5.33 ± 0.01 | Qwen-2.5-Coder-32B:GPU+CPU (qwen2 32B Q4_K - Medium) Qwen/Qwen2.5-Coder-32B-Instruct-GGUF:Q4_K_M |
| 70.55 B 39.73 GiB |
28.16 ± 0.63 26.69 ± 0.14 26.04 ± 0.11 |
1.05 ± 0.00 | Llama-3.3-70B:CPU (llama 70B Q4_K - Medium) unsloth/Llama-3.3-70B-Instruct-GGUF:UD-Q4_K_XL |
| 70.55 B 39.73 GiB |
29.12 ± 0.73 27.27 ± 0.13 26.26 ± 0.18 |
0.97 ± 0.00 | DeepSeek-R1-Distill-70B:CPU (llama 70B Q4_K - Medium) unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF:UD-Q4_K_XL |
| 116.83 B 58.68 GiB |
17.83 ± 1.92 34.91 ± 1.15 36.31 ± 0.30 |
8.24 ± 0.06 | GPT-OSS-120B:CPU (gpt-oss 120B Q4_K - Medium) unsloth/gpt-oss-120b-GGUF:UD-Q4_K_XL |
Same updated list of models as prior (that will fit), from my desktop with an AMD Radeon RX 5500 XL 8GB GPU, and 12-core AMD Zen-3 CPU with 128GB memory.
Again, highlighted the outliers.
| params size |
pp512 pp2048 pp4096 |
tg128 | family / model / spec |
|---|---|---|---|
| 1.24 B 1.22 GiB |
702.28 ± 2.35 685.91 ± 2.55 627.05 ± 0.76 |
40.04 ± 0.02 | Llama-3.2-1B (llama 1B Q8_0) unsloth/Llama-3.2-1B-Instruct-GGUF:Q8_0 |
| 1.78 B 1.76 GiB |
1796.77 ± 0.66 1633.89 ± 0.15 1281.08 ± 0.47 |
98.03 ± 0.17 | Qwen-2.5-Coder-1.5B (qwen2 1.5B Q8_0) Qwen/Qwen2.5-Coder-1.5B-Instruct-GGUF:Q8_0 |
| 3.21 B 3.18 GiB |
597.02 ± 19.65 352.87 ± 4.89 399.31 ± 3.04 |
29.80 ± 0.01 | Llama-3.2-3B (llama 3B Q8_0) unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0 |
| 3.40 B 3.36 GiB |
874.08 ± 0.45 760.75 ± 1.58 591.51 ± 20.78 |
53.69 ± 0.07 | Qwen-2.5-Coder-3B (qwen2 3B Q8_0) Qwen/Qwen2.5-Coder-3B-Instruct-GGUF:Q8_0 |
| 4.63 B 2.43 GiB |
716.55 ± 12.39 674.87 ± 4.08 644.17 ± 0.45 |
37.41 ± 0.01 | Gemma-4-E2B-QAT (gemma4 E2B Q4_0) unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 4.63 B 2.43 GiB |
779.72 ± 44.94 814.93 ± 3.49 774.38 ± 3.09 |
21.71 ± 0.08 | Gemma-4-E2B-QAT:CPU (gemma4 E2B Q4_0) unsloth/gemma-4-E2B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B 3.91 GiB |
416.55 ± 3.99 410.64 ± 0.75 355.50 ± 7.41 |
12.05 ± 0.04 | Gemma-4-E4B-QAT:CPU (gemma4 E4B Q4_0) unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 7.46 B 3.91 GiB |
429.02 ± 10.56 408.08 ± 4.79 245.75 ± 7.60 |
45.67 ± 0.84 | Gemma-4-E4B-QAT (gemma4 E4B Q4_0) unsloth/gemma-4-E4B-it-qat-GGUF:UD-Q4_K_XL |
| 70.55 B 39.73 GiB |
19.56 ± 0.47 18.07 ± 0.02 15.33 ± 0.04 |
0.91 ± 0.00 | Llama-3.3-70B:CPU (llama 70B Q4_K - Medium) unsloth/Llama-3.3-70B-Instruct-GGUF:UD-Q4_K_XL |
| 70.55 B 39.73 GiB |
19.90 ± 0.28 18.12 ± 0.04 15.34 ± 0.03 |
0.91 ± 0.00 | DeepSeek-R1-Distill-70B:CPU (llama 70B Q4_K - Medium) unsloth/DeepSeek-R1-Distill-Llama-70B-GGUF:UD-Q4_K_XL |
| 116.83 B 58.68 GiB |
62.73 ± 1.89 65.86 ± 0.61 64.23 ± 1.06 |
8.42 ± 0.02 | GPT-OSS-120B:CPU (gpt-oss 120B Q4_K - Medium) unsloth/gpt-oss-120b-GGUF:UD-Q4_K_XL |
Reached at least the initial goal.
- Have a datacenter GPU (AMD Instinct MI25) working in a desktop computer.
- Found LLM models suitable for desktop use in development.
- End goal was to use the datacenter GPU to support local development.
- Generated an example project to prove use of the local LLMs in development.
The example project (at present a work in progress) is in:
This takes advantage of carrier-grade external LLMs, as well as local LLMs.
I am still very much on the learning curve with incorporating LLMs into development, but at least now have working local LLMs for much.
This depends on a service to control the MI25 fan:
The llama.cpp setup for my rig is captured in:
At this point, the datacenter GPU is working, and I need get further on learning to make best use of LLMs. ![]()
Should note, thought to ask an LLM about this sort of solution:
This seems to have changed with the update that happened in June because I’m seeing 86-110 tokens per second now in (Generation) Qwen3.6-35b-A3B UD_Q6-K-XL MTP.
From the January drivers to now I see a 30% speedup in MOE using llama.cpp based softwares.
What are you getting with Vulkan?
Without MTP, I see around 110t/s TG on that same model. With MTP, it’s usually around 150-160t/s for code, 130-140t/s otherwise.
I will have to test with the Vulcan and LM Studio again, maybe tomorrow. I have not been using LM Studio and Unsloth Studio does not use Vulcan.
I can tell you that people have switched to ROCM that Vulcan did definitely not have as high performance numbers as your claims or what I mentioned above. It was worse than what I have now.
It will be interesting to see if the updates to the runtimes bring my token generation performance up drastically on Vulcan in LM Studio.