Hi there everyone, long time no see! So, I actually have my system running, but I’m running into results that are as strange as they are a pain.
So, to summarize my journey:
I started this with EPYC 7C13, with 1TB of 2444MHz RAM, and after tweaking a bit with some commands -mla 3, -rtr 1, -fmoe 1 -fa 1 -ctk q8_0, ctv q8_0
In llama-bench with IQ4_KS_R4:
pp512 34.06 +/- 0.77
tg128 3.18 +/- 0.04
Now, that was actually fast enough for myself to go on a limb, and grab myself two RTX’es 3090. And now we need to address the journey that I’ve embarked on because of these. To skip an Iliad-long journey of issues I’ve ran into, as of right now, running on PopOS, with newest Nvidia drivers, newest CUDA, and everything else up to date, I’ve sort of run into a wall that I can’t quite deal with, namely:
A - Despite claiming that the GPU’s are used, this is what I get in llama-bench
root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model | size | params | backend | ngl | fa | mla | rtr | fmoe | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | --: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB | 672.05 B | CUDA | 6 | 1 | 2 | 1 | 1 | pp512 | 36.65 ± 0.68 |
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB | 672.05 B | CUDA | 6 | 1 | 2 | 1 | 1 | tg128 | 3.28 ± 0.09 |
[17:46]
So I am getting almost entirely same performance as I did on just CPU only, which doesn’t seem right, even in the most pessimistic scenario. (If I try to offload a single layer more then 6, the entire thing goes out of memory).
(nvidia-smi sees both 3090’s, and 12.8 for cuda, and 570 for drivers.)
Trying to add some settings that I think could help performance settings seems to lead to failure to load the model. (-ts 1 in this case.)
root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64 -amb 1024 -ctk q8_0 -ctv q8_0 -ts 1
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model | size | params | backend | ngl | type_k | type_v | fa | mla | amb | ts | rtr | fmoe | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -----: | -----: | -: | --: | ----: | ------------ | --: | ---: | ------------: | ---------------: |
main: error: failed to load model '/home/user/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
[17:47]
Alright then, let’s try -ctv and -ctk as that helps drastically when running CPU only, at least in my experience.
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model | size | params | backend | ngl | type_k | type_v | fa | mla | amb | rtr | fmoe | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -----: | -----: | -: | --: | ----: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
And we’re officially at a point where I have no clue what is going wrong. entire issue seems to be based on ggml file, And to be honest, for a reason unknowable to myself, I had some real issues compiling the ik_llama, and ended up having to not use all of the flags that are recommended, ( If I try to compile with -DGGML_CUDA_IQK_FORCE_BF16=1 on, it just won’t work), as well as specifically instructing cmake to compile with 86 in mind, otherwise it fails. (PopOS ships with outdated version of CUDA, but the issue persists even after uninstalling that, and getting newest CUDA kit from Nvidia.)
What is stranger yet, is that the performance gain from GPU’s doesn’t seem to be there, and things just aren’t seemingly working.
root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64 -amb 1024 -mmp 0 -ot exps=CPU
ggml_cuda_init: GGML_CUDA_FORCE_MMQ: no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model | size | params | backend | ngl | fa | mla | amb | mmap | rtr | fmoe | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | --: | ----: | ---: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB | 672.05 B | CUDA | 6 | 1 | 2 | 1024 | 0 | 1 | 1 | pp512 | 33.25 ± 0.48 |
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB | 672.05 B | CUDA | 6 | 1 | 2 | 1024 | 0 | 1 | 1 | tg128 | 2.87 ± 0.07 |
build: b94f3af5 (3806)
Performance somehow regresses with these flags.
And then, well.
This basic command
./llama-server -m /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KSR4-00001-of-00009.gguf --ctx-size 163840 -mla 3 -fa -amb 512 -fmoe --n-gpu-layers 6 -ot "blk.(3|4).ffn.=CUDA0" -ot "blk.(5|6).ffn_.=CUDA1" --override-tensor exps=CPU --parallel 1 --threads 64 --host 127.0.0.1 --port 8080
ends up doing this at the very end
llama_new_context_with_model: KV self size = 5833.12 MiB, c^KV (q8_0): 5833.12 MiB, kv^T: not used
llama_new_context_with_model: CUDA_Host output buffer size = 0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 13584.37 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 14244241920
llama_new_context_with_model: failed to allocate compute buffers
llama_init_from_gpt_params: error: failed to create context with model '/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
^[[A ERR [ load_model] unable to load model | tid="133644715261952" timestamp=1752789224 model="/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf"
munmap_chunk(): invalid pointer
Aborted (core dumped)
[17:56]
Even root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-server -m /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KSR4-00001-of-00009.gguf --ctx-size 163840 -mla 3 -fa -amb 512 -fmoe --n-gpu-layers 6 -ot "blk.(3|4).ffn.=CUDA0" -ot "blk.(5|6).ffn_.=CUDA1" --override-tensor exps=CPU --parallel 1 --threads 64 --host 127.0.0.1 --port 8080 ends up with
[17:56]
llama_kv_cache_init: CUDA1 KV buffer size = 540.00 MiB
llama_new_context_with_model: KV self size = 10980.00 MiB, c^KV (f16): 10980.00 MiB, kv^T: not used
llama_new_context_with_model: CUDA_Host output buffer size = 0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 13833.12 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 14505073536
llama_new_context_with_model: failed to allocate compute buffers
llama_init_from_gpt_params: error: failed to create context with model '/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
ERR [ load_model] unable to load model | tid="130853199867904" timestamp=1752789316 model="/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf"
free(): invalid pointer
Aborted (core dumped)
Now, my gut feeling is that something during compilation goes so wrong as to break entire ik_llama functionality suite, and makes it so the model ends up running only on CPU. However, it does seem to load into GPU VRAM sometimes, though the GPU’s themselves are barely active when that happens, as they are seemingly just used as RAM.
However, before I will dump the unholiness that comes out with compilation when I set flags that work namely:
root@pop-os:/home/adrianna# cmake -B ./build -DGGML_CUDA=ON -DGGML_BLAS=OFF -DGGML_SCHED_MAX_COPIES=1 -DCMAKE_CUDA_ARCHITECTURES="86"
I wanted to make sure I’m not missing something, as I’ve just spent 2.5 days attempting to make this work, and it’s absolutely possible I’m not doing something obvious.
The exact software specs are:
Cuda compilation tools, release 12.9, V12.9.86
Build cuda_12.9.r12.9/compiler.36037853_0
gcc (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
NAME="Pop!_OS" VERSION="22.04 LTS" PRETTY_NAME="Pop!_OS 22.04 LTS"
NVIDIA-SMI 570.153.02 Driver Version: 570.153.02 CUDA Version: 12.8
Thank you very much! Even if just for reading until here.