DeepSeek Deep Dive R1 at Home!

Wendell, I have no idea what you said so I asked chatgpt to tell me what your sentence means and it said this:

ChatGPT said:

It sounds like you’re asking, in terms of tokens per second (t/s), how much throughput you’d need on Yeo RTC Pro (or a similar accelerator) to outperform a CPU‑only AM5 setup.

From community measurements using AM5 CPUs with DDR5 RAM:

  • Prompt‑processing rates of around 40 t/s and token‑generation rates of 10 t/s are typical for 12B models on AM5 CPUs reddit.com+1reddit.com+1.
  • Other users report 8–22 t/s for Q4‑quantized 12B or 4B models on different Intel/AMD setups .

So to beat AM5 + CPU:

Scenario Benchmark to exceed
Prompt‑processing > 40 t/s
Overall generation > 10–20 t/s depending on model size/quant

:white_check_mark: What this means for you:

  • If Yeo RTC Pro is achieving below ~10 t/s, it’s slower than an AM5 CPU.
  • If you’re getting 20–40 t/s, you’re in AM5 CPU territory.
  • To clearly outperform, you’re aiming for ≥ 50 t/s prompt throughput, or 20 t/s+ sustained generation.

:puzzle_piece: How to proceed:

  • Which model & quantization will you run on Yeo RTC Pro? Smaller models (4–12B, Q4) require fewer t/s to beat AM5, while larger ones (20B+) need more.
  • Do you have baseline benchmarks for RTC Pro? If you share its current t/s, we can help assess how far it is from surpassing AM5 CPUs.

Let me know what model and current performance you’re seeing on the Yeo RTC Pro — then we can work out how much further optimization or hardware you’ll need to pull ahead!

Pretty funny since the YEO RTC Pro is an advanced grip designed for recurve archery. How chatGPT will follow a simple typo down the wrong rabbit hole is scary. What if it starts chasing its own tail in the middle of a ‘kill chain’? The defense dept just spent 200 million on each of the major hyperscalars. Can you imagine the hash that the models will make of all military acronyms?

Please don’t use ChatGPT for anything serious. It’s about two years behind SOTA at this point.

1 Like

My threadripper rtx pro is doing ~51 Token/s. Not sure if that is what you are asking tho.

typo, wondered two RTX pros in an am5 system vs two mac pros … t/s wise.

I think the mac clusters were around 40t/s right?

Hi there everyone, long time no see! So, I actually have my system running, but I’m running into results that are as strange as they are a pain.

So, to summarize my journey:

I started this with EPYC 7C13, with 1TB of 2444MHz RAM, and after tweaking a bit with some commands -mla 3, -rtr 1, -fmoe 1 -fa 1 -ctk q8_0, ctv q8_0

In llama-bench with IQ4_KS_R4:

pp512 34.06 +/- 0.77
tg128 3.18 +/- 0.04

Now, that was actually fast enough for myself to go on a limb, and grab myself two RTX’es 3090. And now we need to address the journey that I’ve embarked on because of these. To skip an Iliad-long journey of issues I’ve ran into, as of right now, running on PopOS, with newest Nvidia drivers, newest CUDA, and everything else up to date, I’ve sort of run into a wall that I can’t quite deal with, namely:

A - Despite claiming that the GPU’s are used, this is what I get in llama-bench

root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64 
ggml_cuda_init: GGML_CUDA_FORCE_MMQ:    no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
  Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
  Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model                          |       size |     params | backend    | ngl | fa | mla | rtr | fmoe |          test |              t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | --: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB |   672.05 B | CUDA       |   6 |  1 |   2 |   1 |    1 |         pp512 |     36.65 ± 0.68 |
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB |   672.05 B | CUDA       |   6 |  1 |   2 |   1 |    1 |         tg128 |      3.28 ± 0.09 |
[17:46]

So I am getting almost entirely same performance as I did on just CPU only, which doesn’t seem right, even in the most pessimistic scenario. (If I try to offload a single layer more then 6, the entire thing goes out of memory).
(nvidia-smi sees both 3090’s, and 12.8 for cuda, and 570 for drivers.)

Trying to add some settings that I think could help performance settings seems to lead to failure to load the model. (-ts 1 in this case.)

root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64 -amb 1024 -ctk q8_0 -ctv q8_0 -ts 1
ggml_cuda_init: GGML_CUDA_FORCE_MMQ:    no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
  Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
  Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model                          |       size |     params | backend    | ngl | type_k | type_v | fa | mla |   amb | ts           | rtr | fmoe |          test |              t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -----: | -----: | -: | --: | ----: | ------------ | --: | ---: | ------------: | ---------------: |
main: error: failed to load model '/home/user/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
[17:47]

Alright then, let’s try -ctv and -ctk as that helps drastically when running CPU only, at least in my experience.

ggml_cuda_init: GGML_CUDA_FORCE_MMQ:    no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
  Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
  Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model                          |       size |     params | backend    | ngl | type_k | type_v | fa | mla |   amb | rtr | fmoe |          test |              t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -----: | -----: | -: | --: | ----: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16

/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16

ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error

ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal errorggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16

/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
fatal error
ggml_compute_forward_dup_q: cache_k_l0 (view) -> cache_k_l0 (view) (copy) is of type f16
/home/adrianna/AI/ik_llama.cpp/ggml/src/ggml.c:10819: fatal error

And we’re officially at a point where I have no clue what is going wrong. entire issue seems to be based on ggml file, And to be honest, for a reason unknowable to myself, I had some real issues compiling the ik_llama, and ended up having to not use all of the flags that are recommended, ( If I try to compile with -DGGML_CUDA_IQK_FORCE_BF16=1 on, it just won’t work), as well as specifically instructing cmake to compile with 86 in mind, otherwise it fails. (PopOS ships with outdated version of CUDA, but the issue persists even after uninstalling that, and getting newest CUDA kit from Nvidia.)

What is stranger yet, is that the performance gain from GPU’s doesn’t seem to be there, and things just aren’t seemingly working.

root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-bench --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf -ngl 6 -fa 1 -mla 2 -fmoe 1 -rtr 1 -t 64 -amb 1024 -mmp 0 -ot exps=CPU
ggml_cuda_init: GGML_CUDA_FORCE_MMQ:    no
ggml_cuda_init: GGML_CUDA_FORCE_CUBLAS: no
ggml_cuda_init: found 2 CUDA devices:
  Device 0: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
  Device 1: NVIDIA GeForce RTX 3090, compute capability 8.6, VMM: yes
| model                          |       size |     params | backend    | ngl | fa | mla |   amb | mmap | rtr | fmoe |          test |              t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -: | --: | ----: | ---: | --: | ---: | ------------: | ---------------: |
============ Repacked 551 tensors
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB |   672.05 B | CUDA       |   6 |  1 |   2 |  1024 |    0 |   1 |    1 |         pp512 |     33.25 ± 0.48 |
| deepseek2 671B IQ4_KS_R4 - 4.25 bpw | 367.77 GiB |   672.05 B | CUDA       |   6 |  1 |   2 |  1024 |    0 |   1 |    1 |         tg128 |      2.87 ± 0.07 |

build: b94f3af5 (3806)

Performance somehow regresses with these flags.

And then, well.

This basic command

./llama-server -m /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KSR4-00001-of-00009.gguf --ctx-size 163840 -mla 3 -fa     -amb 512     -fmoe     --n-gpu-layers 6     -ot "blk.(3|4).ffn.=CUDA0"     -ot "blk.(5|6).ffn_.=CUDA1"     --override-tensor exps=CPU     --parallel 1     --threads 64     --host 127.0.0.1     --port 8080

ends up doing this at the very end

llama_new_context_with_model: KV self size  = 5833.12 MiB, c^KV (q8_0): 5833.12 MiB, kv^T: not used
llama_new_context_with_model:  CUDA_Host  output buffer size =     0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 13584.37 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 14244241920
llama_new_context_with_model: failed to allocate compute buffers
llama_init_from_gpt_params: error: failed to create context with model '/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
^[[A ERR [              load_model] unable to load model | tid="133644715261952" timestamp=1752789224 model="/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf"
munmap_chunk(): invalid pointer
Aborted (core dumped)
[17:56]
Even root@pop-os:/home/adrianna/AI/ik_llama.cpp/build/bin# ./llama-server -m /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KSR4-00001-of-00009.gguf --ctx-size 163840 -mla 3 -fa     -amb 512     -fmoe     --n-gpu-layers 6     -ot "blk.(3|4).ffn.=CUDA0"     -ot "blk.(5|6).ffn_.=CUDA1"     --override-tensor exps=CPU     --parallel 1     --threads 64     --host 127.0.0.1     --port 8080 ends up with
[17:56]
llama_kv_cache_init:      CUDA1 KV buffer size =   540.00 MiB
llama_new_context_with_model: KV self size  = 10980.00 MiB, c^KV (f16): 10980.00 MiB, kv^T: not used
llama_new_context_with_model:  CUDA_Host  output buffer size =     0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 13833.12 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 14505073536
llama_new_context_with_model: failed to allocate compute buffers
llama_init_from_gpt_params: error: failed to create context with model '/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf'
 ERR [              load_model] unable to load model | tid="130853199867904" timestamp=1752789316 model="/home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf"
free(): invalid pointer
Aborted (core dumped)

Now, my gut feeling is that something during compilation goes so wrong as to break entire ik_llama functionality suite, and makes it so the model ends up running only on CPU. However, it does seem to load into GPU VRAM sometimes, though the GPU’s themselves are barely active when that happens, as they are seemingly just used as RAM.

However, before I will dump the unholiness that comes out with compilation when I set flags that work namely:
root@pop-os:/home/adrianna# cmake -B ./build -DGGML_CUDA=ON -DGGML_BLAS=OFF -DGGML_SCHED_MAX_COPIES=1 -DCMAKE_CUDA_ARCHITECTURES="86"

I wanted to make sure I’m not missing something, as I’ve just spent 2.5 days attempting to make this work, and it’s absolutely possible I’m not doing something obvious.

The exact software specs are:

Cuda compilation tools, release 12.9, V12.9.86 
Build cuda_12.9.r12.9/compiler.36037853_0

gcc (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0

NAME="Pop!_OS" VERSION="22.04 LTS" PRETTY_NAME="Pop!_OS 22.04 LTS"

NVIDIA-SMI 570.153.02 Driver Version: 570.153.02 CUDA Version: 12.8

Thank you very much! Even if just for reading until here.

1 Like

There are a few things you can change, compile with this for deepseek on ik on cuda:

cmake -B ./build -DGGML_CUDA=ON -DGGML_BLAS=OFF -DGGML_SCHED_MAX_COPIES=1
cmake --build ./build --config Release -j $(nproc)

EDIT: removed -DGGML_CUDA_IQK_FORCE_BF16=1 as it is no longer needed as per ik directly here. Though feel free to try and a/b test speeds especially on older GPUs (3090 and earlier may benefit maybe).

Also all your commands are using -ngl 6 which is not the way to do it with big MoEs. You want to use:

-ngl 99 \
 -ot "blk\.(3|4)\.ffn_.*=CUDA0" \
 -ot "blk\.(5|6)\.ffn_.*=CUDA1" \
-ot exps=CPU \

Start with trying to get that going and holler! Also check out all the discussions on ik_llama.cpp github and the newer model cards on some of my other deepseek flavor models for more example commands.

Eventually crank it up to -ub 4096 -b 4096 if you can for more PP, but you might want to leave that off and go with -rtr instead for a little more TG given your lower RAM bandwidth.

2 Likes

First of all, I don’t know how you did it, but your compiler command worked, whist being same as mine, that didn’t O_O
I have no clue how or why, but at this point, I’m tempted to ask, do you have Ko-fi or something? Because that command solved the issue that stumped me for like two days.

However, I ran into an issue where I’d get a weird error that attempted to overprovision my VRAM to hell and back

ggml_backend_cuda_buffer_type_alloc_buffer: allocating 31329.00 MiB on device 0: cudaMalloc failed: out of memory

That seemingly got solved by parallel 1 (despite that being enabled by default supposedly, that flag somehow made a difference.)

However, I’m running into another weird issue that I know an be solved by some combination of flags and changing files. But I’ve no clues which exact combination it’d be, so:

No matter the settings, to a degree where going from no ctk and no ctv to ctk q4 and ctv q4 makes no difference, I’m still getting OOM.

For example, with ./llama-server --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf --ctx-size 100000 -ngl 99 -ot "blk\.(5|6)\.ffn_.*=CUDA1" -ot exps=CPU -fa -mla 3 --parallel 1 -sm row -ctv q8_0 -ctk q8_0 I get

Tensor blk.60.ffn_up_exps.weight buffer type overriden to CPU
llm_load_tensors: offloading 61 repeating layers to GPU
llm_load_tensors: offloading non-repeating layers to GPU
llm_load_tensors: offloaded 62/62 layers to GPU
llm_load_tensors: CUDA_Split buffer size = 17244.89 MiB
llm_load_tensors:        CPU buffer size = 38317.39 MiB
llm_load_tensors:        CPU buffer size = 42582.45 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 41420.65 MiB
llm_load_tensors:        CPU buffer size =   938.98 MiB
llm_load_tensors:      CUDA0 buffer size =   395.84 MiB
llm_load_tensors:      CUDA1 buffer size = 12445.30 MiB
....................................................................................................
llama_new_context_with_model: n_ctx      = 100096
llama_new_context_with_model: n_batch    = 2048
llama_new_context_with_model: n_ubatch   = 512
llama_new_context_with_model: flash_attn = 1
llama_new_context_with_model: mla_attn   = 3
llama_new_context_with_model: attn_max_b = 0
llama_new_context_with_model: fused_moe  = 0
llama_new_context_with_model: ser        = -1, 0
llama_new_context_with_model: freq_base  = 10000.0
llama_new_context_with_model: freq_scale = 0.025
llama_kv_cache_init:      CUDA0 KV buffer size =  3563.70 MiB
llama_new_context_with_model: KV self size  = 3563.67 MiB, c^KV (q8_0): 3563.67 MiB, kv^T: not used
llama_new_context_with_model:  CUDA_Host  output buffer size =     0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 19190.25 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 20122437632
llama_new_context_with_model: failed to allocate compute buffers 

So even though I can see my VRAM having enough of space, as that GPU is doing nothing else but waiting for that data, ik_llama still throws an OOM. And this is with no layer offloading to CUDA0 GPU, and 24GB of VRAM free. (Well, more like 23.8GB of VRAM free, you get the gist though.)

This happens even when I get really desperate and cut down on context really hard. So with this command ./llama-server --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf --ctx-size 4096 -ngl 99 -ot "blk\.(5|6)\.ffn_.*=CUDA1" -ot exps=CPU

Tensor blk.60.ffn_up_exps.weight buffer type overriden to CPU
llm_load_tensors: offloading 61 repeating layers to GPU
llm_load_tensors: offloading non-repeating layers to GPU
llm_load_tensors: offloaded 62/62 layers to GPU
llm_load_tensors: CUDA_Split buffer size = 17244.89 MiB
llm_load_tensors:        CPU buffer size = 38317.39 MiB
llm_load_tensors:        CPU buffer size = 42582.45 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 40481.67 MiB
llm_load_tensors:        CPU buffer size = 42840.67 MiB
llm_load_tensors:        CPU buffer size = 41420.65 MiB
llm_load_tensors:        CPU buffer size =   938.98 MiB
llm_load_tensors:      CUDA0 buffer size =   395.84 MiB
llm_load_tensors:      CUDA1 buffer size = 12445.30 MiB
....................................................................................................
llama_new_context_with_model: n_ctx      = 100096
llama_new_context_with_model: n_batch    = 2048
llama_new_context_with_model: n_ubatch   = 512
llama_new_context_with_model: flash_attn = 1
llama_new_context_with_model: mla_attn   = 3
llama_new_context_with_model: attn_max_b = 0
llama_new_context_with_model: fused_moe  = 0
llama_new_context_with_model: ser        = -1, 0
llama_new_context_with_model: freq_base  = 10000.0
llama_new_context_with_model: freq_scale = 0.025
llama_kv_cache_init:      CUDA0 KV buffer size =  3563.70 MiB
llama_new_context_with_model: KV self size  = 3563.67 MiB, c^KV (q8_0): 3563.67 MiB, kv^T: not used
llama_new_context_with_model:  CUDA_Host  output buffer size =     0.99 MiB
ggml_backend_cuda_buffer_type_alloc_buffer: allocating 19190.25 MiB on device 0: cudaMalloc failed: out of memory
ggml_gallocr_reserve_n: failed to allocate CUDA0 buffer of size 20122437632
llama_new_context_with_model: failed to allocate compute buffers 

So the problem persists even when context is tiny, and an entire GPU is left idle just for the buffer. (The compilation of ik_llama went fine this time, so I don’t believe this is tied to that in any way shape or form, but who knows.)

But yeah, thank you so much. Honestly, even getting to a point where I have ik_llama successfully compiled for cuda feels like a serious victory.

1 Like

Oh hey great to hear it made a difference! And yes, I’d love to keep working on bringing high quality LLMs directly into the hands of enthusiasts and end-users. I’m scraping by on my own savings plus generous hardware access from Wendell so far, so it would be much appreciated:

By default the pipeline parallel is 4 last I checked. Yeah, compiling it to 1 is a must for multi-GPU use from what I’ve seen. otherwise like you show the CUDA buffers blow up way huge. Keep in mind that --parallel 1 has nothing to do with the “pipeline parallel” llama_new_context_with_model: pipeline parallelism enabled (n_copies=1). Clear as mud yet :sweat_smile: ?

You’re getting closer, and you are right to start off with modest context just to get it up and running. Then you can slowly increase until it OOMs to find the limit. You’re leaving off one other secret weapon from ik’s fork -amb which limits the cuda compute buffers. Without it the buffers blow up huge to fit the entire computation, but he added it to allow a fixed size buffer for MLA computation that just for loops through the big data instead of allocating one huge buffer.

Here is your command massaged for ik’s fork now that you’re compiled correctly:

./llama-server \
    --model /home/adrianna/AI/DeepSeek-R1-0528-IQ4_KS_R4-00001-of-00009.gguf \
    -fa -fmoe \
    -mla 3 -amb 512 \
    --ctx-size 16384 \
    -ctk q8_0 \
    -ngl 99 \
    -ot "blk\.(3)\.ffn_.*=CUDA0" \
    -ot "blk\.(4)\.ffn_.*=CUDA1" \
    -ot exps=CPU \
    --parallel 1 \
    --threads 48 \
    --threads-batch 64 \
    --host 127.0.0.1 \
    --port 8080

If that works, then check out how much vram is free on both cards. Then you can choose your own adventure:

  • More context with -c 32768 up to 160k
  • More TG speed offloading more layers e.g. (3|4) and (5|6) etc

I think your CPU is single socket and has 64 cores? feel free to play around there but this might be about right.

Also if you want some additional support on discord to workshop your command more or explain all the options hmu in dm etc. Cheers and hope you get it going as that is a sweet rig for local LLMs!

3 Likes

Thank you so much! That -amb was everything I was missing! This works far better then I honestly expected.

With no context, in llama_server, I get something between 5.5 t/s to 6 t/s on gen, and something between 40 to 60 on pp, which is honestly way more of an uplift then I expected. (With my luck, I was bracing for something around 4 t/s at the very best.)
(This is on full fat IQ4_KS_R4 mind you. I will test IQ3_K_R4 in a bit. On CPU IQ3 haven’t really changed much, but I wonder how the VRAM util/speed changes between these two with CUDA acceleration.)
I need to play a little bit more with settings and everything before I can definitely say that this is where it’ll stay, but honestly, even if I can’t squeeze much more speed then this, I’m satisfied.

Thank you so much for your help, and if you don’t mind, I will hit you in dm’s when I have a bit of time, I have a few questions 'bout the exact way some flags work or exact ways context and cache stuff goes. (Not to be rude, but by virtue of being so advanced, it feels to me like ik_llama doesn’t have the easiest onboarding experience, though, that might be just myself not being used with software that’s not consumer-facing.)
Thanks again!

1 Like

FWIW, just going to add that @ubergarm’s old IQ2_K_R4 recipe at Huggingface, with the following changes:

blk\..*\.attn_kv_b\.weight=q8_0 like the later Kimi-K2 quants, and,
All routed experts targeted to CPU switched from IQ3_K_R4 to IQ3_KS, and IQ2_K_R4 to the brand-new IQ2_KL,

would produce a 2.930BPW/229.226GiB quant that still does 32K context off of 16GB VRAM. Perplexity is 3.4379 +/- 0.01847, quantized from unsloth’s version of DeepSeek-R1-0528-BF16 with ubergarm’s original imatrix. No surprises with processing and generation performance either, but probably not that much of a difference in perplexity to really matter in any case.

Other interesting tidbits from the experience of cooking quants for a model this large for the first time included the not-very-pleasant discovery that convert_hf_to_gguf.py does not work on fp8 versions of the model, and that the method I’ve found for casting fp8 versions to bf16 versions will OOM on 16GB VRAM. And here I thought quantizing would not be too difficult. :sweat_smile:

2 Likes

@JWNoctis

Oh nice, sounds like you’re able to cook some of your own quants?! Yes, if I were to go back and re-visit some older recipes that mix you describe indeed sounds like some good choices.

If I understand it right, this is the comparison you’re making:

  • JWNoctis/R1-0528/IQ2_KL 230GiB (2.930 BPW)

    • Final Perplexity: 3.4379 +/- 0.01847
  • ubergarm/R1-0528/IQ2_K_R4 220GIB (2.799 BPW)

    • Final Perplexity: 3.5069 +/- 0.01893

That is quite a nice improvement for not much more size! I’ve been re-working my Kimi-K2-Instruct recipes after noticing Kimi-K2 seems even more sensitive to quantization of those attn/shexp/blk.0.ffn.* layers.

If routed experts were the models “body”, and the other tensors were the models “head”, then DeepSeek would have a large body and a small head. Kimi-K2 would have a giant body and a tiny tiny head haha… So not much brains left to compress :sweat_smile:

Did you use unsloths bf16 safetensors or bf16 GGUF? If you use their GGUF it will have already been converted (but for mainline MLA discarding some important data that can be used by ik’s fork).

Given you are trying to run the original DeepSeek fp8_cast_bf16.py I’m guessing you are using the unsloth bf16 safetensors which is good as you can convert from ik’s fork thus keeping the attn_kv_b tensor for optimizing both PP and TG quality and speed fully supporting -mla 3.

And yes there are three methods for running fp8_cast_bf16 to convert the native fp8 safetensors to bf16 GGUF.

  1. eveshiron+triton-cpu method to go in one step omitting fp8_cast_bf16
  2. triton with CUDA arch >=sm89 (for native fp8e4m3 support) which will OOM on 16GB VRAM (maybe 24 even, haven’t tried it)
  3. triton-cpu method which will does not require a GPU and converted kimi-k2 using less than 100GB RAM highwater mark (would be less on deepseek psure).

I figured out the third method in my own experiments. I think I’ve documented the process somewhere, but basically just pip uninstall triton and then build triton-cpu from source like I describe in this gh issue here. Finally update the fp8_cast_bf16.py and change the device 'cuda' to 'cpu' and no more VRAM OOM!

If you decide to release your quant on hf just tag it with ik_llama.cpp so folks can find it! Thanks!

3 Likes

Hey just wondering, it is only me or I can’t see the ik lcpp repo anymore?

2 Likes

No, its not just you. I received the following notification just few minutes before the shutdown.

Namely,

ikawrakow left a comment (ikawrakow/ik_llama.cpp#616)
Btw, what is causing this sudden surge in stars?

There are two possibilities. Its either the “russian hackers” doing their stupid attack, or, alternatively, its ikawrakow who decided to shut everything down because he had enough. Given the circumstances I suggest its the russian hackers.

4 Likes

I suggest you try --tensor-split option. Here is the command to run the Deepseek R1 quant from THIREUS which has best perplexity by size with full 160k context with only three RTX 3090:


export MALLOC_CONF="background_thread:true,percpu_arena:phycpu,metadata_thp:auto,dirty_decay_ms:10000,muzzy_decay_ms:60000"
export LD_PRELOAD=/usr/local/lib/libjemalloc.so

ulimit -n 9999
CUDA_VISIBLE_DEVICES="0,1,2" \
/opt/ik_llama.cpp/ik_llama.cpp/build/bin/llama-sweep-bench \
    --warmup-batch \
    --model /opt/GGUF-Tool-Suite/GGUF-Tool-Suite/DeepSeek-R1-0528.ROOT-6.2478bpw/DeepSeek-R1-0528-THIREUS-BF16-SPECIAL_TENSOR-00001-of-01148.gguf \
    --alias THIREUS/DeepSeek-R1-0528-6.2478bpw \
    --ctx-size $((160 * 1024)) \
    -b $((16 * 512)) -ub $((8 * 512)) \
    --mlock \
    --seed 3407 \
    --temp 0.5 --top-k 0 --top-p 1.0 --min-p 0.1 --repeat-penalty 1.0 \
    -ctk q8_0 \
    -mla 3 -fa \
    -amb 512 \
    -fmoe \
    --split-mode layer \
    --tensor-split 1,2,2 \
    --main-gpu 1 \
    --override-tensor exps=CPU \
    --n-gpu-layers 99 \
    --threads $(grep ^cpu\\scores /proc/cpuinfo | uniq | awk '{print $4}' | xargs -I{} echo "{}-0" | bc) \
    --host 0.0.0.0 \
    --port 8080 \
    --lookup-cache-dynamic /mnt/data/ik_llama.kv.dump

So basically just heads up – why would you run the model that supports say 160k or 128k (as kimi-k2) with a cap of 32k? Doesn’t make sense man. Just add a second or a third GPU and use the models in fullest with the best available quants.

Hey everyone, something strange is going on, ik_llama.cpp and ik’s entire github account seem gone…

I was worried about him and emailed him, he is okay and he did not do this on purpose.

There is a reddit thread going on, he asked me to reply to it and I sent a screenshot of his email

https://www.reddit.com/r/LocalLLaMA/comments/1m4vw29/comment/n47iaq4/

@magikRUKKOLA @Panchovix

And i thought that I had simply forgotten to ssh -A but no the entire upstream is gone!

I hope he can get it back in place soon…

And yes strange about all the sudden stars… it was around 700 stars when i checked last a couple days ago maybe

5 Likes

Missed his account going missing earlier but am glad to hear he is okay, means this will be sorted some time soon. Still strange occurrences, maybe hackers, maybe someone seeing his work as competition.

2 Likes

Standing by to help out if needed; Patrons should have the video on setting up the Q1 on cpu and >= 16gb gpu w/ ik_llama.cpp early next week, so terrible timing.

3 Likes

Oh man, I hope you know what you’re doing…
Here is the thing about Q1(? lol), Q2, Q3 (even some of Q4 from unsloth) – in most of the times they are producing the code which contains the stupid syntax errors – like unescaped double quote or etc. So in the agentic workflow they are annoying – you would have to fix the syntax errors manually. This is not the experience someone would want to have with LLMs. So please make sure to point out that the quntization method is of utmost importance!!

For the R1 quants I am using the one from THIREUS. Yeah, you would sacrifice some speed but there would be no stupid bugs in the code.

Got it? :wink:

Here is the aforementioned quant from the THIREUS:


## Quant mix recipe created using Thireus' GGUF Tool Suite - https://gguf.thireus.com/
# Model name: DeepSeek-R1-0528
# Link to the original model: https://huggingface.co/deepseek-ai/DeepSeek-R1-0528

## Model head & embeddings — qbits: 32 8
output_norm\.weight=f32
token_embd\.weight=q8_0
output\.weight=q8_0

## Special attention kernels — single-quant only (llama-quantize takes care of it) — qbits: 8
blk\.([0-9]|[1-5][0-9]|60)\.attn_k_b\.weight=q8_0

## Multi-headed attention parameters — qbits: 32 4
blk\.([0-9]|[1-5][0-9]|60)\.attn_v_b\.weight=iq4_xs
blk\.([0-9]|[1-5][0-9]|60)\.attn_kv_a_norm\.weight=f32
blk\.([0-9]|[1-5][0-9]|60)\.attn_kv_a_mqa\.weight=iq4_xs
blk\.([0-9]|[1-5][0-9]|60)\.attn_output\.weight=iq4_xs
blk\.([0-9]|[1-5][0-9]|60)\.attn_kv_b\.weight=iq4_xs
blk\.([0-9]|[1-5][0-9]|60)\.attn_q_a_norm\.weight=f32
blk\.([0-9]|[1-5][0-9]|60)\.attn_norm\.weight=f32
blk\.([0-9]|[1-5][0-9]|60)\.attn_q_a\.weight=iq4_xs
blk\.([0-9]|[1-5][0-9]|60)\.attn_q_b\.weight=iq4_xs

## Core FFN weights — qbits: 32 8 6 5
blk\.2\.ffn_gate\.weight=q8_0
blk\.(0|2)\.ffn_up\.weight=iq6_k
blk\.([0-9]|[1-5][0-9]|60)\.ffn_norm\.weight=f32
blk\.[0-1]\.ffn_gate\.weight=iq6_k
blk\.1\.ffn_down\.weight=iq6_k
blk\.2\.ffn_down\.weight=iq5_k_r4
blk\.1\.ffn_up\.weight=iq5_k_r4
blk\.([3-9]|[1-5][0-9]|60)\.ffn_gate_inp\.weight=f32
blk\.0\.ffn_down\.weight=q8_0

## Other tensors — qbits: 32
blk\.([3-9]|[1-5][0-9]|60)\.exp_probs_b\.bias=f32

## GPU-loaded ffn_*_shexp
# ffn_down_shexp (down-projection) — qbits: 8 6 5
blk\.(11|17|19|29|36|39|44|60|2[6-7]|2[0-4]|3[0-1]|3[3-4])\.ffn_down_shexp\.weight=q8_0
blk\.([3-8]|10|12|25|28|32|35|3[7-8]|1[4-6]|4[5-9]|4[0-3]|5[0-8])\.ffn_down_shexp\.weight=iq6_k
blk\.(9|13|18|59)\.ffn_down_shexp\.weight=iq5_k_r4

# ffn_up_shexp (up-projection) — qbits: 8 6 5
blk\.(6|15|18|30|37|39|41|50|54|60|2[1-4]|3[2-4]|2[6-9])\.ffn_up_shexp\.weight=q8_0
blk\.([3-5]|[8-9]|19|20|25|31|38|40|58|4[2-9]|1[6-7]|1[0-4]|3[5-6]|5[5-6]|5[1-3])\.ffn_up_shexp\.weight=iq6_k
blk\.(7|57|59)\.ffn_up_shexp\.weight=iq5_k_r4

# ffn_gate_shexp (gate-projection) — qbits: 8 6 5
blk\.(16|20|29|54|60|5[6-8]|5[0-2]|4[1-2]|4[4-9]|1[8-9]|2[3-6]|3[3-4])\.ffn_gate_shexp\.weight=q8_0
blk\.([3-5]|[7-9]|17|21|40|43|53|55|3[0-2]|2[7-8]|3[5-9]|1[1-5])\.ffn_gate_shexp\.weight=iq6_k
blk\.(6|10|22|59)\.ffn_gate_shexp\.weight=iq5_k_r4

## CPU-loaded ffn_*_exps
# ffn_down_exps (down-extraction) — qbits: 8 5 3
blk\.(51|53|3[2-9]|4[0-9])\.ffn_down_exps\.weight=q8_0
blk\.([3-9]|50|52|60|5[4-9]|1[0-4]|2[0-9]|3[0-1]|1[6-9])\.ffn_down_exps\.weight=iq5_k_r4
blk\.15\.ffn_down_exps\.weight=iq3_k

# ffn_up_exps (up-extraction) — qbits: 8 5 4
blk\.(35|53|55|4[7-8]|5[0-1]|4[3-4])\.ffn_up_exps\.weight=q8_0
blk\.([3-9]|49|52|54|60|4[0-2]|1[1-9]|3[0-4]|2[0-9]|4[5-6]|3[6-9]|5[6-9])\.ffn_up_exps\.weight=iq5_k_r4
blk\.10\.ffn_up_exps\.weight=iq4_ks

# ffn_gate_exps (gate-extraction) — qbits: 8 5 4
blk\.(35|39|41|60|5[0-5]|4[3-9])\.ffn_gate_exps\.weight=q8_0
blk\.([3-7]|9|[1-2][0-9]|40|42|3[6-8]|3[0-4]|5[6-9])\.ffn_gate_exps\.weight=iq5_k_r4
blk\.8\.ffn_gate_exps\.weight=iq4_ks

## Summary of tensor sizes per class
# GPU Total: 11.744 GiB (95.1%) | 12.34 GiB max, if all were q8_0 | 10.39 GiB min, if all were iq5_k_r4
# CPU Total: 477.066 GiB (73.7%) | 647.06 GiB max, if all were q8_0 | 261.68 GiB min, if all were iq3_k
# GPU+CPU Total: 488.811 GiB (84.4%)

## Summary of tensor counts and bpw per qtype
#
# GPU-loaded quants:
# QTYPE         Count   BPW     Assigned GiB    % Assigned      Max GiB (all)
# +f32          361     32.0      0.40 GiB      -               -
# +q8_0         61      8.5       0.51 GiB      -               -
# q8_0          71      8.5       3.07 GiB      55.4%           5.54
# iq6_k         101     6.625     1.60 GiB      37.0%           4.32
# iq5_k_r4      13      5.5       0.27 GiB      7.6%            3.58
# +iq4_xs       366     4.25      5.90 GiB      -               -
#
# CPU-loaded quants:
# QTYPE         Count   BPW     Assigned GiB    % Assigned      Max GiB (all)
# q8_0          46      8.5     171.06 GiB      26.4%           647.06
# iq5_k_r4      125     5.5     300.78 GiB      71.8%           418.69
# iq4_ks        2       4.25      3.72 GiB      1.1%            323.53
# iq3_k         1       3.4375    1.50 GiB      0.6%            261.68
#
# -Average BPW: 6.2478
#
# -Notes:
# - '+' means user-defined pre-assigned tensors and f32 tensors
# - Recipe produced on the 2025-07-16 19:21:22 UTC+0000 using Thireus' GGUF tools (https://gguf.thireus.com/)
# - Script SHA-256: 3c88ec66185ed0999d6be95e1d8e5fb2d22000c404863f0c2fa301a44160f8c3
# - Command used:
# quant_assign.py ppl_results.csv --tolerance 0.01 --cpu-irq-k 1.5 --gpu-irq-k 1.5 --gpu-assign-qtype iq4_xs \
# --cpu-tensors-max-size 500 --gpu-tensors-max-size 95% --exponential-factor 8 --cpu-tensors \
# 'blk\.([3-9]|[1-5][0-9]|60)\.ffn_down_exps\.weight' 'blk\.([3-9]|[1-5][0-9]|60)\.ffn_up_exps\.weight' \
# 'blk\.([3-9]|[1-5][0-9]|60)\.ffn_gate_exps\.weight' --gpu-tensors '.*' --cpu-quants iq4_ks iq3_k iq5_k_r4 q8_0 \
# --gpu-quants q8_0 iq5_k_r4 iq6_k --gpu-assign-tensors 'blk\.([0-9]|[1-5][0-9]|60)\.attn_k_b\.weight=q8_0'

## THE END!
# Saved recipe to file: DeepSeek-R1-0528.ROOT-6.2478bpw-0.0000ppl.488GB-GGUF_11GB-GPU_477GB-CPU.3c88ec6_c3039f4.recipe

It fits 512GB RAM and 72GB VRAM (three GPUs) with 160k context.

And you’re doing what exactly? Trying to run the Q1 with two-channel RAM? lol. I can’t see much sense. The result would be somewhat unsatisfactory (as I pointed out). I mean, yeah, some stupid quant would be able to produce some python code for a bouncing balls inside the heptagon without bugs. But so what? How it would relate to the actual workflow of the software developer?

I am asking these questions because in the 1960’s there already been the time when some researchers overhyped the LLMs and some stupid person like Marvin Minsky came by and pointed out that MLP (multi-layer perceptrons) couldn’t do such and such. And it all came down to the cold winter of AI. Just because of the overhype and some stupid persons like Marvin Minsky. So if you are the influencer or some shit please make sure to notify everyone regarding the opportunities and limitations!!

have you tried ours? it has the best (edit: hah 2nd best. top 5 surely. lol) perplexity on HF. yes it’s not perfect but it’s surprisingly good.

no the video is about context window in this case with q1. it runs at about 5 tokens/sec and can process a 16k context in seconds on the “normal” desktop computer with the ik llama.cpp … I do a few demos around that, not code creation or syntax. I did caution that you are losing important stuff in quantization but it can still be useful.

I also fed it a version of Othello with all the names changed and the reasoning model knew it was Othello and tried to “deconfuse” me that this was Othello.

I then fed it some pride and prejudice and zombies to see if that would trip it up and it answered questions correctly.

3 Likes