There are a few things you can change, compile with this for deepseek on ik on cuda:
cmake -B ./build -DGGML_CUDA=ON -DGGML_BLAS=OFF -DGGML_SCHED_MAX_COPIES=1
cmake --build ./build --config Release -j $(nproc)
EDIT: removed -DGGML_CUDA_IQK_FORCE_BF16=1 as it is no longer needed as per ik directly here. Though feel free to try and a/b test speeds especially on older GPUs (3090 and earlier may benefit maybe).
Also all your commands are using -ngl 6 which is not the way to do it with big MoEs. You want to use:
-ngl 99 \
-ot "blk\.(3|4)\.ffn_.*=CUDA0" \
-ot "blk\.(5|6)\.ffn_.*=CUDA1" \
-ot exps=CPU \
Start with trying to get that going and holler! Also check out all the discussions on ik_llama.cpp github and the newer model cards on some of my other deepseek flavor models for more example commands.
Eventually crank it up to -ub 4096 -b 4096 if you can for more PP, but you might want to leave that off and go with -rtr instead for a little more TG given your lower RAM bandwidth.