My hardware & llama.cpp setup
$ inxi
CPU: 16-core AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (-MT MCP-) speed/min/max: 2000/625/5188 MHz
Kernel: 7.1.4-204.fc44.x86_64 x86_64 Up: 2d 11h 6m Mem: 2.57/124.93 GiB (2.1%)
Storage: 1.82 TiB (39.2% used) Procs: 479 Shell: Bash inxi: 3.3.41
I have setup GTT correctly, this in my kernel command line: amdgpu.gttsize=122880.
I compiled my own llama.cpp, with ROCm & Vulkan dynamically loaded: -DBUILD_SHARED_LIBS=1 -DGGML_BACKEND_DL=1 -DGGML_CPU_ALL_VARIANTS=1.
The issue with ROCm
I was using llama-swap to manage multiple models, but then I noticed it was switching models when technically they should both fit. So I tried to load multiple models using the CLI manually, and it failed with an out of memory error!
0.00.509.134 I srv load_model: loading model 'unsloth/Ornith-1.0-35B-GGUF:UD-Q8_K_XL'
0.01.202.476 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 35905.99 MiB on device 0: cudaMalloc failed: out of memory
0.01.202.482 E alloc_tensor_range: failed to allocate ROCm0 buffer of size 37650156032
0.01.234.557 E llama_model_load: error loading model: unable to allocate ROCm0 buffer
0.01.234.564 E llama_model_load_from_file_impl: failed to load model
0.01.234.572 E cmn common_init_: failed to load model '/path/to/model.gguf'
0.01.234.577 E srv load_model: failed to load model, '/path/to/model.gguf'
0.01.234.580 I srv operator(): operator(): cleaning up before exit...
0.01.236.291 E srv llama_server: exiting due to model loading error
The 2 models I was trying are unsloth/Ornith-1.0-35B-GGUF:UD-Q8_K_XL (38.2 GB) and unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL (22.9 GB). These should easily fit in 120 gigs. Then I noticed, when I list devices, the memory listed for ROCm and Vulkan are different:
$ llama-server --list-devices
Available devices:
ROCm0: AMD Radeon 8060S Graphics (63965 MiB, 125203 MiB free)
Vulkan0: AMD Radeon 8060S Graphics (RADV STRIX_HALO) (123392 MiB, 123052 MiB free)
Both backends detect the entire memory, however the first number is different. I’m presuming that’s the amount of memory that’s accessible by the device. So then I tried to load two models that would definitely fit inside 63 gigs, and it worked!
When I repeat this exercise with Vulkan (I set device like this --device ROCm0/Vulkan0), then I don’t see any issues. I could even load 3 different models!
How do I get ROCm to see and use the entire memory?