ROCm can only use half of the unified memory on my Strix Halo

My hardware & llama.cpp setup

$ inxi 
CPU: 16-core AMD RYZEN AI MAX+ 395 w/ Radeon 8060S (-MT MCP-) speed/min/max: 2000/625/5188 MHz
Kernel: 7.1.4-204.fc44.x86_64 x86_64 Up: 2d 11h 6m Mem: 2.57/124.93 GiB (2.1%)
Storage: 1.82 TiB (39.2% used) Procs: 479 Shell: Bash inxi: 3.3.41

I have setup GTT correctly, this in my kernel command line: amdgpu.gttsize=122880.

I compiled my own llama.cpp, with ROCm & Vulkan dynamically loaded: -DBUILD_SHARED_LIBS=1 -DGGML_BACKEND_DL=1 -DGGML_CPU_ALL_VARIANTS=1.

The issue with ROCm

I was using llama-swap to manage multiple models, but then I noticed it was switching models when technically they should both fit. So I tried to load multiple models using the CLI manually, and it failed with an out of memory error!

0.00.509.134 I srv    load_model: loading model 'unsloth/Ornith-1.0-35B-GGUF:UD-Q8_K_XL'
0.01.202.476 E ggml_backend_cuda_buffer_type_alloc_buffer: allocating 35905.99 MiB on device 0: cudaMalloc failed: out of memory
0.01.202.482 E alloc_tensor_range: failed to allocate ROCm0 buffer of size 37650156032
0.01.234.557 E llama_model_load: error loading model: unable to allocate ROCm0 buffer
0.01.234.564 E llama_model_load_from_file_impl: failed to load model
0.01.234.572 E cmn  common_init_: failed to load model '/path/to/model.gguf'
0.01.234.577 E srv    load_model: failed to load model, '/path/to/model.gguf'
0.01.234.580 I srv    operator(): operator(): cleaning up before exit...
0.01.236.291 E srv  llama_server: exiting due to model loading error

The 2 models I was trying are unsloth/Ornith-1.0-35B-GGUF:UD-Q8_K_XL (38.2 GB) and unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL (22.9 GB). These should easily fit in 120 gigs. Then I noticed, when I list devices, the memory listed for ROCm and Vulkan are different:

$ llama-server --list-devices 
Available devices:
  ROCm0: AMD Radeon 8060S Graphics (63965 MiB, 125203 MiB free)
  Vulkan0: AMD Radeon 8060S Graphics (RADV STRIX_HALO) (123392 MiB, 123052 MiB free)

Both backends detect the entire memory, however the first number is different. I’m presuming that’s the amount of memory that’s accessible by the device. So then I tried to load two models that would definitely fit inside 63 gigs, and it worked!

When I repeat this exercise with Vulkan (I set device like this --device ROCm0/Vulkan0), then I don’t see any issues. I could even load 3 different models!

How do I get ROCm to see and use the entire memory?

Give me a minute I am gonna answer that.

1 Like

GTT is the deprecated way and might not work anymore. You also need to configure TTM.

For example on my 128GB system I have added the following kernel parameters:

ttm.pages_limit=27648000 ttm.page_pool_size=27648000

I think these values give you 108GB that can be used by the iGPU. The value there the 27648000 is the amount of 4KiB pages that are being assigned. You can also install the amd-debug-tool and then use the amd-ttm command to read and set these values automatically.

At least that is my guess why you can’t use the memory.

2 Likes

Interesting, didn’t know about this! I’ll experiment. I installed amd-debug-tool (python3-amd-debug-tools on Fedora).

$ amd-ttm 
💻 Current TTM pages limit: 16375124 pages (62.47 GB)
💻 Total system memory: 124.93 GB
$ amd-ttm --set 27648000
[sudo] password for user: 
❌ 27648000.00 GB is greater than total system memory (124.93 GB)
$ amd-ttm --set 120
🚦 Warning: The requested value (120.00 GB) exceeds 90% of your system memory (112.44 GB).
This could cause system instability. Continue anyway? (y/n): n
🚦 Operation cancelled.
$ amd-ttm --set 112
🐧 Successfully set TTM pages limit to 29360128 pages (112.00 GB)
🐧 Configuration written to /etc/modprobe.d/ttm.conf
○ NOTE: You need to reboot for changes to take effect.
Would you like to reboot the system now? (y/n): n
$ cat /etc/modprobe.d/ttm.conf 
options ttm pages_limit=29360128

This didn’t work though, it didn’t change after reboot. I ended up changing the kernel command line myself: ttm.pages_limit=31457280 ttm.page_pool_size=31457280. This worked!

$ amd-ttm 
💻 Current TTM pages limit: 31457280 pages (120.00 GB)
💻 Total system memory: 124.93 GB

I tested loading multiple models with llama, and it worked this time! Thank you :slight_smile:

Couple of comments

  1. I could not find any documentation on the TTM business in the kernel docs. The post on the Framework forum also links to a very old kernel version (4.14), although the linked page is about TTM, it does not document either of those options! I did a bit of cursory searching in the latest kernel docs, no mention. I presume someone has to dive into the source :person_shrugging:
  2. This behaviour is still a surprising since the Vulkan backend is working without this setting.
2 Likes

I can not tell you the details about that, I learned about this when I upgraded my Framework Laptop to use with AI as well. I got it for the unified memory and was facing the same issue and read about it on the Framework forum. I think amd-ttm writes a config file and you maybe need to regenerate the initramfs.

Though I simply set the kernel parameter and it works reliably.

Okay, thanks again for the help. I’ll search around for details. If I find something, I’ll report back here.

1 Like

Hey mate, not want to be a bother but did you find up what the reason is for this new way of setting things up?

Sorry man, haven’t had the time. I have been distracted with some paperwork and other stuff.

1 Like

llama-swap needs config. I setup each model with maximum concurrent and launch the server with paralell settings to match, along with context settings per model GGUF, I can run four models on my 16GB of VRAM in llama-swap before OOM.

I suggest setting up a config with models and specific settings per model, I’m sure llama-swap is able to load up many models at once, I have a workflow cycling about 8 of them 24/7