I’ve been tuning a Strix Halo box as my long-context local LLM tier and figured the setup might be useful to others running similar gear, or anyone weighing whether to pick one up.
Headline: Qwen3.6-35B-A3B UD-Q8_K_XL through llama.cpp, serving around 25 tok/s with a live observed context of 153,562 tokens. The screencast shows the run sitting under 100W. This is real-world usage, as it’s the current daily-driver chat/high-reasoning tier for my homelab.
Full writeup, screencast embedded, plus an interactive loading/quantization artifact:
Full post: https://kmarble.dev/posts/strix-halo-llm-inference-show-and-tell/
Screencast: https://assets.kmarble.dev/artifacts/strix-halo-llm-inference-loading/strix-halo-llm-inference-screencast-2026-04-26.webm
Interactive: https://kmarble.dev/artifacts/strix-halo-llm-inference-loading/
Download: https://assets.kmarble.dev/downloads/strix-halo-llm-inference-loading-2026-04-26.zip
Hardware/software:
Host: artemis
Machine: Minisforum MS-S1 MAX
APU: AMD Ryzen AI Max+ 395 w/ Radeon 8060S
RAM: 128GB LPDDR5X unified, 124 GiB visible
OS: CachyOS, Linux 7.0.0-1-cachyos
Backend: Vulkan RADV
Mesa: 26.0.5-arch2.2
llama.cpp: b8890, 8bccdbbff
llama-swap: v205, 66639e83f7be4f1354817e45321f50bdb8e3227d
Live llama.cpp metrics at capture:
n_tokens_max 153562
prompt_tokens_seconds 322.685
predicted_tokens_seconds 24.0673
requests_processing 0
requests_deferred 0
Current model launch, shortened:
llama-server
-ngl 999
-fa on
--no-mmap
-t 16 -tb 32
-m Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf
-c 524288
-np 3
--kv-unified
--cache-type-k q8_0
--cache-type-v q8_0
-b 4096
-ub 4096
--cache-ram 16384
--cache-idle-slots
--slot-prompt-similarity 0.8
--mlock
--reasoning on
--reasoning-budget 65536
The RADV APU heap fix is important:
<driconf>
<device driver="radv">
<application name="Default">
<option name="radv_enable_unified_heap_on_apu" value="true"/>
</application>
</device>
</driconf>
Strix Halo will obviously never beat something like a 5090, but it doesn’t have to. It’s the complete envelope: large unified memory, low enough power, no separate VRAM ceiling, and enough generation speed for an interactive local chat/agent loop.
This sits behind llama-swap as an OpenAI-compatible endpoint. My main homelab host, voyager, has a 7900 XTX and runs a smaller Qwen3.6-27B cron tier. artemis is the high-context chat/reasoning box.
The main lesson so far: long context isn’t just a model-card number. You need the model architecture, quant, KV cache settings, backend, memory behavior, and service layer to line up. In this case, Qwen3.6-35B-A3B plus Strix Halo unified memory plus llama.cpp/Vulkan/RADV is a genuinely useful local inference target.
If you’re running similar gear, I’d love to hear what tok/s and context numbers you’re hitting, or any tweaks that pushed yours further. UD-Q3_K_XL is on my list for the next round of testing if anyone’s already been there. Happy to dig into any of the launch flags above if something looks off, too.

