Advice on a local AI server before buying

After seeing how good the cloud models are getting, I feel like this is something we cannot let big tech hold over us so deiced to build a budget local box.

After testing about 20 open models, Qwen3.8-Flash-Next (medium reasoning) was the only one that passed my task without inventing config options when used with a harness that forced doc lookups. So the box is built around that model. GLM-5.3-Flash performed even better but it’s too big for my budget.
Planned build (Netherlands prices):

  • Ryzen 9 9900X, about €300, plus a tower cooler
  • MSI B850 Gaming Plus MAX WiFi, about €170
  • 2×48 GB DDR5-5600, €1,199–1,549. Two sticks only to avoid the four-stick speed penalty.
  • Used RTX 3090, about €1,150–1,500, bought with buyer protection so I can test it in the build before the money is released
  • Case, 850 W PSU and NVMe I already own

That’s about 79 GB of the model in RAM (experts plus the 28.8 GB n-gram table) and about 20.6 GB on the card.
I first had a Ryzen 5 9600 in there, but someone pointed out that the single CCD chips top out around 60 GB/s memory reads, and the speeds I measured were on a two CCD chip. So the 9900X it is.

Questions:

  1. My plan is to run llama-server across both CCDs and pin the background stuff (Syncthing, a sandbox container, an embedding model) to two or three cores. Does that make sense, or is it better to give the model one CCD?
  2. At 96 GB with about 79 GB of mode will 17 GB be enough for the OS, a sandbox container and an embedding model? Should I use --mlock?
  3. Has anyone run Unsloth’s MTP branch with experts on the CPU? What speedup did you get, and does it break the prompt cache on the DeltaNet layers?
  4. Has anyone tried Strata on a 3090 with the UD-IQ4_XS file? On Strix Halo they get 54 tok/s, which makes me doubt the GPU route a bit. The reason I’d still go GPU is adding a second 3090 later.
  5. The second card would go in a chipset PCIe 4.0 x4 slot. Is that fine for splitting layers, or do I need an x8/x8 board?

Super exciting to work on it but I am really inexperienced so this would be my first build. Does it make sense?

Can you stretch the budget to go for the R9700 instead, if you drop to a used Ryzen 5x00 and an X570 board that’ll do x8/x8?

For what it’s worth, the radiance engine is currently running Qwen 3.8 Flash Next on my dual R9700s at ~200t/s decode, 6200t/s prefill (4-bit, with FP8 ngram table and activations, I think). I don’t know how it’d run with just one (I’m using mine right now, so I can’t check), but I’d hazard a guess that it’s significantly faster than a 3090.

BTW, yes…if you want to go for a second card and TP, you’re going to want x8/x8. Layer split doesn’t need that, but at the very least you’re leaving a lot of performance on the table. With R9700s and one in a chipset slot, that card lost ~35-40% performance from the extra hop of latency.

Also…with the advent of all these architecture-specific inference engines, llama.cpp is currently one of the slowest ways to run models (also incredibly limited, given its problems with spec decoding and concurrency).

Has anyone tried Strata on a 3090 with the UD-IQ4_XS file? On Strix Halo they get 54 tok/s, which makes me doubt the GPU route a bit. The reason I’d still go GPU is adding a second 3090 later.

I found the benchmark result. it seems when context size is 200000 (I think this is comfortably large for coding agent), prompt processing is 2000 tok/s and token generation is 60 tokens/sec. unsloth Qwen3.8-Flash-Next 125B UD-IQ4_XS, AMD Ryzen 7 5700X, DDR4-3200 32GB x4, NVIDIA RTX 3090 x1

Also…with the advent of all these architecture-specific inference engines, llama.cpp is currently one of the slowest ways to run models (also incredibly limited, given its problems with spec decoding and concurrency).

I agree 100%