Hello!
I am working on building two machines, one for AI, one for a gaming server. With RAM being super expensive, I am trying to get the most bang for my buck, and I do not want to overspend on hardware that I ultimately do not need. For example, here are two kits I am looking at.
Kit 1:
Kit 2:
What I do not understand, and ultimately what I am trying to figure out, is how these affect real-world performance and if there is enough of a difference in these two kits to meaningfully affect the performance in any significant way at the end of the day that would justify the $230.
My use cases are going to be an Ark Survival Ascended server running 2-4 maps. The second one is going to run local AI. Local AI will be used for coding, having code examined for errors or room for improvement, research, inputting research papers and articles to ask questions about them, and image/video creation. That system will be using a 9950x3d and a 5090. I was not going to build this system anytime soon, but I won a 5090 in a raffle, and now I seem to have little to no reason NOT to do it.
So tech gods, am I good to save some cash?
Bonus question: If I purchase two of the same kits for a total of 128GB for a single system, it should work with no issues, correct?
Local AI will not benefit that much from faster system RAM if you’re running a dedicated GPU. A PCIe 5.0 x16 slot still has significantly lower throughput than DDR5, so shared VRAM bandwidth is going to be bottlenecked by that rather than the RAM itself.
1 Like
You won’t be able to use XMP, but it’ll usually work just fine. I’m currently running 2x DDR5 6000MHz CL36 2x16GB kits of the same SKU at 5200MHz (will not POST beyond that).
If it’s DDR4, you can even run mismatched kits at JEDEC speeds (limited to the slowest kit).
1 Like
Can you go deeper into this? I assume that even if I am bottlenecked by PCIe speeds, it is still pretty darn quick?
Be warned, this ended up much longer than I had planned lol.
The numbers are something along the lines of 80GB/s bidirectional throughput (i.e. real-world bandwidth once you account for overhead etc) for PCIe 5.0 IIRC. That’s roughly equivalent to CL36 DDR5 running at 4800MT/s in 2 channels.
For context, the memory bandwidth on a stock RX 9060XT 16GB is about 320GB/s, and that’s on the low end. On a 5090, you’re now looking at 1.8TB/s membw.
I’m not entirely sure about the following so someone may have to correct me, but:
If you have a multi-GPU setup the memory speeds may potentially have a greater effect since the system memory BW will essentially be split across the slots when it’s feeding all of them with data simultaneously.
That is, however, assuming that you are only supplying them with data from system RAM. Normally in LLM workloads with multiple GPUs, much of that traffic is going to be between the GPUs instead, and also getting bottlenecked by going through the CPU first if you don’t have RDMA enabled.
Anyway, I wouldn’t worry too much about it. For inference it isn’t really all that important, it’s when you get to multi-GPU training that it starts having an impact.
Bottom line is that maximizing your available VRAM will provide much more tangible benefits than the speed of your system RAM, as much less data will need to be moved off of the device if your entire model can fit within VRAM.
If you want an analogy for the difference between RAM and VRAM, imagine the VRAM as a freight train for bulk transport, while the system RAM is more akin to a small transport van used for quicker transports of smaller items. The PCIe bottleneck here is basically the act of transferring the wares from the freight train to the van.
1 Like
That helps. I wanted to do a threadripper build with some RTX 6000 but my wife put a budget on me and RAM just had to go and get expensive. So that is the long term goal, but right now, I will be happy with what I have. Thanks a lot!
1 Like