My home server has an MSI X470 or X570 main board, Ryzen 5700G, 128GB DDR4 3200 (of which 80GB is available to the VM), and an 8gb RTX 2070 Super. I don’t have any free PCIe slots, they’re all either used or blocked. I could free up the second X16 slot by replacing the 3x 1TB SSDs with a single 4TB drive, if it makes sense to spend the ~$600 to do so.
I’ve recently started messing with unsloth and comfy, but I’m getting held back by models not being able to load into the available ram/vram. If I were to invest in a 7900XTX, B70, or R9700 Pro to replace the 2070, would I then be able to run models like Qwen 3.8 reasonably, or would I just run headlong into the next bottleneck, presumably the DDR4 speed?
If this is workable, what GPU is the best bang for the buck? The 24gb 7900XTX and 32gb B70 are around $1300, the R9700 Pro 32gb is $1700. The 4090 and 5090 are well outside by budget, it would make more sense to get an a6000. Used 3090s on eBay are prices starting around $1400, that seems like a bad deal.
As much as I want to try this out at home, I don’t have the kind of money to let me justify spending $2000+ on 128GB DDR5, let alone the motherboard and CPU. So if this is all hopeless, I’d like to know now so I can save my money for whatever is available in a few years, after memory and GPU prices return to more reasonable levels.
Unless you’re going 8+ channels, DDR5 is not going to be able to compete with VRAM. (and even 8 channels is on the slower side in terms of usability)
From what I’ve read, support is the defining characteristic for 32GB cards. nVidia is king, requiring almost no fiddling to get things working, the software on AMD is good enough that you can run almost anything with some problem solving, but the intel is a headache because the software only runs with specific versions of model runners, if at all. If you go multi-card the problems increase to the n-th degree. (look up the RCCL and BattleMatrix problems on this forum)
24GB works, but with limited context size. For me this means limited usefulness. So I’d say that 32GB is realistically the theshold for any useful LLM use, but that might a kind of snobism that I have in me.
First off, note that local LLM inference speeds is directly related to one measurement: memory bandwidth speeds. Your DDR4 AM4 system is only 2 channels, and that’s capped at ~50 GB/s. But you biggest bottleneck for CPU offload are the PCIe speeds (15 GB/s Gen3, 30 GB/s Gen4).
As a comparison, the AMD R9700 VRAM clocks in at over 600 GB/s. An old 3090 has over 900 GB/s and is still highly viable today and a very good choice.
IOW, an Unsloth Qwen3.6 35B A3B model will run around 15-20 tokens/s with a 8GB offload to GPU VRAM - until you grow the context and it will drop like a rock. Slap an R9700 in it, fit it all in the 32GB VRAM, and you’ll see 120-180 tokens/s. That’s the insane difference we’re talking about with an X470 DDR4 system vs “pure VRAM.” Not to mention parallelism of 4x agents or more on a GPU that can increase those speeds.
Next, note that the ideal local “coding" models are around 20-50 GB in memory size. Meaning, you need 24 GB of VRAM (because you need a context window) to 60 GB. There are much larger models, but you need 90-200 GB of VRAM to make them useful.
As I mentioned above, forget about CPU/DDR offloading. It’s not just the DDR4 dual channel limits, but PCIe Gen3 @ x16 is capped at 15 GB/s in the X470. PCIe Gen4 @ x16 is 30 GB/s in the X570.
And the prices are never going to return to where they were. There will be. 20-50% price drop when the bubble bursts, but that doesn’t mean much when things have already jumped 12-15x in price (yes, 1500% increase) - and it’s set to double again in 2027.
Don’t wait for the moment… The moment is gone and won’t ever return. Personally, I will only buy used or surplus going forward. I stopped supporting Nvidia 15 years ago, and AMD recently with their politics.
The rule as of mid 2026 is: you need 24 GB of VRAM to fit all of a model and its context into VRAM, at a minimum. The more VRAM you get from here, the larger context window and more accurate of models you get.
PCIe speeds don’t matter much, so you can run it on a potato - or your X470. Only prompt prompt processing will take a hit which isn’t a big deal unless your a hardcore coder running a half dozen agents. And even then, it just slows things down a little which is reasonable while you grab a coffee.
IMO, go for the largest amount of VRAM you can afford - e.g. the R9700. It will run just fine on the X470. I got mine used on eBay for $1100-$1200 (just got a second one a few weeks ago - after waiting months for a good deal). My ASRock Rack X470 has dual x8 slots in running them in, before moving to my Threadripper 2950X system (dual Gen3 @ x16).
With all that said, you already have an 8GB Nvidia GPU. So take whatever speeds you are seeing, and double it. If that sounds acceptable, then continue reading below.
There is some great work being done for CPU offloading to make better use of low-end machines with GPUs of 8GB of VRAM, and goes up from there - especially for systems with lots of DRAM. It’s still new and is only for Nvidia right now (AMD coming later).
It’s called FreeToken. It’s a new LLM Engine (replaces llama.cpp/vLLM).
But note that it’s still just a barely usable system - instead of Qwen running at 20 t/s, it may now run at 30 t/s. And then slows down as you grow the context window.
If you can afford to move to 24GB or 32GB VRAM GPU, do that instead as it will be well over 10x faster, if not more in parallel.
FreeToken is targeting systems that can’t run 24 GB of VRAM, to maximize all the bandwidth efficiently for the absolute fastest the system can perform. Workstations are outside affordability for this context.
It’d be relevant to know if that 2nd slot is connected to the chipset (so limited to x4 lanes), or to the CPU (meaning x8 lanes).
Nonetheless, that 5700G would limit the PCIe Gen to 3.0.
Qwen 3.8 27b is a dense model, so ideally you want all of it on VRAM, otherwise your speeds would go down the drain if you were to offload parts of it into regular RAM. If you can fit it entirelly in VRAM, then your RAM speed would be almost insignificant.
DRAM speeds become relevant when you do hybrid offloading, and that’s a good strategy for MoE models that only have a small part of weights active at once, meaning that even at regular RAM speeds you can have good enough speeds for it to be usable.
With that said, any 24GB+ GPU would run Qwen 3.8 27b reasonably with the right quants. With Q4 you should be able to have ~100k of context.
On the other hand, Qwen 3.8 Flash has just arrived. If you were able to dedicate more RAM to your VM, you could likely try it out since it “only” has 6B active weights:
3090s are usually the best bang for the buck, but those prices are absurd. The 3090 will be faster than the 7900xtx, so I don’t think it’s worth it.
The B70 has more VRAM but is considerably slower than the 3090 and has more issues software-wise.
The R9700 pro is also slower than the 3090, and not as good sw-wise as the green option, but it’s not as bad as Intel and has 32GB of VRAM as well.
I’d try to find a 3090 closer to $1000 or less than that, or a R9700 if that’s not feasible.
I don’t think it’s worth it to change your platform unless you were to actually upgrade your RAM amount. That’s what I did when I went from my 5950x+128GB into my current 9950x+256GB, but at the current prices this makes no sense and you’d be better off buying a more expensive GPU instead.
Whoa…I just increased from Q2 to Q4 bit model quality, doubled my context from 131K to 262K, doubled my context precision from Q8_0 to full FP16, while more than doubling my original NVFP4 CPU Offloading speeds on my 16GB laptop (went from 20-25 t/s to 45-60 t/s) for a 20 GB model!
I would purchase an R9700, look around, and you should be able to find one for $1,400 in the U.S. They are by far the best thing for the buck in terms of having great performance, modern architecture, and large VRAM for a reasonable price. You can definitely run Q4 over the famous 27B model and get a lot done.