B70 or R9700? (help decide)

Ah, that’s it. You can fit it all into VRAM - whereas I am offloading a 35B model with 4B active layers into only 16GB VRAM.

The OP mentioned whichllm models on a small GPU, so MoE offloading still applies for their reading pleasure.

So technically, with 4 concurrent processes, I think that’s basically two requests per GPU with a pair of R9700s - you lucky son of a … (where’s the mods?)

If you wanted to replicate my setup, use --cpu-moe to offload the experts to your CPU over PCIe Gen4 @ x8.

The laptop is also dual channel DDR5 to help here.

Weirdly, it’s not - there’s barely any increase in GPU usage showing in nvtop with four requests relative to just one. Maybe it’s just less idle time in the pipeline?

Basically, I have no idea how or why it’s working like that, but I’m not complaining that it does.

The token processing is still sequential to my knowledge, because I’m using -sm layer. If it was -sm tensor, then the work would be more properly distributed; however, that actually slows everything down under Vulkan because they haven’t optimised it for non-CUDA/ROCm yet (and there’s no word on when/if they will), and ROCm still sucks for token generation with MoE models.

yes, I already run some local models as explained earlier in this thread. but I do feel slightly limited by the vram. hence the desire to upgrade. which is why I made this thread to ask for advice.

I don’t think I really care about agentic work. At least not at this time. I’m fine programming on my own as I have for many years without LLM’s, and I am fine with deliberately going to the prompt with pasted snippets and questions when and if I need.

I have experimented with copilot IDE (Visual Studio) integration at work, and it’s intellisense suggestions were so inconsistent and bad (suggesting nonsense) that I ended up disabling it in favor of regular native intellisense which is good enough autocomplete for my needs. And then I copy paste functions and stack traces into a separate copilot prompt when I need. But that is at work, so slightly different context and usage compared to home use which is what this thread is about.

Quite frankly I like the separation IDE from LLM, as so far my experience with integration has been counterproductive.

But perhaps I can revisit this once I have play around with the setup and models and settings. etc…

Problem with gaming on the R9700 is noise as others have mentioned. My current setup is very quiet and my old setup was rather loud so I do value the silence. And even you have mentioned:

So I am not really interested in gaming on it if it will be this loud. Hence the question of do I go x8/x8 to run both reasonably, with gaming only on the 6750, and LLM only on the 9700.

I do not currently own a garage, and running a cable long enough to separate monitors from the loud machine is an entirely different problem. Maybe involving optical. Not a can of worms to open today.

I have no problem picking up an x8/x8 board if this is a real problem. Or suffering the noise of the 9700 while gaming till cooling is better.

I am sure glad I started this thread, as clearly it is a little more complex than I thought as I mainly wasn’t aware of the PCI-e limitations initially.

Indeed… I still have some reading to do.

Thank you all for your valuable advice.

2 Likes

( ͡° ͜ʖ ͡°)

2 Likes

I don’t think you have read this thread if you are still suggesting nvidia at this point.

I have already had plenty of success with ollama (even tho I should clearly shift to llama.cpp, will play around this weekend when time find). So far I have had no problem running local models without CUDA. I haven’t spent lots of config time, it pretty much “just works” for my use case.

My limitation is VRAM (and PCI-e lanes). Not CUDA.

And AMD drivers on Linux are objectively better than nvidia.

I’d also argue that The gap between a R9700 and other RTX options with the same vRAM is closing when you look at it purely from $/GB of vRAM. Yes, a 5090 is objectively better with more than double the memory bandwidth and mature software stack in Linux and Windows, but for the price of one 5090 you can get two R9700s.

The other reason is between better drivers, llama support, and tuning for your use case you will not really miss that extra Tok/s when you can either run larger models/context for the same cost and arguably get more done.

2 Likes

For the price of one 5090, I could get two R9700s, AND an x8x8 mobo to host both more properly than the current mobo, and still have a few pennies left over.

1 Like

This is is exactly my problem. I find myself wanting (not needing, mind) more raw performance, as much as dual R9700s is already pretty baller in absolute terms…there’s so much of a price leap to get to the next rung on the ladder that I simply can’t bring myself to do it.

Hell, even a pair of 5090s doesn’t really count as that next step, because of all the cooling problems given Nvidia’s rule that consumer cards can’t be 2-slot to force people to their pro cards for AI. That means that the next step in CUDA-land is, really, the Blackwell RTX 6000 Pro…priced at approximately £divorce.

With that said…there’s one interesting possibility: the blower-style 3090 cards from China. There seems to be a glut of them on eBay at the moment, and while they’re not as cheap as they were…48GB total VRAM is pretty sufficient for most models at the moment, and they’re faster than the R9700s. Interestingly, when I had a pair of non-blower 3090s, I noticed that they didn’t suffer from the x4-from-the-chipset problem when running llama.cpp under CUDA. They don’t represent an upgrade for me (performance, yes, but a downgrade in capability from 64GB), but they might be a good place to start for you?

1 Like

I guess my final question would be, is how much impact is there going from x16 to x8, for either card?

Do games / LLMs benefit from the full x16 over x8, and by how much?

Games: not really, no.

LLMs: Sometimes, it depends.

Yeah, that’s a crappy answer, but it really is quite complex. If you’re using tensor parallelism, and it’s well-optimised, and you’re using a dense model, then…yeah, you can expect a bit of a performance delta between x16 and x8. In practice, if you’re using the B70 or R9700, then layer split with llama.cpp on Vulkan x8 is probably going to be faster than tensor split on SYCL or ROCm at x16, whether you’re using vLLM or llama.cpp…as long as you don’t go above 4 parallel requests on a regular basis.

In other words, don’t worry about it, because it probably won’t be an issue. Maybe.

2 Likes

Is water-cooling not a viable option? I know for production you don’t want to play it fast and loose (LTT…), but with thoughtful planning and running two pumps in series for redundancy, you can have a viable solution with alerts and should be leak free as long as you do yearly replacements of hoses and o-rings.

Ok, so the only board that is in stock at my local that apparently supports x8/x8 according to the support guy on the phone is the

MPG B550 GAMING PLUS

Just going through the info trying to validate this. MSi’s website says:

This looks like only x4, so I don’t think is correct?

I’ve had two liquid-cooling failures in my three decades of this hobby/career, and both were painfully expensive. As far as I’m concerned, when dealing with anything expensive that’s likely to be running for a significant portion of every day - especially when it’s critical - a big honkin’ lump of metal is by far the safest approach.

All depends on your tolerance for risk, I suppose. And maybe whether you can find somebody at the insurance company who’ll take back-handers.

Correct.

You really need to be carful with finding a AM4 (even AM5) MB that does true 8x/8x.

The best place is eBay if you can find a good seller. And it’ll largely be x570 boards. X570 Aorus Pro Wifi being a prime example, but I think the AsRock Taichi is another.

DS is correct. Games don’t really benefit at all from x16 to x8, unless it’s PCIe 3.0 as the bandwidth for 4k can saturate an x8 Gen3 bus IIRC. GN did a video about it years ago, going all the way down from x16, x8, x4, x2, and x1. IOW, “not really, no.”

LLMs, yes, usually a little bit in my opinion. Most noticeably is prompt processing I’ve seen with the same GPU drop about 30% PP in pcie3 and pci4. Inference token gen - a bit as well.

The most sensitive issues with PCIe speeds are:

  • Communications from multiple GPUs to either other. If you stick to a single GPU, not really an issue.

  • CPU offloading for when you want to run a really large MoE model and offload most layers to system ram. On a dual-channel DDR4 system, you’ll hit memory bottleneck first though. IOW, this is very slow anyways on an AM4 system and x16 to x8 Gen4 doesn’t matter at all for MoE offloading in my testing. Basically, Gen4 x8 is pretty close to DDR4 3200 in dual channel.

    Note that I’m working on a customers AM5 right now and this is a different story with faster 6400 ram on dual channels…. Yes, Gen4 x8 could be a bottleneck in AM5. Still testing.

  • Prompt processing, or aka, “how fast can you upload 300MB of PDFs into the model and process them (OCR, image to text, etc).”

    A simple chatbot is not really stressing PP.
    A long coding session with 200,000 lines of code, really stresses PP.

    This PP number is most critical in Agents and/or longtail coding sessions that build up huge contextes.

Training, I haven’t tried training over pcie3 but I’d suspect training will take a nose dive.

I vote for ASRock x570 Tachi mobos if you can find one on ebay. Otherwise, whatever else has x8+x8 split.

3 Likes

Would you happen to have some numbers on how these changes affected compute and rendering/gaming performance? I would likely “have to” use the card both for gaming and for inference, so this info would be super useful for me.

Thank you for all the super valuable information everyone, great thread! <3

I use the card for both, but I don’t play many games. Currently, I only play Warframe, and the card if fairly quite without any reduction to the power level, it might not be the same for more demanding AAA games.

I can’t test what the different would be in a game, I’m using two VMs one for games and one for AI. They are using different amdgpu drivers. It’s only the AI VM that that has the dkms driver that can undervolt, and it’s only the steam VM that can use a connected display.

This has been my biggest restraint from buying the next GPU for my system. Custom waterblocks are slim to none. I was eyeing a tesla with a waterblock already on it strictly for the saving of time and money.

As someone who slaps the PCIe bandwidth walls of the RVII daily, LLM’s are like 5k on the bus

I actually do training on the RVII but custom wrote the framework for the card. Stil a WIP and not on existing models, but the work can be done, and having 16GB @1tb/s is killer for speeds, just slow for large sized models. Trick is never use the bus, do the work on the card, entirely.

1 Like

Submitted a bug report for it not detecting the RAM and BW:

Using this PR

the output is now:

AA Index fetch failed, will use fallback: __NEXT_DATA__ payload not found

╭──────────────────────────────────────────────────────────────────────── Hardware Info ─────────────────────────────────────────────────────────────────────────╮
│ GPU 0: Navi 22 [Radeon RX 6700/6700 XT/6750 XT / 6800M/6850M XT] — 12.0 GB — BW: 320 GB/s                                                                      │
│ CPU: AMD Ryzen 9 5900X 12-Core Processor — 12 cores (AVX2)                                                                                                     │
│ RAM: 125.7 GB                                                                                                                                                  │
│ Disk free: 802.0 GB                                                                                                                                            │
│ OS: linux                                                                                                                                                      │
╰────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

                                               Recommended Models                                               
┏━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━┳━━━━━━━┳━━━━━━━━━━┓
┃   # ┃ Model                                    ┃ Params ┃ Quant  ┃ Published  ┃ Downloads ┃ Score ┃ License  ┃
┡━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━╇━━━━━━━╇━━━━━━━━━━┩
│   1 │ openai/gpt-oss-20b                       │  21.5B │ Q3_K_M │ 2025-08-04 │      8.2M │  74.4 │ apache-… │
│     │                                          │ (3.6B… │        │            │           │       │          │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   2 │ Qwen/Qwen3-14B                           │  14.8B │ Q5_K_M │ 2025-04-27 │      2.2M │  72.9 │ apache-… │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   3 │ microsoft/phi-4                          │  14.7B │ Q5_K_M │ 2024-12-11 │    891.5K │  70.5 │ mit      │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   4 │ Qwen/Qwen3-8B                            │   8.2B │ Q5_K_M │ 2025-04-27 │     12.8M │  64.9 │ apache-… │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   5 │ deepseek-ai/DeepSeek-R1-Distill-Qwen-14B │  14.8B │ Q5_K_M │ 2025-01-20 │    585.0K │  60.1 │ mit      │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   6 │ google/gemma-3-12b-it                    │  12.2B │ Q5_K_M │ 2025-03-01 │      2.8M │  59.6 │ gemma    │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   7 │ Qwen/Qwen2.5-14B-Instruct                │  14.8B │ Q5_K_M │ 2024-09-16 │      2.4M │  54.9 │ apache-… │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   8 │ deepseek-ai/DeepSeek-R1-Distill-Llama-8B │   8.0B │ Q5_K_M │ 2025-01-20 │    416.0K │  54.8 │ mit      │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│   9 │ Qwen/Qwen3-Next-80B-A3B-Instruct         │  81.3B │ Q5_K_M │ 2025-09-09 │    382.0K │  66.7 │ apache-… │
│     │                                          │ (3.0B… │        │            │           │       │          │
├─────┼──────────────────────────────────────────┼────────┼────────┼────────────┼───────────┼───────┼──────────┤
│  10 │ openai/gpt-oss-120b                      │ 120.4B │ Q5_K_M │ 2025-08-04 │      5.1M │  65.9 │ apache-… │
│     │                                          │ (5.1B… │        │            │           │       │          │
└─────┴──────────────────────────────────────────┴────────┴────────┴────────────┴───────────┴───────┴──────────┘
  Top pick confidence: Medium (direct benchmark, gap +1.5)
  Benchmark reference: 2026-05 curated snapshot; live AA / LiveBench / Aider merged when reachable.

Which seems a little more accurate.