With AI workloads, it’s not that bad. Especially in a rack server.
For gaming in a tower system where it’ll be hammered in ways that even heavy AI workloads wont, it’ll scream at 100% and even with headphones it’ll be a pain.
With AI workloads, it’s not that bad. Especially in a rack server.
For gaming in a tower system where it’ll be hammered in ways that even heavy AI workloads wont, it’ll scream at 100% and even with headphones it’ll be a pain.
Hmm. The noise makes me second guess selling the 6750 then.
Should I still go with a true 8x/8x Gen4 mobo if I want to keep the 6750 for gaming and the R9700 for AI and leave them both in?
Would one card / workload prefer the higher bandwidth? EG: if I kept the current mobo should I have R9700 in slot 1 and 6750 in slot 2, or the 6750 in slot 1 and R9700 in slot 2? Will that have a major impact?
You can reduce the power limit by 100W, sclk by 500MHz, and adjust the fan curve to 30%.
I’ve done that to my R9700 and went from being painfully loud to almost silent.
It doesn’t affect the performance of the token generation, but it will reduce prompt processing by ~15%.
Having the card sitting in a tower 1m away from me, I was more than happy to sacrifice a little performance to not have to listen to the blower fan.
According to this post, looks like they are working on block for the sapphire version:
But for the price of a waterblock + radiator etc, probably makes sense just to upgrade the mobo to have x8/x8 per card and not have to deal with water. And also having to wait for them to release it.
I am glad I asked here, because I did not know of this potential issues.
Gaming in the 16x, AI in the slower slot. Once the model is loaded in the r9700’s vRAM, it really doesn’t care as long as you are not pooling the two cards (which you can’t, they use different arch).
Get the R9700, use it for AI until the waterblock in available.
This is what I was hoping to hear.
Thank you so much for the insight.
I’ve bought MSI for my last few cards for no good reason.
Are Sapphire a good brand? They seem to be the best price atm near me.
Thanks for your help.
All 9700s are the same other than small layout choices. Buy what’s cheapest and available.
Darn, my local computer hardware shop had 2 in stock a few days ago when I started this thread. Gone now.
Guess I’ll go newegg. ASRock Creator version is cheaper than what was available locally, so I guess that’s the winner.
Incorrect. When the R9700 is running in a slot attached to the chipset, the extra hop causes 30-40% performance degradation. In real terms, it acts like a <100W GPU in that use case.
Talking to the llama.cpp devs, the explanation is that the kernel dispatch from the CPU is very sensitive to PCIE latency.
You absolutely don’t want to do this; it’s the reason I downgraded my AI server from a 13th gen Intel CPU to a 5700X and ROG Strix X570 - to get the x8/x8 split from the CPU without spending a fortune.
You don’t need an x8 slot for the R9700.
I’m using the x4 (CPU connected) slot on my X670E, and it seems to work just fine for inference workload.
As I said…it’s not the bandwidth, it’s how it’s connected. Very, very few motherboards have a CPU-connected x4 slot, to my knowledge…?
As a caveat, I only ever tried that setup with Vulkan; it’s entirely possible that it’s a different story when using ROCm (but not guaranteed). However, even if that’s the case, you’d be stuck with lower token generation rates on MoE models (which are, oddly, roughly the same performance as the Vulkan-attached-to-chipset case).
Lots of great advice in this thread. Here’s my 2 cents.
That R6750 XT supports the RDNA2 framework, making it a viable candidate for local Ai - just limited in VRAM. Our 6800 XT is pretty bad ass on models that fit within 16GB.
Don’t count it out. Especially it’s low power.
For agenetic work (agents), you will want to have a few different models loaded for different tasks, and can assign smaller Image Generation or high speed intellisense models to this RX 6750.
Yes, that mobo limits you to only a single x16 to the CPU. So the RX6750 can be placed in the Chipset x4 slot and run models with small contexts, like image generation models. Going through the chipset will destroy your PP, but for small contexts or one-shot image generators, thats not an issue unless you are loading like 10 images into the context - then you need to wait a minute. Just don’t CPU offload anything on the smaller GPU.
Also, the R9700 is a badass gaming cars too! About 30-40% faster than the 6750! So just use that - as that’s what I do (game on my r9700 when not in use).
IOW, think about gaming and playing with the r9700 and think about the r6750 for smaller support models.
More below…
Answered above by others.
However, I will say that for agenetic work, you want to stay off of the CPU (aka CPU Offloading).
If you don’t want agents and just want a Chatbot to bounce code snippets off of, then YES absolutely use CPU offloading - but ONLY on the GPU with x16 Gen4 lanes or higher which will limit you to around 15-24 TG. The trick for performance is to use --mlock --n-cpu-moe # because mlock will help you with the slow M.2 drives by loading the all the weights into DDR4 - bypassing the m.2 entirely once the model is loaded. And don’t let context spill into ram, make sure to use --ngl 999 and --fit to auto resize the context to fit all in VRAM.
Even Gen5 m.2 8GB/s is way too slow when compared to DDR4 dual channel 60 GB/s and PCIe Gen4 x16 @ 32 GB/s (64 two-way). This is why MoE models should be pushed all into DDR ram, so it pulls from Ram and not the M2 when it accesses more layers.
This could not be screamed loud enough. Get off of Ollama.
Also:
llama.cpp as you can really push the R9700 with mROC and Vulkan really hard and get significant performance boosts over vllm. Llama.cpp falls flat with simultaneous access by multiple agents (see below).vllm as you’ll see slightly less single session performance, but aggregate multiple sessions will yield much higher throughout.It’s f*cking loud. All blowers have that high pitch. Some people aren’t bothered, but absolutely most are. Especially family members all the way across the other end of the house!
You’ll hear it through the door, but not too much.
Yes, undervolt the R9700 to quiet it down until the waterblocks come out. Or, put it in the garage.
I think this is reasonable advice, but with a caveat: vLLM only really overtakes llama.cpp when you go past 4 concurrent requests. Personally, and given that I only very rarely have more than four (sub-)agents running, and I like experimenting with models without having the overhead of spending half an hour trying to find the right magic incantation to get vLLM to even launch with the latest models, I find that overall llama.cpp saves me a lot of time.
Also, with the new router mode, it has Ollama’s model-switching functionality built-in. That’s a huge benefit, and if the model you’re switching to is already in the system RAM cache, it takes seconds - as opposed to doing something similar with vLLM, in which case it’s often minutes before the model is warmed up and ready to use (at least, in my experience).
True, some llama.cpp setups can handle a couple concurrent requests. However, and I admit I may not have had all the right settings, in my experiments it fell by 50% with two concurrent inference requests, and crawled at 70% less with 3.
It may be completely hardware dependent and the R9700 may have yielded more concurrent pipelines.
My tests were on a PCIe Gen4 x8 laptop with Quadro A5500 (aka 3080 Ti) and 16GB VRAM, and DDR5 memory. With llama.cpp, I got impressive 25-30 TG with a particular model offloading to CPU (qwen3.6 35b a4 Q2) with the settings earlier mentioned (getting off the Gen4 M.2 and putting it all into DDR5).
However, as soon as I started a second chat… it fell on it’s face with the huge drop in perf to around 12-15.
I’ll admit I didn’t look into more concurrent settings for llama-serve at the time which could have helped.
True true. I’m still burning days figuring out the best vLLM settings.
I’ve been using llama-swap for a while now. Originally for it’s container isolation, but lately because it will run any command I want. llamacpp, custom llamacpp, vllm, and even batch files to kick off a docker compose setup (long story). So it’s for anything you can run on a CLI, like vllm, and gives you the ability to switch cleanly.
The router is also experimental ATM. While most of us don’t care, it should be noted you are relying on the router to mfree everything and not leave any strangling in memory when switching models. This is why I like llama-swap: it does a full PID kill, and starts a completely new PID.
I tried the router for a bit and while nice, I was limited to strictly the single llama-server binary - I could not experiment with other llama builds nor vllm. I did like the simplicity of the config file though, but llama-swap’s cfg files is simple yaml.
Another thing about llama-swap: once you do spend the time to figure out all the little flags of llama.cpp/vllm and get something good going - it’s a simple copy and paste into the llama-swap’s yaml cfg as all it does is run CLI commands.
There is one downside to llama-swap and that’s the logging, or rather how to view the results of the failed model swap. Only way I found was to curl and stream an API endpoint served up by llama-swap (not at my desktop at the moment, so can’t recall the command - ping me later if anyone wants it).
Point is: llama-server’s router is awesome if you are 100% set on your one llama-serve instance. However, for vllm, multiple llama-server builds, and even wrapping functionality within a bash script, this is where llama-swap provides all the functionality.
unfortunately, Its just getting harder and harder to ignore the benefits of CUDA, even with the associated “tax”
Unless you have lots of time to spend in extra config time.
Well, you’ve prompted me to actually test it (I haven’t tested parallel requests for quite a while). Here’s a log from four concurrent queries, based on a single-request baseline of 110t/s:
slot print_timing: id 0 | task 1372 | n_decoded = 257, tg = 59.40 t/s
slot print_timing: id 2 | task 1405 | n_decoded = 276, tg = 56.84 t/s
slot print_timing: id 1 | task 1269 | n_decoded = 420, tg = 59.95 t/s
slot print_timing: id 3 | task 1464 | n_decoded = 277, tg = 64.15 t/s
slot print_timing: id 0 | task 1372 | n_decoded = 433, tg = 59.10 t/s
slot print_timing: id 2 | task 1405 | n_decoded = 452, tg = 57.45 t/s
slot print_timing: id 1 | task 1269 | n_decoded = 596, tg = 59.49 t/s
slot print_timing: id 3 | task 1464 | n_decoded = 453, tg = 61.79 t/s
slot print_timing: id 0 | task 1372 | n_decoded = 608, tg = 58.82 t/s
slot print_timing: id 2 | task 1405 | n_decoded = 627, tg = 57.61 t/s
slot print_timing: id 1 | task 1269 | n_decoded = 771, tg = 59.16 t/s
slot print_timing: id 3 | task 1464 | n_decoded = 621, tg = 60.06 t/s
slot print_timing: id 0 | task 1372 | n_decoded = 776, tg = 58.17 t/s
slot print_timing: id 2 | task 1405 | n_decoded = 802, tg = 57.70 t/s
slot print_timing: id 1 | task 1269 | n_decoded = 938, tg = 58.49 t/s
slot print_timing: id 3 | task 1464 | n_decoded = 795, tg = 59.54 t/s
Effectively, total throughput goes to ~235t/s with four concurrent requests. To save space here, when it’s two concurrent requests it’s ~77t/s each, so 154t/s total; by way of comparison, the same model (Qwen 3.6 35B A3B Q8_0) only manages around 70t/s in vLLM on this hardware for a single request.
Worthy of note is the fact that this is not possible when running MTP - it exactly halves, then quarters the output (this is known behaviour, and something they’re going to tackle in a later refactor). I have a feeling it would look better just running ngram-mod speculative decoding, but that introduces extra variables that might muddy the waters.
I think this is the same behaviour as llama.cpp’s router mode; I don’t actually use the web view, because I prefer Open-WebUI’s extra bells and whistles. It just talks to the parent llama.cpp process, which then launches a new llama-server process for each model that’s loaded, and kills them when you request a new model and it’s already hit the maximum configured. It also lets me set up alternate model hosts (for when I can be bothered to battle with vLLM again…it happens periodically).
From there, per-model config is just a case of creating a config.ini.
I think that covers pretty much all of the functionality of llama-swap - certainly everything I need to use, at least, and it means one less moving part to manage. The only annoyance I have with it is that it’ll blindly try to load a new model whether there’s VRAM available for it or not, so if I accidentally try to load two big models into memory without restarting the service, it starts swapping to system RAM and I have to kill it. I’m sure there’s an easy way to stop it doing that, but I haven’t mustered the energy to deal with it yet.
Not criticising your approach, BTW - I just find it interesting that we’ve both approached the exact same problem from two completely different angles, and got pretty much the same results ![]()
My bottleneck was the PCIe Gen4 x8 pipe to an Ampere GPU.
I’m curious how you got so much throughput. Dual R9700s @ x16 Gen4 perhaps? ![]()
Remember, I was CPU offloading MoE due to not enough VRAM - which tanks with concurrently over a single x8 Gen4.
Yeah, I think I read in another experimental thread that MTP won’t work on CPU offloaded models?
Yeah I could have worded that better. llama-server’s router mode fires off child PIDs that it kills.
However, I believe you are still limited to that one single llama-server instance. No vLLM nor other llama-server variants.
Ditto! Yep, we have the same goals but completely different backend approaches. lol
OoenWebUI (and I’m playing with others) all point to the same OpenAi endpoint served by llama-swap.
This biggest advantage to me for llama-server’s Router mode was auto-discovery! By carefully organizing your own /models/ directory into two tiers with hardlinks (synlinks) to actual Ramalama model pulls (within the snapshots), you can have the Router discovery all the models on disk.
The biggest advantage here was the ‘auto loading’ of mmproj files for vision support.
However, I ran into several limitations and posted in GitHub issues - where they said they won’t fix or issues left open for a year. Things like when using –models-dir, it won’t pickup individual model preset.ini files with you specifying a model global preset.ini. when not specifying preset, it didn’t pickup the global models.ini even though it lived in the root models-dir as specified. And the model matching was very annoying, requiring exact case and name - wildcards are broken ATM (hint: it’s the Directory name, not filename, if you want mmproj auto load support). Argh…
It ends up creating a lot of duplicates with what was auto discovered and the params I was setting and trying to match. I have a family that uses our Ai setup, so it needs to be clean and not confusing. There was also the issue of not being able to list the same model multiple times (with matching) - you’d have to move the model out of auto-discovery path, and make up a new name for each instance in the models.ini.
Then there is the issue of hosting multiple Quants of the same model. Auto discovery would randomly pick just one or the wrong one, and match up the mmproj. Solution to that workaround was to separate each individual quant into it’s own dir - duplicating the mmproj file hardlinks!
Just a lot of edge cases that while seemed cool, didn’t work with all the bugs. Just made life annoying and I wanted to get work done.
And when I found myself manually specifying all the models and params in models.ini, I shook my head and said, “since auto discovery had so many issues requiring a full models.in, then this is no different than llama-swap, where I have to specify everything anyways - but I don’t have to format the params there as I can just copy and paste my actually CLI lines. What the point of router?”
Hence, the move back to llama-swap as the auto-discovery isn’t ready for real filtering and matching.
So yeah, I really did try router and pushed it to it’s limit. Lol
I should mention that I am still in the debugging phase of figuring out the best params to use across devices and machines. And having to shutdown llama-serve to make a single tiny change just slows me down, while llama-swap monitors the config file for changes while running. Making quick tweaks and reselecting model in Openwebui fast.
Correct. I was thinking of it solely from the standpoint of bandwidth, not latency. This is a valid callout that I was overlooking. And agree that Ollama needs to go the way of the Ender 3 (to pull for the 3D printer world), it had its place but there are much better options now.
My only pushback is, we may start sending OP down an optimization rabbit hole for a card they don’t even own yet. Decision paralysis bites hard when you are looking to drop $900+ on a card.
The r9700 is a safe bet for their constraints. The current card is acting as a bridge for gaming until a better cooling solution is available.
Nope - just two R9700s, each on PCIE 4.0 x8 from the CPU. In truth, I hadn’t really examined the numbers until prompted by this thread. I mean, I knew that in theory it would behave roughly like that from other people’s benchmarking months ago, but being totally honest…I’m almost as surprised as you are that it didn’t choke ![]()
EDIT: They’re using Vulkan, so there’s not even any peer-to-peer communication between them; it’s literally just two GPUs getting a lot of instructions from the CPU.
Yep, that’s true, if you’re using the built-in web UI. I only do that when I’m testing alternate branches (my Open-WebUI instance stays static and un-bothered by the chaos caused by testing weird PRs).
That’s fair. I tend to configure --mmproj only for the models I want to have vision capability, and…well, that’s exactly one at the moment, so it’s not a massive overhead for me. All the models I use have very specific configs for their exact use case, and I tend to create duplicates s I can test individual features, eg MTP on/off etc.
Excuse me, this is the Internet. You’re supposed to argue your position until the mods get involved. Please rectify.
(apologies for the bluntness of my original response, BTW - that’s something I tend to only spot on re-reading after I’ve accidentally offended someone)