I’ve been tinkering with locally-run LLMs in recent months and have a question that has come up in my adventures. My workstation is a Xeon W system with two RTX 3090 GPUs currently running Ollama with Open-WebUI front end. Models run well on the GPUs and, to my surprise, small models also run serviceably well (about 10 to 12 tokens/second) on the CPU. That is, unless I run both simultaneously, in which case both are horrendously slow.
My question is what specifically is the bottleneck I’m experiencing here? Am I saturating PCIE lanes? Is the GPU workload still flowing through the CPU throughout the processing? Software limitation? Is there a possibility of removing this bottleneck?
It’s not a huge problem and I suspect it’s just a hardware limitation, but I can see use cases where having LLMs running in parallel would be valuable.
I’ll just close with a big thanks to the members of this forum and the L1 crew, from whom I’ve learned a ton in recent years! Thank you! You’re an awesome community!
In short, besides the ethics issue, there’s the issue of poor performance when they dropped llama.cpp.
I personally don’t have much experience with dual GPUs, so I’ll have to defer to the experts.
But the 3090s should be kick ass with the following models:
Qwen3.6 27B Q4 should fit on a single 3090. Q8 will easily fit if you split the layers (not tensors). With llama.cpp’s Cuda, it would yield around 70-100 t/s.
Qwen3.6 35B A3B Q8 can be split across the two GPUs. I think I’ve seen 140 t/s from others posting.
You should always target models that will 100% fit in VRAM. So 24+24 = 48G , which is huge. You’ll want to target 27B to 35B models at rich Q8 quants. Anything larger and you’ll be CPU-offloading - which would bring your speeds down to those 10t/s. The bottleneck being the slow DRAM system memory.
So most likely it’s a configuration thing if these are the sizes you are using. If you are trying to run a 70B or larger model, uh, don’t if you want speed. Besides, Qwen3.6 27/35B models meet or exceed 70B and larger models now from last year.
Wendell had mentioned “ComfyUI” in a recent video, so I’ll suggest that to get started. It uses llama.cpp under the hood.
Wow, what a great read. Thank you for sharing that article. I had read several vague comments alluding to some of what it discusses, but definitely wasn’t aware of the whole story. I also wasn’t aware of any of the technical issues coming out of the Ollama system either. It really puts things in a new context. I have actually been planning to give llama.cpp a spin sometime soon, as it’s often cited in instructions on using Hermes and Turnstone (thanks, @wendell), both of which I really want to try.
On the Qwen models, I’m with you there as well. I keep hearing that Qwen 3.6 works really well with the harness systems. I take it from your comment that Qwen 3.6 models would likely outperform Deepseek-R1? For my purposes, I’m mostly going to be looking for quality over speed.
With this convo, I’ll have to try out the llama.cpp platform in my weird GPU and CPU in parallel use case to see if it functions any differently. I’ll also look into that ComfyUI.
Since you mentioned Deepseek-R1 and CPU and GPU parallel, I’m going to stop you right there. I’ve read a half-dozen other posters say the same thing: “I’m ok with slow, I want quality over speed.” I said the same thing last year when I got started. However eventually everyone comes around to realize: dumping 1000W+ of heat and noise to crawl at 10-15 t/s is absolutely not worth it with a system like yours.
Those that don’t care usually dump the GPUs and PC and move to a DGX Spark or a Strix Halo platform. Their unified memory approach let’s you run large models a little faster because of their 270 GB/s (Spark) or 200 GB/s (Strix Halo). More about these numbers below…
But in short, when you have monster GPUs in a PC slot setup like yours (doesn’t matter if you are on AM4 or a pimped out sTRX platform), you absolutely want to avoid any CPU/DDR offloading. You want to only run models that fit in 48GB of VRAM. Those statements will save you months of frustration.
So you are at a fork in the road: can you drop $3000-$5000 on a new Spark or Strix Halo build? If yes, go that route (and I’ll buy your 3090s hehe). If no, then read below as to why you need to reset expectations…
You asked what could be the bottleneck in your OP - the answer is “Inference speed is directly proportional to Bandwidth speeds.”
You didn’t mention your hardware, so let me throw some numbers to skip ahead as to why you’re looking at the wrong models to run.
Now with all that said, your LLM inference speeds will be bottlenecked by the lowest common denominator listed above.
That’s you’re bottleneck. Now compared that speed with nearly 1000 GB/s if you just stick to what fits within the GPU 48GB VRAM. That’s 67x faster than, say, PCIe 4.0 x8.
GPU to GPU communications isn’t effected much over PCIe links when talking LLM Inference. Usually only the PP is effected a bit, and depending on if you are running agents or not.
So the bottom line is if you are sticking to your 3090 GPUs, you’ll want to set expectations and move to what fits all into 48 GB of VRAM across the two cards. With that mindset, it’s pretty easy to get setup with ComfyUI for a couple of GPUs for insane performance with Qwen3.6 35B Q8 GGUF.
*If you have a Mobo with 4 memory slots, this doesn’t exactly mean it has quad channel. Only workstation and server platforms have true 4/6/8/12 channels to the CPU. Consumer platforms are all dual channels only.
Doh, I should have asked.. what do you want from your local Ai?
Coding assistance? Then what I said above is the correct approach. Qwen is currently the best coding model that ranks right up there within 3-8% of quality benchmarks, is ultra fast, and has great PP with the right setup for agents in parallel.
Or maybe you want a local family chatbot, kitchen voice assistant for cooking, anonymous search*? If yes, and you want a deep reasoning model like DeepSeek V3 or GPT or some other larger general purpose/reasoning model? For those, you want a different system like the DGX Spark or AMD Strix Halo platforms to handle 100+ GB models.
Personally, my family switches between Qwen3.6 3.5B and Gemma4 27B models for their chatbot/kitchen assistant. Mostly because I only have a single P40 I dedicate for those. It’s tied to SearchXNG which basically searches the web for them. So they don’t need ultra large models for a good chat and experience, as the search results usually keep the facts straight.
Yeah, at the moment, I’m primarily interested in coding assistance (I’m very much a noob at coding). I hadn’t been aware that PCIE bandwidth was so important here. I could’ve sworn than I had read early on that it explicitly wasn’t. I’ll look back and try to find what I read on that way back when.
As for my workstation, here’s the goods:
Dell 5820
Intel Xeon W-2155 (10-core)
128 GB ECC DDR4 RAM in quad channel (as reported by the BIOS)
2x Nvidia RTX 3090 (turbo blowers), 48 GB VRAM total, both are seated in PCIE 3.0 X16 slots running at X16
System is a virtual machine running on Proxmox with GPU passthrough - I’m sure there some additional performance hit there a well
As far as I am aware, the Deepseek-R1 model I run is the 70B model that is 43 GB, which should fit in VRAM across both cards. There doesn’t seem to be any overlflow to system RAM and nvtop reports less than 100% VRAM usage. When running two models on parallel, Ollama reports 100% GPU for one model and 100% CPU for the other.
In short, it’s time to overhaul ALL the software and more tinkering. I might at some point consider a beefier system but wanted a, um, “cost optimized” option first. With all the new things that have come out since I built this system, there’s more playing to do before I jump to something else.
Do you have any experience with LLMs in a VM as I have mine? Is that a speed killer as well or just primarily the PCIE speeds?
Only if you are using “cpu offloading”, which means the models layers are split between GPU and your CPU and DRAM.
Otherwise, no, PCIe bandwidth is not really an issue. There are people running 8x GPUs on a single motherboard with x1 to x4 PCIe Gen4 and Gen3 lanes.
This is 100% acceptable for coding models that fit within your 48GB VRAM. It’s only a few seconds longer to load the model into VRAM at first run, and then PCIe is not really used.
And then it will scream! I think you can expect around 150 t/s on a single GPU and around 130 t/s across both with Qwen3.6 (the larger the model, the slower it gets too). Older coding models, like Qwen2.5 would hit around 300 t/s. But don’t use that, 3.6 is far superior with its reasoning and thinking contexts.
The rest of the system doesn’t really matter, as long as you are not offloading any layers to CPU/DRAM (logs will tell you this).
I must be missing something. All the massive R1 models I see are well over 150 GB in size!
Maybe you are picking a massively distilled model?
EDIT: ah, found the distilled version. Which is around 40GB for Q4. Yeah, don’t use that. We have much better coding models now.
So yes, it most likely is offloading to CPU/DRAM and killing your performance. Quite amazed you got it working really. 43GB VRAM sounds about right to leave 5 GB for context window, and the rest gets offloaded to CPU DRAM - at 15 GB/s bandwidth.
Anyhoot, DeepSeek R1 is old school and we have much better coding-focused models these days, like Qwen 3.6.
Yes, extensively and explicitly.
VMs and Containers have no impact on speeds. If anything, it makes much easier to isolate the driver stacks.
If for coding, then yeah standup an llama.cpp instance. For newcomers on the Linux CLI, is recommend Ramalama but you already have a VM. So might as well just spend the time learning the documentation and running llama.cpp directly (llama-server is what you want to read up on).
Personally, I pull docker versions of llama-server so I don’t have to deal with compiling it.
Do searches here for keywords like llama-server, qwen, 3090, etc. You should find quite a lot of examples to get started. Then study the params and advanced techniques, like --kv-unified if that fits your agent workflow.
For a quick and powerful test, install ComfyUI in the VM (shutdown/remove ollama) and try out Qwen3.6 35B. That’s maybe a 5-10m test.
I found the 70B Deepseek model I’ve been using. It’s here. And you’re right - it’s a distilled model (with Llama). The main info page definitely doesn’t state that, so as the article you shared points out, users are a bit mislead in that regard.
It’s been a while since I’ve fired the system up, but I will definitely do that this weekend, if not sooner. I’ll give the unsloth model a run on a llama.cpp setup and report back on my results. I’ll also experiment with how the new setup performs with GPU and CPU models in parallel per my original post.
Many thanks once again for all the great info, @eduncan911. I have no doubt you’ve saved me many hours in learning on my own (possibly the wrong way… again)!
PS - I’m glad there’s no issue. I was also able to find my original source for my understanding that PCIe lanes/speed don’t matter much. For anyone interested, it’s Tim Dettmer’s blog on hardware for AI uses. His info is also the main source for why I chose RTX 3090s for my rig as a cost optimized solution.
I second this, stick to models that fit inside your VRAM pool. There’s limited utility for those extremely large MoEs running hybrid inference. Good for churning out a detailed implementation plan over a few hours but super slow for anything interactive.
Thanks for the links, @lambda . If I’m reading correctly, hybrid inference is using both the GPU and CPU to run a single model. I’m personally less interested in doing that due to the huge performance hit as I am in running two separate models, one on the GPUs/VRAM and one on the CPU/system memory.
I am interested in the fixed Qwen chat template though. Looks like it fixes some problems with Qwen’s original and has instructs for install in llama.cpp that should be easy. I’ll give that a spin once I get the new AI VM set up. Is there anything special I need to know about using an alternate chat template like this? May I assume it’s just plug and play and there’s nothing much else to it?
Nope. You just specify it in the params and it overrides the one within GGUF. I use the same template but didn’t want to overwhelm you during the transition in just getting up and running first with a smoke test.
At this point, read up on what GGUF encapsulates so you know what makes it easy, and what you may want to override. Llama.cpp requires GGUF files.
Okay, so I installed llama.cpp, Qwen 3.6 27B in a new VM then did some testing. Sharing here for anyone interested.
My “old” setup was a single VM running Ollama, Deepseek-R1:70B (distilled with Llama running on 100% GPU), Qwen3-coder (running on 100% CPU) and interfacing using Open-WebUI. The “new” setup is a new VM (cloned from the first one) running llama.cpp instead. The interface for the purpose of the tests was llama.cpp’s cli interface using the –verbose switch to get token speed. I couldn’t get llama.cpp to run two instances - one for a GPU model and another for a CPU-only model - so I ran the simultaneous GPU model + CPU model test in two ways:
using llama.cpp for the GPU model and using Ollama & Open-WebUI for the CPU model and
using the “new” VM for the GPU model and the “old” VM for the CPU model.
In brief, I think the problem with my original simultaneous runs (i.e. the basis for my original post) was a software issue, not a hardware limitation. When running the simultaneous GPU model + CPU Model test within the same VM, the CPU model became worthlessly slow. However, when running them in two separate VMs, they both performed fine. Full dataset is in the expandable section below.
The Deets
In other news, I had an absolute hell of a time with the Unsloth Qwen models running in llama.cpp. I had gotten them to run initially, but at some point they started failing upon submitting the first prompt. It throws a CUDA error - “the function failed to launch on the GPU.” I cannot figure this out for the life of me. It may have to do with the CUDA drivers and I’m wondering if I need to install a much older CUDA version. I tried both the newest 13.3 and the older 12.9. Qwen models from ggml-org work with no issues.
Lastly, I downloaded the fixed chat template for the Qwen models. Like running the models, it seems to work fine with the ggml-org models, but fails with the unsloth models. There’s not really any user output, it just vomits out a bunch of XML at me with or without an error, so I guess I don’t technically know it’s working.
Hmmm… It appears that the expandable section in my previous post decided to lose all of its content. Damn. Is anyone else able to see the data I put there?
Instead of wrangling with Ollama + OpenwebUI, just use LM Studio. It’ll download the required backend/inference bits for you.
The performance drop could be due to number of threads being allocated by default. You can use llama-server to run llama.cpp via an endpoint. For benching, use llama-bench.
Ignore the chat templates for now. You can address that part later, remove them for now.