I have 4x 6000 Blackwell and serve 10 devs and 40 office workers with it

I’m the Linux admin in a company that does print stuff for their customers. Think invoices, medical reports and so on. Privacy is extremely important, and our customers are demanding 100% on-premise only IT.

In 2023 i started playing with llama.cpp on some of our beefy compute nodes. It ended in Deepseek R1 671b oneshotting a working flappy bird clone at 0.7 t/s decode speed over night.

C-level ignored my pleas for dedicated hardware, until end of 2025, when i got a 2U Gigabyte with 4x 6000 Blackwell MaxQ. I was quick to setup OpenWebUI and a bunch of models for testing.

Our business is slow, not much happens. There is no software left to write, no sudden changes, no need for a proper AI strategy. This isn’t the heart of Meta or Google, it’s the same business as ist was 20 years ago.

However, users started to use ChatGPT and other AI services for their office work, reformatting emails for example. C-level allowed it, as long as the data is sanitized. However, that just isn’t possible, as we have seen a few times. But our local AI wasn’t ready, until Qwen 3.5 122b happened. Since then we are 100% local. We then went to 3.6 27b and now 3.8 27b. We want to use Deepseek V4 Flash 0713, but can’t get it under control yet. People prefer 2.8 27b.

Other apps we run, besides OpenWebUI, are Vexa AI, a Teams bot that uses speech to text to make transcipts, FIM models in VSCode, Hermes or OpenCode inside VMs on the developer workstations, Qdrant for RAG databases, ComfyUI, and N8N. We are thinking about getting another box for redundancy and load balancing.

6 Likes

It’s a pretty straight setup, but if there is interest i can share some details.

I run both DSv4 and Qwen3.827b as well as other models within my mixture of 6 pro 6000s and a 5090. I recently set up DSv4 on 4 pro 6000s and am producing over 850 tg/s across 16 concurrent clients. Around 100t/s for an individual client. PP spikes in the 10s of thousands. Basically it finishes prompts faster than i can type them.

I’m using SGlang rather than vLLM which was a challenge to get set up. If it wasn’t for GLM5.2 i probably wouldn’t have finished it. Qwen3.8 is too gd stupid to get it set up for me.

DSv4 is pretty darn smart and i’d still be using it now if I wasn’t able to get Kimi K3 going. If I were running a business behind my workstation there would really would be no other good option than than DSv4. qwen3.8 is definitely not the answer.

Qwen3.8 is great for repetitive, grunt work tasks that still elude the capability/feasibility of traditional automation. A great example, I have a task that I perform a lot that is actively hostile to automation. Qwen3.8 handles it with no problem, every hour like clockwork. Sometimes it takes ten minutes, sometimes thirty. Sometimes it struggles to get the answer, but it always does eventually. That task doesn’t need frontier intelligence, or even a priority spot in the rack. I would bet the average IT job is mostly made of tasks like this.

Agreed, DS4F is a pretty good option, but i can’t seem to get it stable and predictable. 3.8 27b will deliver what it can, and it will do so even in the 20th round. DS4F gets unstable the first time Hermes/OpenCode compacts or somewhere in the 6th or 10th round during a chat. It starts leaking tool calls, switches to Chinese and becomes unusable. A problem in my setup i have not been able to solve so far. The user preference right now is firmly on 3.8 27b (i use OpenWebUIs Arena feature) for chat, office work and tiny RAG (knowledge in OpenWebUI).

  • Getting DS4V stable with two cards is my prio #1.
  • Next i have to deliver a few ideas for redundancy, because right now we have none. It’ll take some time to turn into hardware.
  • After that i’ll setup some kind of routing layer, i’ll probably skip LiteLLM because of it’s history with supply chain attacks, but maybe they solved the problems for good.

And then i wait for the next model that blows what we have so far out of the water, at the same hardware requirements. I’m hoping for Qwen 3.8 122b, one of the Qwen developers dropped a “we will release another model in the 3.8 range and it will be larger than 100b“. Could be 122b-A10b or 397b-A17b.

How do you run it? We use llama.cpp with a Q8 quant, MTP and tensor split over two 6000 blackwell. Users love it for office work (reformat this email, how do i do this or that in Excel, research this or that and write a Wiki-Article).

Ah yes, i also used llama.cpp at first with DSv4 and also found it super unpredictable. Its important to note that while you are dealing with technical issues related to DSv4 on llama.cpp, you’re also dealing with performance issues.

llama.cpp doesn’t handle concurrency very well and handles context very poorly. So while with llama.cpp your first prompt is lightning fast, each additional prompt builds ctx and you progressively get slower. I use llama.cpp quite a bit because of my weird tensor splitting but if its possible to fit a model within a power of 2 number of GPUs i definitely avoid it.

With SGlang, it loads up addressing a full 1m ctx from the start, you don’t decide that like with llama.cpp. This means that the size of the model more than doubles in VRAM which is important. You’ll probably have to drop down to a quant if you want to fit it on two cards. It fits perfectly on 4 pro 6000s though.

With Sglan, most importantly, there is almost no impact to speed as you run up your prompts. When i was at 500k plus ctx built up about 5 hours into a big prompting/planning session, i was down to about 80 tg/s at the very slowest.

Here are my test figures with Sglang, DSv4, 4x pro 6000s across multiple concurrency setups

Workers | Reqs | Out tok/s | Tot tok/s | Req/s | Avg lat | Avg out/req | Err
───────────────────────────────────────────────────────────────────────────────
4 | 52 | 317.8 | 352.2 | 0.9 | 4938.7 | 367 | 0
8 | 96 | 567.3 | 631.3 | 1.6 | 5331.1 | 355 | 0
16 | 144 | 851.5 | 947.5 | 2.4 | 6522.2 | 355 | 0

I run it at Q8 on dual 9700 pro cards or at Q4 on a single B70 pro depending on the complexity and urgency. This is with Vulkan llama.cpp, layer split. I’ve never gotten tensor split to work.

Maybe i wrote that wrong, i use llama.cpp for Qwen 3.8 27b. For DS4F i use vLLM. I don’t think it’s a llama.cpp or vLLM problem either, it’s a templating/tool calling problem. Performance is great, i get 170 to 200 tokens generated for one client. That is why i need to get it robust so it can replace 3.8 27b.

Ah ok, no I made that connection myself.

I never ran DSv4 on vLLM so I can’t speak to it but it sounds like you are having the same issues i had with llama.cpp. Random tool call failures were so frustrating. It would randomly forget it could recall previous sessions or anything beyond its own memory. I tossed DSv4 within a day of its release for this reason. Its only because of SGlang that I went back to it and i’d still be using it today if not for K3.

I’ve been running DS4F 0731 pretty reliably on 2 x PRO 6000s for the past few weeks on long-running agentic workloads and while I’ve seen some corrupted CoT, I haven’t run into problems w/ it even running for days on tasks.

I am running vLLM but with PR #50276 (merged in nightly now, not in stable - it’s a must). Also, running DeepGEMM to get decent speeds. I’m running it TP2/EP2 - it has a 700K FP8 KV token pool and I run it at a 256Ki window (which may be part why it’s been quite stable for me - 384Ki is probably undertrained). It’s decently fast (short c=1 9K tok/s prefill, 111 tok/s decode).

That being said, I think especially if your users like Qwen 3.8 27B and it does the job there’s no reason to switch. Running the official FP8, at similar TP2, while the prefill is a bit lower (7K), the short decode is actually higher - 174 tok/s w/ DFlash2, and it has 3.23M of token pool. It also runs significantly faster at concurrency (2X throughput w/ ShareGPT c=32).

In my coding/agentic benchmarks, Qwen 3.8 27B Medium scores a bit higher than DS4F 0731 (at any reasoning level), although so far, subjectively I personally prefer DS4’s reasoning style better and I haven’t run Qwen 3.8 on enough agentic workloads to really see it in action yet.

1 Like