As the title says - show off you Ai cobble job! What you run, with what backend / frontend?
I figure a lot of us are running local so why not show N’ tell? (let the jealousy flow through you…)
As the title says - show off you Ai cobble job! What you run, with what backend / frontend?
I figure a lot of us are running local so why not show N’ tell? (let the jealousy flow through you…)
I’ll start it off:
aigoat@aigoat
OS: EndeavourOS x86_64
Host: Relion 2908GT (0100)
Kernel: Linux 7.2.4-arch1-2
CPU: Intel(R) Xeon(R) E5-2697A v4 (32) @ 3.60 GHz
GPU 1: ASPEED Technology, Inc. ASPEED Graphics Family
GPU 2: NVIDIA RTX A4000
GPU 3: NVIDIA RTX A4000
GPU 4: NVIDIA RTX A4000
GPU 5: NVIDIA RTX A4000
GPU 6: NVIDIA RTX A4000
GPU 7: NVIDIA RTX A4000
GPU 8: NVIDIA RTX A4000
GPU 9: NVIDIA RTX A4000
Memory: 31.32 GiB / 125.76 GiB (25%)
Swap: 0 B / 138.34 GiB (0%)
Disk (/): 889.43 GiB / 1.73 TiB (50%) - btrfs
————
omnigoat@omnigoat
OS: EndeavourOS x86_64
Kernel: Linux 7.2.3-1-cachyos-custom
CPU: AMD Ryzen 5 5600X (12) @ 4.88 GHz
GPU 1: NVIDIA Quadro RTX 5000 [Discrete]
GPU 2: NVIDIA Quadro RTX 5000 [Discrete]
Memory: 9.41 GiB / 31.21 GiB (30%)
Swap: 511.88 MiB / 68.84 GiB (1%)
Disk (/): 69.27 GiB / 883.03 GiB (8%) - btrfs
This little guy is running an asus x470, pcie 3.0 x8 for both cards with a nvlink between them. It is a great 256k @q8 kv - Qwen3.8-27b q4 unsloth gguf runner. ~50ish tps via llama.cpp, not tea bag.
Oh, you wanna see my meath lab?
Epyc Rome 7532
128GB DDR4
Various NVMe totaling ~5TB of storage
4x AMD V620
All in an open air crypto-mining bench. No, this does not trip breakers, but it comes close.
Current stack is a LiteLLM frontend to vLLM (I’ll do a write-up on how I optimized vLLM soon), serving Qwen3.8-flash-next at int4 quant, with the n-gram table offloaded to system memory. Gets me ~70TPS of single-request inference and ~90TPS peak for 2-3 simultaneous streams. This is without MTP, because MTP is not reliably performant with multiple simultaneous requests being served. Still trying to figure that one out.
Also, yes. I know this is a crime against cable management. No, I will not fix it.
Pretty boring compared to some of the posts here.
Dual 5060ti 16g, 64 gig ddr4, 5700X on an MSI board, Meshify C case
Ubuntu 26.04 (desktop even), llama.cpp (llama-server)
Qwen 3.8 27b, Q4 weights, Q8 kv, sampler settings for coding, full 256k context.
This machine runs a scheduled ralph loop, 12 hours a day through pengy-cli in one shot mode iterating on the loop. Can keep the machine occupied most nights building some big rock items I need for the larger universe of junk running here.
That’s too nice… I don’t know if I can sit at your table… ![]()
If it helps, the office it lives in is REAALY messy right now.
(Imagining an immaculate shining beacon of clean, with an open bag of chips on the desk…)
Have this for a while, I’m just too busy to tinker with it.
2x AMD EPYC 7532
512 GB DDR4 3200
Some storage ![]()
2x 3080 20GB (Total 40GB)
Ooph, that ram tho - that’s a college fund right there.
I wound up popping an AMD Instinct MI210 (64GB HBM2, 1600GB/s) into my existing main Proxmox node, and passing it through to a VM.
So it is less of an “AI rig” and more of a large hypervisor that also has a VM that does AI stuff, as one of its uses.
Turns out that despite having a proper server with good airflow front to back, that just wasn’t sufficient for cooling the thing.
I tried one of these 3d printed 80mm fan shrouds I found on eBay, but it was too large and didn’t fit:
So I wound up ghetto modding it by cutting up a plastic cutting board and making my own shroud and using foil tape to hold it together.
It looks a little ghetto, but it gets the job done.
Sadly I forgot to take a picture of the ghetto mod.
The tech priests require art and evidence of your devotion.
You’re gonna leave us hanging without proof of this described handy work? I can’t be the left to be the only one showing off my superior cardboard skills… ![]()
I’ll be sure to take a pic the next time I have to open it up.
But it ain’t pretty. I promise. There is an 80mm San Ace fan ghetto mounted to the opening on the card using a little hole I drilled in the rear support bracket to secure it in place with zip ties.
Then I shaped a kind of trapezoidal box out of plastic cutting board material, and taped that together using HVAC foil tape, also using some tape to seal it so no air leaks outside the GPU.
The little 80mm fan wound up sandwiched straight up against the chassis fan wall, so it is not ideal, but it DID work.
Id;e is down to 46C which is quite good for one of these, and it rarely hits over the high 60’s under AI workloads. Pre-fill is when it gets the hottest.
Have you seen my rig? ![]()
Pretty doesn’t factor in, what matters is if it LLMs.
We can’t all be that awesome ![]()
Hey, I’m thoroughly jealous of the MI210. That’s a lot of GB/s
Plus, if it’s jank, might break, and runs hot, it’s worthy of respect in my book.
It is, and it comes in handy sometimes (2500 tokens/s pre-fill and 80 tokens/s eval on Gemma4:26b) but it is a little bit more impressive on paper than it is in reality (something I didn’t realize until after I bought it)
It turns out AMD’s ecosystem is nowhere near as optimized as Nvidia’s, so when it comes to efficiency, I rarely get more than 60% of the theoretical max throughput, which is a bit of a bummer. But it is still quite nice.
I did also realize after the fact that this earlier generation of Instinct cards lacks FP8 support which also stings a little…
I do really appreciate the high speed pre-fills though (though I admit, I don’t fully know how that compares to other hardware out there)
A ~30b class dense model at Q8_0 usually evaluates at about 25-30 tokens/s, which is less exciting than I had hoped for when I forked over the cash for the thing, but that’s what happens when you first get into stuff. You learn along the way, make mistakes, and some of them are expensive.
If I had to make a decision all over again, for the money I spent, I might get a couple of DGX sparks. They would evaluate much slower due to their much lower memory bandwidth, but damn would it be nice to have 256GB of RAM.
It’s just the cost of learning ![]()
Have you built an optimized stack for it? You might want to take a look at this repo, if you’re just throwing llama.cpp at it and crossing your fingers, AMD stacks really need a bunch of custom patches and optimizations, especially for the older stuff. That’s how I went from 20t/s on my stack to up to 90.
My 4x v620s does about 1200TPS prefill with qwen3.8-flash-next. You’re getting excellent prefill performance out of it. (to be fair, you do have 3x the memory bandwidth of one of my GPUs though)
I appreciate the info, but yeah, the figures I am seeing are with llama.cpp compiled with HIP support and all of the special gfx90a flags in it (which was a pain in the ass due to all of the dependency conflicts)
I tried a few different versions of ROCm (my theory was that more recent ones might be more optimized for more recent GPU’s and older ones might be better, but that turned out to have negligible effect)
At least it did better this way in llama.cpp than in vLLM, which I never got working satisfactorily due to performance kernels in it not supporting the gfx90a architecture.