So I’m late to the party guys. A couple months back finally started paying for Codex and wow. But. I don’t want to get API fees so I’d like to stay local. I need a GPU as I don’t wanna use my laptop 5080 and my server is an RX5700 I use for Plex transcoding. I’m looking at the Strix Halo stuff but am wondering if I get two AMD R9700 32gb cards (card vs card roughly equivalent perf to 9070 XT). This is way faster and will have 64gb VRAM for about $3k, and since my server is a threadripper (I know) I have enough slots and lanes to go to 4 cards at 128gb should the need arise.
Question: Can I use 2x 32gb cards roughly equivalent to one 64gb card just twice as fast or does the model have to fit on one GPU?
the idea would be to up the vram total to be able to load a larger model / larger context / overall less quantized llm - at speed. with 2x r9700 you will be having a good time with Qwen3.8-27B. 4x r9700 then Qwen3.8-Flash-Next, Deepseek v4 Flash 0731 / GLM 5.3 flash (with some ram overflow) is probably where you will be aiming at - depending on use case (eg how many users need access to the ai?)
I’m getting my dual r9700s setting up on the old ThreadRipper 2950 DDR4 256GB 4-channel build as well for giggles. I want a baseline before investing into the CPU my “needs a CPU” ThreadRipper WRX80 8-channel of same ram build.
You didn’t mention what you wanting to do with local Ai.
Just Codex? One thing I learned from a single R9700: you want to run multiple models, not just one. The latest hotness is to run a smart router that picks the best model for the type of request. For example, I use Qwen 2.5 4B for Intellisense as it is nearly instant on response, running at 250 t/s insanity. Though now I offload that to the CPU and DRAM only.
Then there’s “always listening" live voice models, the router model itself, vision models for desktop app and/or web browser testing, bigger 7B but without thinking configured, and so on. You want these hot, loaded in memory at all times with concurrency.
Hence the 2nd R9700 @ $1200 (buy used, not new) to increase VRAM. I last used multiple llama runners under llama-swap but currently setting up a smart router (there are a couple different ones) and moving to Turnstone agent harness.
I’m still working out the security angle so not ready for the closed-loop workflows just yet. Want to nail down the smart router stuff first and optimize model parameters for concurrency over the PCIe 3.0 x16 (aka Gen4 x8) buses first - again, as a baseline before moving to dual Gen4 @ X16.
I have a R9700 and a W6800 in my machine. I occasionally run one model on two cards that would otherwise fit on just one, only because it’s marginally faster while processing batches. More useful if you are running, say, Claude Code with subagents that run in parallel but without the ability to direct subagents to use individual GPUs (or maybe I’m too dumb to figure out how CodeRouter works). This way, llamacpp doesn’t drop context checkpoints for subagents because I can offer more KV cache, and enable more parallel decodes. Not spending an eternity processing every single prompt makes things faster by proxy. You could also have your KV cache in RAM, so probably not that big of a deal when you have that instead.
llamacpp will let you use tensor offloading that allegedly gives you more speed, but I didn’t really notice a difference with Qwen 3.8 27B compared to dialing down context length and running it on one card. I think there was some difference in prompt processing, but I really couldn’t be sure. Also note that I am on a 5950x, one of my cards is a W6800, and I’m not using PCIe peer-to-peer. You might get different results on your hardware, though I doubt it’d be by much since two cards just adds more overhead as your system tries to keep them in sync.
IMHO, mine is definitely an edge case of unique needs and poor support from Claude Code. It’s not worth it just for the speed. You might see a bit of a difference if you’re doing layer offloading and running many requests in parallel, but not to a degree that would be worth the extra $$$. I’d use the extra VRAM to run better models instead, or just use the one card if you’re fine with 32G.
Just me, this is for me at home. Obiously I got the bug, and would like to start playing with agents and such but I don’t want to use a pay-per-use type API, I’d like to keep it local.
Right now my big use is I’m using it to code, go through my entire self hosted stack and optimize, secure and I even got it to make a grafana dashboard (actually a series of dashboards). I’m truely impressed it installed Loki and promethius, set it up, created a dozen different dahsboarss with a dozen or more metrics each and I’m just now realizing this is more than a parlor trick “make me an image of a spaceman rising a unicorn on the moon”.
Also, as to the “once as fast” comment - does that mean TWO cards gives you twice the ram, but only the speed of a single card? Ugh.
No, it depends on the model and it’s support in your inference engine of choice. In general both vLLM as well as llama.cpp provide tensor parallelism. This is basically equivalent to both GPUs working in parallel.
As a note, it does not work always and when a new model architecture comes out the people will need to add these feature for this model first. But Qwen3.8-27B is well supported. Qwen3.8-Flash-Next for example is still lacking the tensor parallelism as of today (if I am not mistaken, have not checked the past two days).
I’d say go with llama.cpp first, you can either use the Vulkan or the ROCm backend with those AMD cards. Either work when you set the correct parameters during compile time. They both have their advantages, you might try them out as needed.
I’ve been looking all over, and cannot find anything close to that price. In face, new price is $1600-ish and used prices at least where I looked (ebay) are way over that. Is that the “stalk facebook marketplace for the unicorn” deal, or am I missing an awesome marketplace somewhere?