So what does having 2 cards buy me vs just 1 card. I am trying to figure out if I should order a 2nd card for my workstation build. I am in the cybersecurity field and want to learn AI. I am also interested in potentially using AI to make long for high quality videos. I might also like to experiment with some high level local inference like trying to run deep seek R1. I will have a TR 9880 and 384 GiGs of RAM. Would like to know if this is a worth wile upgrade or if I am just wasting my money. Does having a 2nd card open up a lot of extra utility? I can get the card for about 7K via and EDU discount.
Yeah vram is king.
Depends what you want to do ofc but yeah get it.
Keep in mind it’s 600W each if you take the workstation edition. It heats up the place at full blast for some times
Initially I was just using the second RTX Pro 6000 in my workstation to run larger models faster. Now that the initial novelty has worn off I find myself using smaller specialized models spread across the two cards. Ex card 1 is running re-ranking, embedding, and an image model, card 2 is running a large language model, as an example of one of my workflows.
I imagine as I add additional cards I could build more complicated workflows but for now this seems to be working for the things I want to do.
Generally, same for me.
GLM/GLM-Air may be a bit of a game changer for me - they’re very good. Air’s speed and quality may move me completely away from the thiccbois, and just use that as my go to local LLM. Run that on one card, and a combo of other things including desktop, games, image gen, and vision models on the other makes for a really nice experience right now.
What’s great is, I still have an option to load up a big Kimi or Deepseek setup if I need it for something. I have it all dockerized so it’s super easy to spin up/down.
If you can swing it, multiple 6000s is a lot of fun.
They are basically 5090 with enough VRAM (according to what we are trying to do with them, anyway), but the default workflow through WAN 2.2 is still going to take around 20 minutes for that 5s clip.
Being able to run multiple 30B models at the same time? very convenient
I’m even considering running two models at the same time and maybe make them try that kaggle chess thing…
I’m getting two with the idea that at the very least I can have a huge model using one, and the other one is free for GPU-intensive non LLM work!
That’s probably a pretty silly justification though:)