I’m trying to decide on an SR-IOV GPU solution for my homelab/primary workstation, and I do not want to deal with maintaining hacks and/or license servers. AFAIK, the Arc Pro cards are the only solution at the moment? I have other GPU resources for local AI, but I also have to consider the potential future use of any GPU for AI for all the obvious reasons.
I have loved my Hydra configuration of Proxmox (single computer, multiple GPUs and multiple USB controllers passed through to different VMs for multiple physical workstations from one host-workstation). But the cost is inefficient use of monitor and desk real estate depending on what I’m working on at a given moment.
Previously I had 3 physical workstations, 2 w/1x monitor, 1 w/3x monitors. I moved my Win11 VM to a Testla P4 using Sunshine/Moonlight @4k60 ~20ms latency and am happy with the performance, now to enable that on all VMs and consolidate all monitors and primary GPUs (4090/5090) to my daily driver VM (bazzite/cachy) and access all the other VMs via Moonlight instances as needed. If there are better solutions for this, please suggest.
From reading various threads, I think this is essentially the current state of things RE Intel Arc Pro B-series lineup, please correct me:
B50, 16GB ~$400:
2 VFs (12 with old firmware), no official comment from Intel on final count
Running multiple concurrent or higher resolution/fps streams can get stuttery or bog GPU down, possibly due to 70W power limit
B60, 24GB ~$700:
ASRock: 24 VFs
Sparkle: 7 VFs, and Sparkle claims Intel is limiting to 7 for B60, so ASRock might nerf their current 24
Theoretically should not have same multiple 4k streaming encode/decode issues as B50 due to higher power/capability/etc
B70, 32GB ~$1000:
4 VFs… ookay, wtf ?!?
ASRock vs Intel versions of card?
Could potentially be used later for low-perf AI due to 32GB vram, or used now depending on how much control of how vram gets split between VFs. For now, VFs are priority for me but efficient use of card by also running a small LLM is interesting.
My current thought process:
I think >= 3 VFs are needed, less important VMs can suffer with non-accelerated VNC/RDP/etc.
Near-zero latency needed for 4k stream, between local VMs connected via ConnectX4 SR-IOV. Gaming/high perf graphics are only used on daily driver VM with dedicated GPUs passed through, not important for VFs.
B50 with 4 Vfs would be perfect price/perf, but restricting to 2 VFs feels like a waste of a PCIE slot, and I am leary due to reports of stuttering/perf issues with more than one 4k stream encode/decode.
B60 is roughly double the $ of the B50, but is reported that Intel is expected to support at least 7 VFs on the card. Theoretically this higher W card should not have any issues with multiple 4k stream encode/decode as reported on B50.
And then there is the B70, which could theoretically also be used with lighter weight LLMs, but the current 4 VF limit seems bizarre? So theoretically I could minimize my VF count to only what is needed to maximize VRAM available for low-perf AI. So theoretically more efficient use of resources in one single very precious PCIE slot.
I am currently leaning toward the ASRock B60, but since its $700-ish and a B70 is “only” a few hundred more, I have to consider that one as well. If the B50 had 4 VFs and could easily handle multiple 4k streams, then it would have been a no-brainer for my current needs.
At the moment I’m working on the long term solution for high performance near-zero latency remote-ish desktop for my VMs, but any purchase in the current $$$ market needs a little extra consideration before pulling the trigger.
Or other suggested solution(s)? Advice / $0.02 appreciated.
Asrock B60 owner. original FW was 24vf, 2nd latest FW dropped it to 7vf. Did need new FW to work in unraid/pass through to dockers.
Regarding AI, qwen3-coder just benchmarked 30t/s, just above gpt-oss @ 25, and qwen3.5:27b/qwen3.6:35b/deepseek:35b/gemma21b all around 5-10t/s.
some of this might be down to bad config on my side, but it gives you a bit of an idea.
I have an ASRock B60 too, posted some benchmarks in another thread of how llama.cpp performs on this card.
It depends a lot on how you’re using SRIOV. For me this is on my main desktop. I am only running one VM that uses the VFs at the moment so if I was buying today I’d probably go B70 just because VRAM gets pretty tight when running VF and LLM at the same time.
My setup runs ubuntu, then has B60 in a pcie 4.0 x4 slot and I have a rtx pro 4500 card in the pcie 5.0 x16 slot. (z790 mobo)
Because I want to keep maximum vram available for the rtx pro I set mutter to use the B60 as primary card. I then use vramfs before creating the VF so that I have room to spare for smaller 7-9B class LLMs to run on the B60.
I tried running a graphics benchmark in the windows VM using VF while also running an LLM in docker in the host and normal desktop UI stuff in the host as well. It didn’t fall over but the worst is that the host processes seem to take precedence over the VF so the framerate tanked inside the VM.
The fact that everything ran though in a setup that is hybrid workstation + VMs is very cool though!
The B60 is weak for running larger models.
If I had a wish it would be able to set a VRAM allocation when creating the VF.
It is freaking awesome that you have that working where the host is still able to use the monitor ports and retain some vram while sharing GPU to clients. I’ve wanted that config for many years and it is finally possible, well, not for me yet.
Yes, my perfect virtualization workstation has the physical GPU on the host OS used for physical monitors with same card/PCIE slot providing VFs for VM accelleration that are then displayed via the host’s monitors, and zero extraneous vram burned on $$$ gpus dedicated to AI. Mix and match VMs dynamically for whatever workflow I am (supposed to be) working on at a given moment. ConnectX4(+) SR-IOV between all VMs and host/external network. Maximum efficient use of monitor and desk real estate, and low-to-zero non-AI/game vram usage on 4090/5090. AFAIK Sunshine/Moonlight is still the best perf for RDP/etc scenario.
I run proxmox on my workstation for all the normal reasons, but primarily to be as efficient as possible with $$$ hardware, plus I haven’t landed on a long term daily driver distro, so still using a daily driver VM for easy switching/swapping passed-through hardware around. I was a die hard Mint enthusiast until they dropped support for KDE plasma and I’ve been distro hopping ever since. I even did a Linux From Scratch build for the heck of it. Currently been on Bazzite for a few months now and so far it has been pretty rock solid stable while still being up to date enough for gaming, cachy is next on my list. I am not a real gamer, but for the price I paid for the 4090 and all the rest of the gear, it is about efficient use of hardware when I’m not playing with AI (my excuse for buying it in the first place).
So, long story short, my updated thinking:
B60, “cheaper”, assume only ever used as a fancy but respectable graphics adapter card for multiple vms, plus the host if I ever find a distro I’m willing to install bare metal on my only beefy desktop/workstation box. Could be used later for light weight llms.
OR
B70: 3 VFs for vm low-latency RDP/etc, 1 VF for dedicated low latency AI audio pipeline STT/TTS intended to be real-time if such a thing is feasible with this card.
Either way, the 2 VF limitation on the B50 is a deal breaker for me, though otherwise would probably have been pretty darn perfect for mainly rdp/etc type use, I think at least.
Is this sane, or are there any other sane-er solutions to consider?
I went with the B70. Long story short, the tipping point between the B60 and B70 for me was getting the most out of very limited PCIE slots. Thought I’d give the B70 a chance for light AI workloads and get more off my other cards, while also accelerating some other VMs via SR_IOV.
ASRock B70 with firmware as of 8517 (installed via Win11):
7 VFs.
The blower fan is slightly quieter than the R9700.
I was pleasantly surprised at the B70’s performance for AI, but keep in mind I went into it expecting the worst. Also, my 4090 is used for monitors too, so not purely AI in the quick benchmark. For some very simple super quick adhoc testing for a quick compare:
lm studio
B70 running on Win11/Vulkan, 4090 running on Bazzite/Vulkan
Qwen3.5-27b: B70: 23 t/s, 4090: 31 t/s
Qwen3 30B A3B 2507: B70: 60 t/s, 4090: 100 t/s
I had to hack my old BIOS to add support for Resizeable BAR so Windows will work. The link is floating around in other threads too, but thought I’d drop it here again.
After enabling Resizeable BAR, I saw very small improvement in LLM model load time with my non-Arc gpus (didn’t have before/after experience for Arc), but there were a couple models that I think sped up about 20% for inference, which seems bizarre to me since raw inference should not be affected by ReBAR. It may have just been a weird fluke, but I guess I’ll see.
so i have been having a weird issue. if i shard 2gbs for a windows vm and leave llama.cpp idling with mlock, the system consistently crashes after a random x hours, that can be 10 or 2