Radeon Pro AI R9700 fan noise? Which model to buy?

Hi all,

I’m looking at adding two R9700 cards to my desktop to fit larger LLMs.

Currently running an Intel Xeon 8368 w/ 256GB DDR4 and an RX 7900 XTX for AI/compute & RX 7600 XT for rendering the desktop without being choked by compute workloads.

I’ve read some worrying remarks on reddit about the fan noise on these blower cards. I’ve had blowers before (Nvidia Founders Edition cards) and those weren’t that bad, but the AMD cards are said to be much louder.

My workstation is in my desk which is a quite place in general. I don’t mind some noise when under load but I don’t want the noise level of a datacenter GPU in my desktop either.

Can anyone who has tried these cards describe the noise level (maybe compared to other cards or hardware) and what brand of card they got? I’m also curious if there is much of a difference between the different vendors in terms of fan noise / behaviour.

1 Like

I have 2x of the ASRock R9700. The PC lives at my desk as well. I’ll ramble my thoughts below.

At full power and full load, I find them to be disruptively loud. They are not high-pitched and annoying, so if you wear nice isolating headphones, you might not really be bothered. My headphones don’t have much isolation (and I’m often not wearing headphones) so I’m not keen to run them at full tilt. YMMV there.

That being said, I reduce the power in the Adrenaline software by -30%, and then they’re pretty quiet even under full load together. The way I use them is not especially sensitive to time/throughput/raw performance so this is fine for me; I honestly haven’t even benchmarked to see what the performance difference is, the noise (and power/heat) reduction is way more impactful than any tok/s in my case.

Can’t comment on different vendors - I have a local Micro Center and was able to get them open box, so they were simply the best deal by far ($1110 before tax). For me personally, it has been a long time since I’ve used any high-wattage blower cards so it’s hard to compare. I was expecting a blower at 300W to be loud so this wasn’t shocking (and I doubt any of the other vendors would manage to be much quieter, but don’t take my word for it).

Still, I was very pleasantly surprised by 1) how they are much more wooshy than whiny, that’s a big plus and 2) how much quieter they are with power tuning. You can probably be a lot less aggressive than -30% and still tamp the noise down considerably.

Without any tuning, they are (each) definitely louder and more disruptive than my 5090 at full load, but maybe less annoying than my older EVGA 3080Ti at full load (because the latter has much more of a “whine” to it). With tuning, if I’m not wearing headphones, I can tell they’re working, but they easily fade into background noise. When the A/C is running I can’t hear the PC at all. So, there’s a wide gulf to play around in between “barely notice” and “too loud” (for me at least).

TL;DR: Stock behavior on an ASRock model is probably not well suited for an otherwise quiet environment, but headphones with good isolation may help a lot. That said, there’s a lot of wiggle room in power tuning and undervolting. I run a set-and-forget -30% power limit with no bells or whistles and then they’re barely noticeable, so I imagine there’s a lot of room to find your sweet spot.

3 Likes

Thanks for the info! What are you using them for and how is performance when you need to span models across both cards?

My implementation is still maturing, but I’ve been putting a few models through their paces to decide what I’m keeping. With LM Studio at least, models spanning cards works pretty seamlessly (my cards are in 8x/8x Gen 5). Since my use cases are not really time sensitive, I’m using CPU offload pretty aggressively with 192GB DDR5 as well, in order to fit in larger models.

Can’t quite reach Qwen3.5-397B, but quants of EDIT: (Qwen3.5-122b-a10b@Q8_0, not Qwen3-235b; I actually was disappointed in the outputs of the latter, even though it ran fine), the new Nemotron 3 Super, and Minimax M2.5 all work for me. My plans are typical LLM stuff, I know it won’t beat ChatGPT, but being private and owning my own data, and being able to use it with work docs and such without exposing those docs to OpenAI will (hopefully!) make it super useful anyway.

I have another smaller setup for recording calls and transcribing them, then this machine will soon be summarizing/extracting info from the transcripts and filing them away, (again - hopefully!) building a personal use knowledge base around my job without needing to document it all myself.

Performance across cards will depend heavily on the software you’re using. vLLM might be able to do good things with the tensor parallelism, but llama.cpp doesn’t (at the moment).

There’s also the consideration that Vulkan on llama.cpp is faster than ROCm on vLLM even without tensor split, so…it’ll depend on the models and your specific setup.

For reference, even with one of my cards in a slot hanging off the chipset, I get ~80t/s running Qwen3-Coder-Next-80B running across both cards with llama.cpp and --split-mode row, and 110t/s with Qwen3-Coder-30B across both (I have both models loaded, the former for agentic stuff and the latter for inline completion).

Neither card goes over 70W-ish when run like that at the moment, mainly because of the lack of tensor parallelism. Unfortunately, the PR that’s currently in-progress for tensor split on llama.cpp doesn’t fully work with Vulkan, and when it does it’s actually slower than row and layer split. However, the performance is currently adequate for my purposes, so…when it finally arrives and gets fixed for Vulkan, I’ll take the boost as a happy thing.

My system is 16x Gen 4 on 4 slots, so same bandwidth. Hopefully it will work just as well. I have 8 channel DDR4-3200 (~220GB/s system RAM), but a single CPU (Intel Xeon Platinum 8368), so CPU offload doesn’t work great for me with my current setup.

ggml: backend-agnostic tensor parallelism by JohannesGaessler · Pull Request #19378 · ggml-org/llama.cpp · GitHub is this the PR you’re talking about?

@wimm - that’s the one :slight_smile: If you want to give it a go, you should bear in mind that the branch is a long way behind master.

I wish I had 8 channel memory! I’m working with measly dual-channel on an Intel Core Ultra 265K. As mentioned, the speed is fine for my purposes — I’m perfectly fine waiting 5-15mins for a response to complete while I’m doing other things, as long as the output is of sufficiently high quality. In that context even 5-10 tok/s is plenty, so I’ve just been aggressively chasing the largest models that will fit with big context, without much regard for throughput.

I’m much more inclined to use LLMs in a “fire and forget” manner where I send off what I want and I come back later to find it done. That said, I’m not using it for coding, and I’m sure that if I were to start coding with it at some point I’d prefer more real-time interaction, and settle for a smaller model suited to the task.

But the whole “prompt 5-10 times and combine the outputs and keep refining and babysit the LLM to get what you want” feels self-defeating to me. I usually find that whatever I want done would be faster to do myself than to provide my full attention to coaxing a lesser LLM to an output I can use/keep. I use them to offload things from my human RAM to machine RAM so I can focus elsewhere :grinning_face_with_smiling_eyes: This way I can “double-task” and get more done, with the added benefit of also satisfying the ADHD need to explore random questions/subjects throughout a day without getting obsessive over it (dump the question out into an LLM and get back to whatever I was doing, forget I ever asked the LLM about it, find the response again later… :sweat_smile: )

Not a knock on anyone who finds a lot of value using them in other ways, this is just the only way I have personally found value in them. For all I know it’s a skill issue. But when I am able to get this “flow” to work, it works really well for me.

For those wondering about my original question: I bought 2 Gigabyte R9700 and the noise is absolutely fine so far.

I haven’t managed to get a stable and high speed inference solution out of it yet, but that’s for another thread :slight_smile:

1 Like

If I get some time in the next couple of days I can put up a guide to my setup if you want?

I’m using Vulkan, rather than ROCm, with llama.cpp set up as a service that always points to the latest downloaded version of llama.cpp, running in router mode for runtime model selection and with per-model configuration.

I usually use it to run Qwen3-Coder-Next-80B-REAM for agentic work with 256k context, alongside Qwen3-Coder-30B-A3B-REAP with 32k context for inline completion. During the day, at least (at night, it’s batch processing text for my forum).

3 Likes

That’s awesome, I’d love to see it.

Are yours Rev1.0 or 1.1? I have a couple of Rev 1.0 and have seen some issues with my temps. I have since found that Gigabyte introduced 1.1 quite early as they had serious issues with the early boards.

Did you ever get around to it?

I’ve run some Qwen3.5 and now Gemma4 on mine but performance has been rather lackluster for me.

How do I check this?

I didn’t - I got overtaken with actual work (the b’stards).

For what it’s worth…Qwen3.5 can be quite slow (eg the 27B model is around 25-30t/s), unless you’re using the 35B-A3B model (can get up to 130t/s, depending on the quant). I haven’t managed to get self-speculative decoding working yet, but I’ve found that the current 3.5-series models aren’t as good as Qwen3-Coder-Next for code-based work.

Gemma4 support in llama.cpp is pretty new, I’d give it a couple of weeks before the real optimisations start coming in.

What do you call “lackluster” performance?

“Lackluster performance” is really subjective and depends on your use case. I for example find even 10 tok/s totally usable for everything I want - faster than that usually isn’t even useful to me, as it’s working faster than I am.

Like digitalscream I think my qwen3.5 was getting 20-30t/s. Gemma4 is getting 15-20 (31b dense, q6_k). EDIT: Wanted to note that Gemma4 MoE rips pretty fast, though. I don’t remember the stats because I found the outputs weren’t hitting the quality I was looking for but dense Gemma4 does.

For me, Gemma 4 26B A4B Q4_K_M runs around 85-90t/s on dual GPUs, 117t/s on a single GPU.

That’s about the same speed as Qwen3-Coder-Next-80B. I suspect that there are more optimisations to be made.

EDIT: By way of comparison, Qwen3-Coder-30B-A3B Q4_K_M runs at 110t/s on dual GPUs, 140-150t/s on a single GPU. That might seem like similar performance, but on dual GPUs it’s the difference between “can be used for inline suggestions” and “feels like it’s getting in my way”.

1 Like

you won’t hear it unless you’re doing something, then you really hear it

I like to put 90mm fans with rubber feet on the back of GPU’s where the core is to help keep cool but that might be too redneck engineering for most peoples workstation

Can anyone else share experience trying to limit the noise R9700 blowers?

I am planning to get an Asus R9700 (apparently has less annoying blower sound). My issue is that my workstation is in the living room, and I doubt that the typical blower sound is compatible with other family members hoping to have a restful evening when I start my coding hobbies.

I understand that by limiting the card power, it is possible to reduce noise significantly but not sure if that is enough. Would be grateful for any experience that people could share!