Radeon AI PRO R9700 x2 — card drops to Code 31

Four days into this and I’ve eliminated most of the hardware. Posting in case anyone

recognises it, or has an R9700 that behaves the same way.

Two Radeon AI PRO R9700s. Whichever card has **no display attached** dies within 5

minutes of idling or so. Device Manager gives **Code 31 (`CM_PROB_FAILED_ADD`)** —

*“Windows cannot load the drivers required for this device. {Operation Failed}”*.

The kicker: **while the card is in Code 31, GPU-Z still reads `PCIe x16 5.0 @ x16 5.0`

off it**, along with the VBIOS build string, device IDs and memory size. The card is

awake on the bus, trained at full Gen5, answering config-space reads — and the driver

simply won’t attach to it.

CPU Threadripper PRO 9955WX

Board ASUS Pro WS WRX90E-SAGE SE |

RAM 192 GB DDR5 ECC RDIMM

GPU 2× Radeon AI PRO R9700 32 GB (ASRock)

OS Windows 11 25H2, 26200.8973 |

Driver Adrenalin 26.7.1 (`32.0.31035.1003`), also reproduced on 26.6.1 |

I’m at a loss here, I did install Ubuntu 26.04 and the system ran just fine no issues, I have another workload that requires windows but its just not stable at all with the video cards in.

1 Like

Can you put your hand in the card and see if it’s crazy hot? There might be a bug where the fan doesn’t ramp if no display is attached

I had this same problem. I cannot explain it, but it is not the hardware. I went through two sets of cards trying. It is 100% the drivers. No amount of fiddling, even running windows in a vm under proxmox with passthrough could fix this.

I’m very sorry to say, you will just have to run linux. But Linux is faster and works better anyways, so… I don’t see this as a huge loss personally. If your workload hard requires windows, my condolences to you.

2 Likes

Run a full DDU (Display Driver Uninstaller) from Safe Mode and if that doesn’t work, downgrade to Windows 11 Enterprise LTSC 2024.

Hypothesis there being that 25H2 might have altered some kernel/driver APIs and/or behavior, which breaks the PRO drivers?

One weird workaround I haven’t fully proven, but it seems to work, is having a monitor plugged into each card. The system is stable, no issues. BUT I have not fully proven it; I’ve been playing part swap too much.

@wendell The card seems to be just fine temperature-wise, not super hot. I’ve not even gotten to the point where I can perform any type of workload.

@Strawberry This is exactly where I feel I’m at. I started running Claude to track what’s going on, and it’s more than once said the same thing: give up, use Linux…. I’m just stuck with Windows for some specific tasks. I was about to RMA my video cards trying to figure out what in the world is going on.

@Ewout I think this is a valid test that I’ll run. I have run DDU several times. I also reinstalled Windows to get a clean install. Same issue. Trying to go to Windows 11 Enterprise LTSC 2024 might be the option I need to try next to rule out Windows being Windows…..

Below are the notes from Claude. I’ve been using it to help me track what I normally would be able to figure out…… but this is a head-scratcher….

Radeon AI PRO R9700 (dual, Windows 11): headless card dies at idle — code 31,

silent 0x1B0 VIDEO_MINIPORT_FAILED_LIVEDUMP, AddDevice returns 0xC0000001

System

  • Threadripper PRO 9955WX, ASUS Pro WS WRX90E-SAGE SE (BIOS 1203, stock defaults), 192 GB DDR5 ECC
  • 2× AMD Radeon AI PRO R9700 32 GB (ASRock-built, 1002 7551 / 1849 5413, VBIOS 023.008.000.068.000001 on both)
  • Windows 11 25H2 (26200), Adrenalin 26.7.1 (32.0.31035.1003), clean install after full removal
  • Resizable BAR enabled. Slot 1 card drives a monitor; Slot 3 card is headless (compute only, llama.cpp Vulkan)
  • Same hardware previously ran Ubuntu + ROCm 7.14 for weeks under near-continuous load with zero issues

Symptom

A headless, idle R9700 drops dead: Device Manager code 31 (“Windows cannot load the drivers
required for this device” / {Operation Failed}), anywhere from seconds to ~42 minutes after
boot, or ~13 minutes after a compute load stops. Only a reboot brings it back.

While “dead,” the card is still on the bus: GPU-Z reads full config space, PCIe x16 5.0 @ 5.0,
correct VBIOS/IDs/VRAM — but clocks are 0 and no driver is attached. So it’s not falling off
the bus; the driver fails to re-attach.

Under sustained load it never fails (38 and 382 clean llama-bench iterations; ~207–209 tok/s
tg128, above my own Linux numbers). Load and an attached display both prevent it. With displays
on every card: 10h+ clean. Went headless again: failure within minutes. Both cards have failed
this way at different times (whichever was headless), so it is not one defective card.

The part Windows hides

Nothing appears in Event Viewer or Reliability Monitor. But (elevated!)
C:\Windows\LiveKernelReports\WATCHDOG\ had been silently collecting live kernel dumps
for days:

VIDEO_MINIPORT_FAILED_LIVEDUMP (1b0)
Arg1: 1  (Add device failed)
Arg2: ffffffffc0000001  (STATUS_UNSUCCESSFUL)
FAILURE_BUCKET_ID: LKD_0x1B0_dxgkrnl!DxgCreateLiveDumpWithDriverBlob
Stack: PnP device restart -> dxgkrnl!DpiAddDevice -> 0xC0000001

Five of these across two motherboards. Disassembly of dxgkrnl shows this specific dump type
only fires when the failure is at/before the miniport call, and the status is amdkmdag’s
own DxgkDdiAddDevice return — i.e. the AMD driver refuses to re-attach to the card after
something kills it at idle. Each dump was preceded within minutes by other GPU livedumps
(code 141 VIDEO_ENGINE_TIMEOUT bursts), so the AddDevice failure looks like the failed
recovery, not the original death. WER never uploads these reports, and they never appear
in Reliability Monitor — if you have this problem you will not know unless you look, elevated.

Separately there have been 0x116 VIDEO_TDR_FAILURE bugchecks under/after load — possibly a
second, distinct bug. AMD Crash Defender has also fired once (“display driver now operating
in safe mode”).

Tested / eliminated

  • Motherboard: swapped for an identical new WRX90E — same fault, same dump signature on both boards
  • Third GPU: RX 7900 XTX removed — persists
  • Driver install: full AMD cleanup + clean 26.7.1 — persists (also seen on 26.6.1)
  • One bad card: both R9700s have failed the identical way (timestamp-matched to the dumps)
  • Link/slot/contact: card holds Gen5 x16 with readable config space while faulted
  • Thermal/power: 85 °C hotspot at 299 W under load, benchmarks above spec, weeks of Linux load
  • ULPS/ASPM registry tweaks: did not prevent it (GPU-Z reports ULPS N/A on these cards)
  • BIOS at stock defaults throughout; no WHEA errors ever logged

Possibly related (found searching later)

  • llama.cpp discussion #23443 — R9700 VRAM fully evicted after ~15 s idle on Adrenalin 26.5.2+,
    inbox Windows driver clean. Same card, same “driver does something aggressive at idle” smell.
  • r/AMDHelp thread — headless secondary R9700 getting disabled mid-inference (Crash Defender),
    worked around by switching runtime Vulkan → ROCm.
  • The same WER bucket family (...Status_0xC0000001_Driver_amdkmdag_failed_DdiAddDevice...)
    shows up on RX 580 / 6600 XT / 7700 XT / 7900 XTX systems, often idle/sleep-related.
  • AMD has acknowledged an RDNA4 MES firmware bug causing abnormal idle behavior after compute
    workloads (fix queued on the Linux side; firmware ships with drivers).

Not yet tried

  • Older driver A/B (Adrenalin 26.3.1, the last release before AMD’s idle-power rework, or the
    inbox 32.0.22042.x driver) — machine is being kept stock for an open AMD support case
  • ROCm/HIP backend instead of Vulkan; DP/HDMI dummy plugs (expected to work as mitigation)
  • VBIOS update (ASRock publishes none for this card)

Questions

  1. Anyone else running an R9700 (or dual RDNA4) headless on Windows — have you checked
    C:\Windows\LiveKernelReports (needs admin)? You may have these dumps without knowing.
  2. Anyone on a pre-26.5 driver or the inbox driver with a headless R9700 that stays up?
  3. Anyone with dummy plugs on compute R9700s — stable long-term?
1 Like

I run a 9070xt as primary and R9700 as secondary (no DP/HDMI attached) and I do not run headless on Windows, but, I do notice R9700 is not recognized most times that I boot into windows.

I primarily use Linux for work dev and gaming but I do need windows dual-boot for Battlefield 6 and I use the 9070xt so R9700 not being recognized is a non-issue for me. It is a little peculiar though.

try disabling the thunderbolt/usb4 in bios on the WRX90e?

Here are some things I tried.

Adding a third GPU.
Adding a third, Nvidia GPU.
Moving the cards apart a slot.
Moving the cards apart two slots.
Swapping the cards.
Putting the cards in literally every physically possible position on a WRX80 SAGE motherboard.
Forcing PCIE Gen 3
Forcing PCIE Gen 2
Removing every other PCIE device.
Enabling Rebar + 4G decoding + 41, 42, 43, 44 bit space (ASUS lets you do this)
Disabling Rebar. (This nearly worked actually, but the cards fell apart under load still)
Running Windows in a VM instead of bare metal.
RMA’d one card.
RMA’d the other card.
Tried every driver released for the 9700 (the last six months or so).

I spent two weeks on this. Windows just DOES NOT LIKE these cards. I wish I had an explanation.

2 Likes

@wendell I have disabled it… The system is now stable; I still need to run more tests, however. Even though I have disabled USB4, it appears ReBAR is now disabled no matter the setting in BIOS. I’m going to flash the latest BIOS, as this is a brand new motherboard.

1 Like

make sure CSM is disabled – disabling USB4 should not disable rebar, that’s interesting

2 Likes

Quick follow-up.

I thought it was the motherboard, I mean I guess it could still be….. I disabled CSM, disabled USB4, no change. Swapped in a brand new
ASUS WRX90
: same fault. Last ditch, moved everything over to an ASRock WRX90: same fault again.
All of this on Adrenalin 26.7.1.

Three motherboards across two vendors, identical issue, so I stopped thinking it was the moterboard and put my first first ASUS WRX90 back in.

I rewatched @wendell video to see if I could find any clues to what version of the driver he was running with the 4x R9700s on Windows, so I took a shot in the dark and installed 26.3.1, roughly around the time his video was release

It’s been stable ever since. Neither R9700 has dropped to code 31. Currently sitting at 20+
hours of uptime
with both compute cards headless and idle, which is the condition that used to kill them inside 20 minutes.

At this point I’m waiting for AMD to release a newer driver than 26.7.1 to see if that resolves the issue. I’m also thinking of reaching out to AMD through some of my industry contacts but my pool is small there.

BELOW IS FROM THE ASSISTANCE WITH AI! (Claude)

Working back through the versions

Check the Driver Store version, not the marketing version. It’s what
(Get-CimInstance Win32_VideoController).DriverVersion reports:

Adrenalin Driver Store version Result
26.7.1 32.0.31035.1003 faults
26.6.1 32.0.31019.2002 faults
26.5.1 32.0.31007.1017 hard crash 14 minutes after install
26.3.1 32.0.23033.1002 20+ hours clean, identical config

26.3.1 is the last release of the 230xx branch. 26.5.1 is the first of 310xx. There is no
26.4.x at all, a seven week gap between them. Every 310xx release I’ve tried faults, including the
very first one. The last 230xx release is fine.

Same machine, same board, same BIOS, same headless topology, same USB4 state for all of these. The
only variable that moved was the driver.

What the dumps say

Five full bugchecks in C:\Windows\Minidump, every one naming amdkmdag.sys:

  • 0x116 VIDEO_TDR_FAILURE x3
  • 0xA0000001 x2

That second code is the vendor-defined range, which means AMD’s driver called KeBugCheckEx
itself, deliberately, having decided something was unrecoverable. Windows doesn’t do that on a
driver’s behalf.

The livedumps are VIDEO_MINIPORT_FAILED_LIVEDUMP (0x1b0) with STATUS_UNSUCCESSFUL. That’s the
display miniport failing to bring up a device that is provably healthy: a card sitting in code 31
still reports PCIe x16 5.0 @ x16 5.0 and answers config space reads.

Worth noting which sub-code you get, because it isn’t the same on every version:

  • 26.6.1 and 26.7.1: Arg1 = 1, Add device failed
  • 26.5.1: Arg1 = 2, Start device failed

Same NTSTATUS, one PnP phase apart. On 26.5.1 the driver gets past attaching to the device node and
dies at IRP_MN_START_DEVICE instead. One defect showing up at two adjacent points in the same
device bring-up sequence.

And one of those 26.5.1 dumps fired 11 seconds into a boot. That’s normal boot enumeration, not
idle. So I’d stop calling this an idle bug. It looks more like a device bring-up defect that idle
happens to expose, because idle forces a re-start that boot would otherwise only do once.

Windows or AMD?

Five kernel level records, five naming AMD’s code. It isn’t ambiguous.

The fair qualifier is that it’s AMD’s driver on Windows specifically. The fault lives where their
driver meets dxgkrnl’s PnP and power model. Linux has no dxgkrnl, which is exactly why the same
silicon ran ROCm for weeks without a hiccup. Microsoft owns half the interface, but the code that
fails is AMD’s.

What I suspect changed, and why I can’t prove it

The release notes are no help. 26.5.1 lists two game fixes and nothing else. AMD doesn’t document idle or power management internals.

But look at the driver package:

amdkmdag.sys   112,456,720 bytes

That’s 112 MB, which is not a kernel driver’s worth of code. On Windows AMD embeds the GPU’s own
firmware inside amdkmdag.sys
rather than shipping separate blobs the way Linux does in
linux-firmware/amdgpu. There are no mes_*.bin or smu_*.bin files anywhere in the package because they’re compiled in.

So changing Adrenalin version also changes the microcode loaded onto the card at init. What differs between 26.3.1 and 26.5.1 may not be driver code at all.

The piece I’d look at is MES, the MicroEngine Scheduler, the on-GPU firmware that manages queue scheduling and hands off to power management. AMD has publicly acknowledged an RDNA4 MES firmware bug causing abnormal idle behaviour after compute workloads, with a Linux side workaround and new firmware promised. That firmware ships per driver release. My fault is an RDNA4 card, idle, after compute work, failing to come back. Same shape.

Candidate, not proof. I can’t disassemble firmware out of a 112 MB signed binary and neither can anyone outside AMD.

Workaround

Adrenalin 26.3.1, and block Windows Update from replacing it:
Settings > System > About > Advanced system settings > Hardware > Device Installation Settings > No.

2 Likes

Last update, I upgraded to 26.8.1, I had the same issues. I just switched over to Ubuntu 2604 and made it a dedicated AI workstationfor now. I’ll try troubleshooting this another time. Super frustrating issue.

2 Likes

It’s not just the Code 31 error; there are also issues with AMF decoding on the second card, and the fan speed display constantly reads zero, making it impossible to adjust the speed. None of these issues occur when switching to Ubuntu 26.04.

Super strange issue… I didn’t have the fan speed issue though.

If its a driver + no display glitch, you can use dummy display dongle, and then have windows just ignore the dummy display. The GPU will then behave like it has a display plugged in.