MSI X870E Carbon / 9950X / 2x48GB - RAM bizzare issues

DDR5 module detected as “2GB ghost DIMM” after S3 sleep on AM5: root cause found (SPD hub stuck in 2-byte addressing) and an in-band fix, no power cycle needed

Board: MSI MPG X870E Carbon WiFi (MS-7E49), BIOS 1.AA3 (2026-06-26, latest)
CPU: Ryzen 9 9950X3D
RAM: G.Skill F5-6400J3239G32GX2-TZ5RS (2x32GB DDR5-6400 CL32, SK Hynix A-die), XMP + manual 6000 MT/s, DIMMA2 + DIMMB2. Stable for months, Prime95-clean.
OS: dual boot, Ubuntu 26.04 (kernel 7.0.0-28) / Windows 11 25H2. The OS turns out to be irrelevant, see below.

This is the same failure discussed in [level1techs thread 229940] and
[MSI forum / “bizarre RAM issues” thread 222454]: after suspend-to-RAM,
one DIMM turns into an unknown 2GB module (“Devices Changed (CPU or Memory)
or CMOS have been Cleared” at POST), the system runs with 34GB instead of
64GB, warm reboots don’t help, and only a full power-off cures it. It hits
MSI, ASRock and Gigabyte AM5 boards, so it smells like AGESA rather than one
vendor’s UEFI code. I spent a few days instrumenting my machine and I believe
I can now say precisely what breaks, where, and how to fix it in-band. Long
post, but every claim below is reproducible with i2c-tools and I kept the
raw dumps.

TL;DR

  • The SPD hub (SPD5118, Montage silicon) of one DIMM gets its MR11 register
    (I2C legacy mode, address 0x0B) written to 0x08 = “2-byte addressing mode”.
    This is volatile state inside the hub, powered from the standby rail, so it
    survives warm reboots and S3, and only VDDSPD power loss resets it.

  • A hub in 2-byte mode ignores 1-byte page selects and serves EEPROM page 0
    for every read. All legacy SPD readers (BIOS SMBIOS code, CPU-Z, HWiNFO,
    Linux spd5118 driver) then decode page 0 as if it were the whole SPD: the
    result is a deterministic “2GB, 1 rank, serial 00206200, no part number”
    ghost. The EEPROM itself is intact; nothing to RMA.

  • The bit is written by firmware, not by the OS: I caught it being set during
    a POST with Linux completely idle (details below). The S3 resume path is
    just the most frequent trigger, matching everyone’s reports.

  • Disassembly of both the AMD ABL (PSP-side memory init) and MSI’s
    MsiOcMemSPDPei module shows the same defect pattern in two independent
    implementations: the SPD page number is computed as offset >> 7 with no
    “& 7” mask and (in 2 of 3 ABL write paths) no bounds check, and is then
    written to MR11. For any offset >= 0x400 the value written is exactly 0x08:
    bit 3 lands in the legacy-mode bit and latches the hub into 2-byte mode.

  • Fix without power cycle, from Linux (this is, as far as I know, the first
    documented in-band recovery):

    # find the stuck hub: 0x50-0x53 on the piix4 SMBus bus, one per DIMM
    sudo i2cget -y 1 0x53 0x0b b     # returns 0x08 -> stuck in 2-byte mode
    sudo i2cset -y 1 0x53 0x0b 0x0000 w
    sudo i2cget -y 1 0x53 0x0b b     # must now return 0x00
    # warm reboot; BIOS notices the "changed" DIMM, retrains, full capacity is back
    

    The write-word trick works because for a hub in 2-byte mode the SMBus
    sequence [0x0B, 0x00, 0x00] parses as “write 0x00 to MR 0x000B”. Do NOT
    enable PEC (no ‘p’ suffix on the mode): the extra PEC byte would be
    consumed as data and auto-increment into MR12.

Symptom and history

Trigger sequence on my machine: hours in S3 under Ubuntu, resume, then a warm
reboot within minutes. Sometimes the next POST is normal and fast; sometimes
it takes much longer, ends in the “Devices Changed” prompt, and from then on
BIOS Memory-Z shows 34816 MB with DIMMB2 as an unknown 2GB module. The state
survives any number of warm reboots and OS switches. This has been happening
since I built the machine in March 2025, through every BIOS MSI has released
since, up to and including the current 1.AA3 - so it is not a recent AGESA
regression but a long-standing defect. And it is not the modules: this same
kit previously spent about 16 months (late 2023 to March 2025) in an Intel
system (i9-14900K, ASUS Z790 Hero) without a single incident, same DIMMs,
same Montage SPD hubs. It is intermittent: the
journal shows an identical 19.5h-S3 + resume + reboot sequence on Jul 17 that
did NOT trigger it.

Notable: the broken system still trains both DIMMs at 6000 MT/s and runs
stable, because with Memory Context Restore the training data comes from
saved context. Capacity, SMBIOS and the SPD contents are what go wrong.

Diagnosis

In the broken state, Linux says:

spd5118 1-0053: Adapter does not support 16-bit register addresses

That driver message means: MR11 bit 3 (SPD5118_LEGACY_MODE_ADDR) reads as 1
and the SMBus controller (piix4) can’t do 2-byte addressing, so the driver
gives up. Reading the hub raw confirms it:

$ sudo i2cget -y 1 0x53 0x0b b
0x08                              # bit3=1: 2-byte mode, page bits = 0

MR0/MR1 still identify a healthy SPD5118 (0x51/0x18, vendor 0x06:0x32 =
Montage, same silicon as the twin module). A 1-byte i2cdump of the stuck hub
shows the mechanism directly: the lower half (0x00-0x7F) returns the MR
registers, the upper half returns EEPROM page 0 - and page 0 is what every
page-select-blind reader then decodes as the entire 1024-byte SPD. That
produces the famous ghost, byte for byte:

Ghost value reported Actual origin (SPD page 0)
Manufacturing location “18” byte 2 = 0x12
Date “year 2002, week 4” bytes 3-4 = 02 04
Serial 00206200 bytes 5-8 = 00 20 62 00
Part number empty bytes 9+ = 00 00 …
“2 GB”, 1 rank, x8 density/org bytes read from wrong page
DDR5-4800, JEDEC timings intact page 0 really does hold base JEDEC data

Same ghost in Windows (CPU-Z, HWiNFO, Task Manager) and Linux (dmidecode),
because everyone is reading the same stuck hub live at each POST.

Per the SPD5118 datasheet, MR11 has no in-band reset: Bus Clear and Bus Reset
preserve registers, and only VDDSPD < 0.3V for >= 1ms clears it. That is
exactly why only power-off cures it and why it survives S3 (DIMM standby
power stays up) and warm reboots.

Who writes the bit? Firmware, with the OS ruled out

First I audited the Linux spd5118 driver (the only OS-side MR11 writer, it
uses MR11 for page selects and rewrites it via regcache_sync on resume): all
its writes mask the page to bits [2:0] or preserve bit 3, so there is no code
path that can set the legacy-mode bit. But the decisive evidence came from an
accident while testing the fix:

  1. Jul 23, 20:44:11 - I cured hub 0x53 (write MR11=0), verified 0x00.
  2. 20:44:49 - clean shutdown for a warm reboot. 38 seconds, no suspend, the
    driver idle the whole time.
  3. The next POST was visibly long (BIOS saw the “changed” DIMM, invalidated
    the memory context, re-read the SPDs and retrained), and came up with…
    34GB again. But now with the roles swapped: 0x53 probed fine and 0x51 -
    the DIMM that had been healthy for months - was stuck with MR11=0x08.

So the bit was planted by firmware during a POST, with no S3 and no OS in the
loop. Linux is exonerated as the writer; S3 resume is just the most common
window where the vulnerable firmware path runs. (I then cured 0x51 the same
way, rebooted again, and got 64GB back, both hubs healthy - two in-band
recoveries, zero power cycles.)

Where in the firmware (disassembly of BIOS 1.AA3)

I extracted the AMD ABL from the PSP directory (entry type 0x30, zlib
streams) and MSI’s SPD-related UEFI modules from E7E49AMSI.1AA3.

ABL (ARM, runs DDR init incl. S3 resume): the SPD access routine at
offset 0x765f0 of the decompressed ABL computes the page as
ubfx(offset, 7, 8) - an 8-bit field, no “& 7” - and writes it to SPD hub
register 0x0B over its DesignWare-style I2C master. That page value reaches
the MR11 write through three paths; two of them (0x7667a-0x766b0, the first
one silent with no status check, and 0x766be-0x766d6) have no bounds check,
while the third (0x7672e) does have a cmp 0x400 / bls guard - so the authors
knew about the boundary and guarded only one of three sites. For an offset in
[0x400, 0x47F] the computed page is exactly 0x08. Related strings in the same
code: “Cannot write SPD Hub page registers for DIMM at 0x%x”,
“Mem CtxSaveRestore enabled, skip spd read”, “ABL General S3 Init”.

MsiOcMemSPDPei (x86 PEI, runs every POST to feed SMBIOS/Memory-Z): same
defect, independent implementation - shr ebx, 7 with no and bl, 7 (module
offset 0x698), feeding an EfiSmbusWriteByte helper with Command = 0x0B
(helper at 0x235e, several callers with the same pattern).

I have not traced the exact call chain that produces an offset >= 0x400 in
the field (that part remains a hypothesis), but the arithmetic is not in
question: two unrelated SPD drivers in this firmware compute an unmasked page
from a byte offset and write it into the one register where bit 3 flips the
hub into a different addressing mode. Nothing anywhere in the firmware
strings or code suggests 2-byte mode is ever used intentionally.

For what it’s worth, the Linux driver contains an independent observation of
the aftermath (drivers/hwmon/spd5118.c, commit a852162efbff): “some PC BIOS
versions will not change the addressing mode on a soft reboot”, written after
someone hit hubs left in 2-byte mode. And the driver’s own 16-bit support is
being removed upstream in 2026 (“testing was limited and there are no known
users”), so future kernels won’t even limp along with a stuck hub - one more
reason to get this fixed at the source.

How to check your own system (Linux, 30 seconds)

sudo apt install i2c-tools
dmesg | grep spd5118        # "Adapter does not support 16-bit register
                            # addresses" on 1-005X = that DIMM's hub is stuck
sudo i2cget -y 1 0x5X 0x0b b   # 0x08 = stuck; 0x00-0x07 = normal
# (the stuck hub is never bound to the driver, so i2cget just works;
#  for a healthy hub you'd need modprobe -r spd5118 first)

Then the two-command fix from the TL;DR, then reboot. Expect one long POST
with a possible “Devices Changed” prompt - that’s the BIOS retraining with
the real SPD data again.

What I’d ask MSI / AMD

  • Audit every SPD hub page-select write for the missing mask/bounds check
    (offsets above; happy to share the full analysis and dumps).
  • Defensively write MR11 = 0x00 for each DIMM early in POST before the
    SMBIOS/Memory-Z SPD reads. That single write would make the failure
    self-healing on the next boot instead of persisting for weeks.
  • If the unguarded paths are in the ABL, this needs to go to AMD - the
    multi-vendor reports (MSI/ASRock/Gigabyte) point that way.