Guide + prebuilt fix: ROCm math libraries segfault on gfx906 with ECC disabled (Vega II / Radeon VII / some MI50s)

This is for anyone running the budget-compute favorites — MI50s, Radeon VIIs, or Mac Pro 2019 Vega IIs under Linux — who’s hit the maddening situation where the GPU looks perfectly healthy (rocm-smi fine, rocminfo fine, your own HIP kernels compile and run fine) but anything touching rocBLAS, rocSPARSE, or rocALUTION segfaults instantly, which also takes out everything built on top of them (hipBLAS/hipSPARSE, PETSc-HIP, sparse ops generally).

Root cause: gfx906 exists in two ISA feature variants — sramecc+ (SRAM ECC enabled, the typical datacenter MI50/MI60 config) and sramecc- (ECC disabled, which is what Vega IIs, Radeon VIIs, and some MI50s report). Recent ROCm releases ship prebuilt math-library kernels only for sramecc+. When the HIP runtime can’t find a code object matching your card’s exact feature string, it doesn’t return an error — it segfaults. Your own hipcc-compiled kernels work because they’re compiled for exactly what’s in the machine, which is what makes this so confusing to diagnose: half the stack works, half crashes, and nothing tells you why.

Check your card:


rocminfo | grep -o 'amdgcn-amd-amdhsa--gfx906[^ ]*' | head -1
```

If it prints `gfx906:sramecc-:xnack-` and you have the segfault signature above, this is your bug. With `AMD_LOG_LEVEL=3` you'll see `No compatible code objects found for: gfx906:sramecc-:xnack-`.

Two things that *don't* work, to save you the attempts: `HSA_OVERRIDE_GFX_VERSION` (the mismatch is the sramecc flag, not the version number) and reinstalling/upgrading ROCm (the new packages have the same gap).

**What does work:** rebuilding the three foundational libraries from AMD's MIT-licensed sources with the target pinned to your card's actual ISA, installed to a user-local prefix. No root at any point, and your system ROCm is untouched — the fix is purely LD_LIBRARY_PATH ordering, so uninstalling is rm -rf one directory.

I've packaged the whole thing: https://github.com/Intermountainh8ter/rocm-gfx906-sramecc-fix

- **Option A**: prebuilt release tarball (built against ROCm 7.2.3) — extract to ~/local, export one LD_LIBRARY_PATH line, run the included validation script
- **Option B**: numbered build scripts (01_clone through 05_validate) that clone AMD's sources at the matching release tag and build in the right dependency order with the right cmake wiring (rocSPARSE needs to find *your* rocBLAS, not the broken system one — the scripts handle that). Budget 2-4 hours; Tensile's kernel generation for rocBLAS is nearly all of it.
- The validation suite is the actual diagnostic toolkit I used, from a minimal HIP hello through the exact one-file segfault reproducer up to a full preconditioned CG solve on the GPU — so you get a definitive yes/no on your own card, not just "it installed."

**Performance, honestly stated:** on 2x Vega II vs a dual-Xeon 56-core host, the GPU path wins from ~1-2M DOF upward (~4.4x at 4M DOF on a Jacobi-CG Laplacian solve, double precision, results bit-identical to a scipy reference). Below ~1M DOF the CPU wins — these are memory-bandwidth-bound problems and a dual-socket Xeon has a lot of bandwidth. So this is worth it for real sparse-solver workloads, not for making small problems faster.

Licensing: it's all AMD's own MIT-licensed code, rebuilt; license texts ship in the repo and tarball. Unofficial, not AMD-affiliated.

Happy to answer questions if anyone hits variant weirdness on other gfx906 cards — the scripts take the ISA string as a variable, so the same approach should extend.

Wow, I thank you for this as a person who is collecting some of this hardware right now.

you can reflash your GPU (check sramecc on different VBIOS)