vLLM + ROCm + Kappa-20B-131K-MXFP4 on RX 7900 XT (gfx1100) — Triton kernel compile failure

I’m attempting to run Kappa‑20B‑131K‑MXFP4 on an AMD RX 7900 XT (gfx1100) using vLLM nightly with ROCm.

The model appears to load normally and the engine begins initialization, but startup consistently fails once Triton attempts to compile the MXFP4 fused‑MoE kernels for the AMD backend.

This failure occurs across multiple environments (nightly containers, ROCm dev containers, and local source builds), and appears to happen specifically during Triton’s AMD compilation pipeline.

The failure occurs during the Triton AMD lowering stage:

ConvertTritonAMDGPUToLLVM → RuntimeError: PassManager::run failed

This reproduces consistently when:

  • running vLLM nightly containers
  • running ROCm dev containers
  • building vLLM from source

The failure appears specific to the MXFP4 fused‑MoE Triton kernels on the ROCm backend.


Model

https://huggingface.co/eousphoros/kappa-20b-131k-mxfp4

The model card states:

MXFP4 is natively supported in vLLM (nightly)

Example command from the model card:

vllm serve /path/to/kappa_20b_131k_mxfp4 --trust-remote-code

Docker example from the model card:

docker run --gpus all -it --rm \
  -v /path/to/kappa_20b_131k_mxfp4:/model \
  --ipc=host -p 8000:8000 \
  vllm/vllm-openai:nightly --model /model \
  --trust-remote-code \
  --chat-template /model/chat_template.jinja \
  --served-model-name "kappa_20b_131k_mxfp4"

Minimal Reproduction

Running inside AMD ROCm container:

docker run --rm -it \
 --network=host \
 --ipc=host \
 --device /dev/kfd \
 --device /dev/dri \
 --group-add video \
 -e HSA_OVERRIDE_GFX_VERSION=11.0.0 \
 -v /home/joe/models/active/vllm:/model \
 rocm/vllm-dev:nightly \
 vllm serve /model \
 --trust-remote-code \
 --chat-template /model/chat_template.jinja \
 --served-model-name kappa_20b_131k_mxfp4

Expected vs Actual Behavior

Expected

vLLM should successfully initialize the engine and compile the MXFP4 fused‑MoE kernels for the gfx1100 (RX 7900 XT) target.

After compilation, the server should start normally and expose the OpenAI‑compatible API endpoint.

Actual

Engine initialization fails during Triton compilation of the MXFP4 kernels.

The failure occurs in the Triton AMD backend pipeline during:

ConvertTritonAMDGPUToLLVM

The resulting error is:

RuntimeError: PassManager::run failed

Startup begins correctly:

Automatically detected platform rocm
Resolved architecture: GptOssForCausalLM
Using max model len 131072

The failure occurs once Triton begins compiling kernels.


Trimmed Error (Root Cause)

WARNING Using legacy triton_kernels on ROCm

ConvertTritonAMDGPUToLLVM

RuntimeError: PassManager::run failed

The stack trace shows failure inside the MXFP4 MoE Triton kernels:

vllm/model_executor/layers/quantization/mxfp4.py
gpt_oss_triton_kernels_moe.py
matmul_ogs.py
triton/compiler/compiler.py

Specifically during the Triton AMD pipeline stage:

ConvertTritonAMDGPUToLLVM

Hardware

GPU

AMD Radeon RX 7900 XT
Architecture: gfx1100
VRAM: 20GB

CPU

AMD EPYC 7713
122 cores

RAM

377GB system RAM

Operating System

Ubuntu 24.04.3 LTS
Kernel 6.8.0-101-generic

ROCm Runtime

HIP version: 7.0.51831
LLVM: ROCm clang 20

GPU detection:

rocminfo | grep gfx
Name: gfx1100
amdgcn-amd-amdhsa--gfx1100

Containers Tested

ROCm nightly container

rocm/vllm-dev:nightly

Versions:

vLLM: 0.17.0rc1.dev
Triton: 3.4.0

ROCm DSFP4 container

rocm/vllm-dev:dsfp4_0215

Versions:

vLLM: 0.9.2rc2.dev
Triton: 3.4.0

Both containers fail in the same Triton compilation stage.


Observations

Key observations:

  • The crash only occurs when the AMD GPU is exposed to the container.
  • Without GPU device flags, the container starts normally.
  • The model loads successfully before Triton begins kernel compilation.
  • The failure consistently occurs inside Triton’s ROCm backend during kernel lowering.

Additionally Triton prints:

Using legacy triton_kernels on ROCm

which may indicate the newer kernels are not active.


Diagnostics

Triton / Torch environment

docker run --rm rocm/vllm-dev:nightly bash -lc "python - <<'PY'
import torch, triton, os, platform
print('torch:', torch.__version__)
print('triton:', triton.__version__)
print('hip runtime:', torch.version.hip)
print('cuda field:', torch.version.cuda)
print('python:', platform.python_version())
print('ROCM env:', {k:v for k,v in os.environ.items() if 'ROCM' in k or 'HSA' in k or 'HIP' in k})
PY"

GPU detection

docker run --rm rocm/vllm-dev:nightly bash -lc "rocminfo | grep -E 'Name:|gfx'"

Torch environment dump

python -c "import torch; print(torch.utils.collect_env.get_pretty_env_info())"

Suspected Root Cause (Triton AMD backend)

The failure appears to occur during Triton’s AMD compilation pipeline when lowering the MXFP4 fused‑MoE kernels.

The crash consistently happens inside the Triton AMD backend during:

ConvertTritonAMDGPUToLLVM → convertScaledMFMA → PassManager::run failed

The failing kernels originate from the MXFP4 fused‑MoE path used by the GPT‑OSS model family:

vllm/model_executor/layers/fused_moe/gpt_oss_triton_kernels_moe.py
triton_kernels/matmul_ogs.py

Given the assertion failure in MFMA.cpp referencing operand layout requirements, this may indicate an incompatibility between:

• Triton AMD backend
• gfx1100 (RDNA3 / Navi31) target
• MXFP4 scaled‑MFMA kernel generation

This reproduces consistently across:

• vLLM nightly ROCm containers
• AMD ROCm dev containers
• local vLLM source builds

If this is expected behavior for MXFP4 on gfx1100, clarification on current support status would also be helpful.


Questions

  1. Should Kappa‑20B‑131K‑MXFP4 currently work on RDNA3 (gfx1100)?
  2. Is the Triton ROCm backend missing support for this kernel path?
  3. Is a specific Triton / ROCm combination required for MXFP4 models?

Closing

The model itself is extremely interesting — particularly the 131k context length and persona system.

If this failure is a Triton ROCm issue rather than a model compatibility problem, I’m happy to test patches, alternate Triton builds, or debug configurations to help narrow it down.

Try this:

MODEL="eousphoros/kappa-20b-131k"

docker run --rm -it \
 --device=/dev/kfd --device=/dev/dri \
 --group-add video --group-add render \
 --ipc=host --shm-size=16g \
 -p 8000:8000 \
 -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
 -e HF_HOME=/root/.cache/huggingface \
 rocm/vllm-dev:rocm7.2_navi_ubuntu24.04_py3.12_pytorch_2.9_vllm_0.14.0rc0 \
 vllm serve "$MODEL" \
 --host 0.0.0.0 --port 8000 \
 --served-model-name kappa-20b-131k \
 --max-model-len 131072

and see if that works?

Quick follow-up.

I tested the suggested container:

rocm/vllm-dev:rocm7.2_navi_ubuntu24.04_py3.12_pytorch_2.9_vllm_0.14.0rc0

Hardware:
RX 7900 XT (gfx1100, 20GB)
EPYC 7713
Ubuntu 24.04

Test 1 — plain model
eousphoros/kappa-20b-131k

Result:
Server initializes but fails loading with GPU OOM. Lowering max_model_len did not materially change memory usage, so this appears dominated by model weights rather than KV cache.

Test 2 — MXFP4 model
eousphoros/kappa-20b-131k-mxfp4

Results:

  • At max_model_len=131072 → startup fails due to insufficient KV cache memory
    Required: ~3.05 GiB
    Available: ~1.35 GiB
    vLLM estimates max context on this setup at ~56k tokens.

  • At max_model_len=32768 → server starts and responds to requests successfully.

Example request works via the OpenAI endpoint:
curl http://127.0.0.1:8000/v1/chat/completions …

Important note:
Using this newer ROCm 7.2 Navi image, the earlier Triton compile failure

ConvertTritonAMDGPUToLLVM → PassManager::run failed

does NOT reproduce anymore.

So current status on RX 7900 XT:

plain kappa-20b-131k → fails due to VRAM
kappa-20b-131k-mxfp4 → runs successfully up to ~32k context
131k context → fails due to KV cache memory limits on a 20GB card

If there are tuning knobs for KV cache memory or MXFP4 memory usage I should try, I’m happy to test further.

1 Like

it seems as though on 7900xt mxfp4 is being cast to 8 bit, or something like that. 4 bits * 20 billion parameters shouldn’t be 18something gb vram consumed

2 Likes

I have the same hardware, also problems. Fyi there is an rocm container from vllm vllm/vllm-openai-rocm - Docker Image , but this also doesn’t work in my case…

1 Like

Thanks for the directional nudge, Wendell. After some more testing, it no longer looks like this is simply “the MXFP4 model is getting cast to something else.”

What I found instead:

  • The checkpoint config still reports quant_method: mxfp4

  • In the direct Hugging Face Transformers path, Triton was present, but kernels_available was False

  • In that state, a direct AutoModelForCausalLM.from_pretrained(...) load printed the BF16 fallback warning and dequantized during load

The key checks were:


docker exec -i goofy_greider python - <<'PY'

import torch

from transformers.utils.import_utils import is_accelerate_available, is_triton_available

from transformers.utils import is_kernels_available

print("torch.cuda.is_available:", torch.cuda.is_available())

print("accelerate:", is_accelerate_available())

print("triton>=3.5:", is_triton_available("3.5.0"))

print("kernels_available:", is_kernels_available())

PY

That came back as effectively:


cuda=True, accelerate=True, triton=True, kernels_available=False

And the direct Transformers load was showing the BF16 fallback warning:


MXFP4 quantization requires Triton and kernels installed: CUDA requires Triton >= 3.4.0, XPU requires Triton >= 3.5.0, we will default to dequantizing the model to bf16

At that point I tried the narrowest possible fix, only on the HF kernels side:


docker exec goofy_greider python -m pip install --no-cache-dir -U kernels


docker exec -i goofy_greider python - <<'PY'

from huggingface_hub import snapshot_download

snapshot_download("kernels-community/triton_kernels")

PY

Then I re-ran the gate check and confirmed:


kernels_available: True

After that, the BF16 fallback warning disappeared, and direct Transformers MXFP4 GPU load succeeded:


VIDEO_GID=$(getent group video | cut -d: -f3)

RENDER_GID=$(getent group render | cut -d: -f3)

docker run --rm -i \

--device=/dev/kfd --device=/dev/dri \

--group-add "$VIDEO_GID" \

--group-add "$RENDER_GID" \

--ipc=host \

rocm-vllm-mxfp4-kernels:try1 python - <<'PY'

from transformers import AutoModelForCausalLM

model_id = "eousphoros/kappa-20b-131k-mxfp4"

model = AutoModelForCausalLM.from_pretrained(

model_id,

trust_remote_code=True,

device_map="cuda"

)

print("MODEL LOAD FINISHED")

PY

That completed successfully.

So at that point the first HF-side blocker appeared fixed.

Current state now:

  • direct Transformers MXFP4 load works on the 7900 XT after adding the missing kernels dependency

  • direct Transformers MXFP4 generation still fails later, now inside the downloaded community Triton kernels path, specifically matmul_ogs.py

The next-step generation test was:


VIDEO_GID=$(getent group video | cut -d: -f3)

RENDER_GID=$(getent group render | cut -d: -f3)

docker run --rm -i \

--device=/dev/kfd --device=/dev/dri \

--group-add "$VIDEO_GID" \

--group-add "$RENDER_GID" \

--ipc=host \

rocm-vllm-mxfp4-kernels:try1 python - <<'PY'

import torch

from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "eousphoros/kappa-20b-131k-mxfp4"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)

model = AutoModelForCausalLM.from_pretrained(

model_id,

trust_remote_code=True,

device_map="cuda"

)

inputs = tokenizer("Reply with exactly: MXFP4 works", return_tensors="pt").to("cuda")

with torch.no_grad():

out = model.generate(**inputs, max_new_tokens=12)

print(tokenizer.decode(out[0], skip_special_tokens=True))

PY

That now fails later in the kernel path with:


TypeError: unsupported operand type(s) for *: 'NoneType' and 'int'

coming from:


target_info.num_sms()

inside:


kernels-community/triton_kernels/.../matmul_ogs.py

So at this point the issue looks narrower than:

  • “ROCm can’t do MXFP4”

  • or “the model is just getting silently widened to 8/16-bit”

What it currently looks like is:

  • vLLM on the newer ROCm 7.2 Navi image can already run the model at 32k context on this card

  • the Hugging Face fallback was due to a missing kernels dependency, not missing Triton broadly

  • after fixing that, MXFP4 GPU load works in the direct Transformers path

  • the remaining bug now appears to be in the kernels-community/triton_kernels inference path on this ROCm target, specifically target-info / device-metadata handling during kernel execution

TL;DR:

  • first HF-side blocker was missing kernels

  • fixing that removed the BF16 fallback and allowed native MXFP4 GPU load

  • the remaining failure is now in kernels-community/triton_kernels during inference on this ROCm target, not the original fallback gate

Would it help to open an issue on the vllm github? I read a little bit online about the gpt-oss20b model and vllm/rocm and it sounds rather rough in terms of software support

1 Like

That was my read too, so I’m holding off on opening a fresh issue until I finish validating the ROCm/gfx11xx work that’s already been identified on the vLLM side.

I did get past the original compile failure with one more container tweak, but that still didn’t look like the intended MXFP4 result. The model would start, but the VRAM footprint stayed much heavier than expected, so it didn’t look like a proper lean 4-bit path.

From there I went digging through the existing vLLM issues and PRs to see whether this had already been reported, and that’s how I found what looks like the same ROCm/gfx11xx trail already sitting in their backlog. An older issue looked adjacent (#26303) [Feature]: `Mxfp4MoEMethod` support on ROCm · Issue #26303 · vllm-project/vllm · GitHub, but the more relevant one for this seems to be (#33906) [Bug]: mxfp4 (gpt-oss moe) on AMD rocm (W7900/gfx1100) breaks · Issue #33906 · vllm-project/vllm · GitHub and the associated Triton-branch update.

I’m rebuilding and validating that locally on my 7900 XT now. With luck that’s the real fix. If not, I should at least come out of it with a much cleaner repro and enough evidence to open or update an issue usefully.

1 Like

Triton rocm support isn’t great (understating it), I have direct contact with the vendors and it’s still been an issue. My next steps would be a solid repro and raising this with the vllm and maybe even the triton teams, AMD won’t help you. I can’t see anything problematic with what you most recently posted. VLLM tries their best but ROCM can be a mess. Going off my prior experiences it’s probably a ROCM api or function that shifted under them without any comms. I won’t even get into all the issues with ROCM on MI300As because I might pull my hair out.

1 Like

Quick update.

I went ahead and validated the merged vLLM Triton/base-image fix locally on the same RX 7900 XT (gfx1100) box.

Result:

  • it does help the older gfx11xx ROCm/Triton side enough for the rebuilt image to build and serve successfully
  • but it does not fix the remaining behavior I’m seeing with eousphoros/kappa-20b-131k-mxfp4

After rebuilding from the merged fix, Kappa still loads heavy on my 7900 XT:

  • Model loading took 14.3 GiB memory
  • with max_model_len=131072, runtime still only reports about GPU KV cache size: 29968 tokens

More importantly, after making the relevant vLLM warnings loud in the logs, the current ROCm path is still explicitly reporting:

  • MXFP4 linear layer is not implemented - falling back to UnquantizedLinearMethod
  • MXFP4 attention layer is not implemented. Skipping quantization

So my current read is:

  • the merged Triton/base-image fix was real and worth validating
  • but it does not resolve the remaining Kappa / gpt-oss MXFP4 behavior on ROCm/gfx1100
  • the current heavy memory footprint appears to come from partial MXFP4 support in the active vLLM path, not just the older Triton branch issue

So, not back to square one, but also not “fixed.” At least the problem space is narrower now.

So to try things I converted the model to gguf to use it with llamacpp. This seems to work, but I can’t setup personas. I can not comment on performance or any downsides to that approch …

1 Like

Good idea, I may follow suit when I get feeling better and back on this.

After a lot more digging, this looks less like a simple container/branch bug and more like a missing ROCm feature path in vLLM for this specific GPT-OSS/Kappa MXFP4 setup.

Kappa’s MXFP4 path is a pretty specific one: the quantized part is on the expert MLP / MoE side, and the training writeup explicitly calls out a Triton matmul_ogs path chosen because GPT-OSS uses a hidden size of 2880 that does not play nicely with some of the other FP4 routes. In other words, this is not just a generic “4-bit model should work everywhere” case. 

At this point I think the fix, if/when it happens, needs to land upstream in vLLM/ROCm rather than in more local container fiddling. There is even an open vLLM ROCm feature request specifically around Mxfp4MoEMethod support, which lines up pretty well with what I’ve been seeing in practice. 

On the plus side, the model itself does not seem to be the problem. Like Greedence, I was able to convert Kappa to GGUF pretty easily and get it running outside this stack, so the weights/model look healthy. I’m still testing behavior and parity, but the MoE structure and the mixed tensor story appear to have come across. 

I’m going to park the vLLM/ROCm angle here for now. If someone upstream wires in the missing path, I’d be very interested in testing again.

No. For AMD, MXFP4 is only supported in hardware starting with CDNA-4 (MI350 series) GPUs.

gfx1100 - RDNA 3 - int-8, bf16, fp16.
gfx115x - RDNA 3.5 - int-8, bf16, fp16.
gfx12xx - RDNA 4 - fp8, int-8, bf16, fp16.

If something is able to load the model on gfx1100, it’s likely upscaling it in software. I think you’re seeing that. There are no 4-bit options on gfx1100. The only 8-bit option is int-8.

My GPUs are gfx1100. So I’d be thrilled to find out I’m wrong about this. All these new models working in 4-bits… it’s frustrating. My current strategy is to create my own Q8_0 from the original bf16 model.

I’m going to try to follow through what you did. Even if it does need to dequantize to a larger type, so long as it still fits and runs, that’s a win.

I pulled that custom docker, and added in the missing parts in a new Dockerfile:

FROM rocm/vllm-dev:rocm7.2_navi_ubuntu24.04_py3.12_pytorch_2.9_vllm_0.14.0rc0
RUN pip install --no-cache-dir -U kernels
RUN python -c "from huggingface_hub import snapshot_download; snapshot_download('kernels-community/triton_kernels')"
RUN python -c "from huggingface_hub import snapshot_download; snapshot_download('eousphoros/kappa-20b-131k-mxfp4')"

I’m using this command to try running it:
podman run --rm -it --device=/dev/kfd --device=/dev/dri --group-add keep-groups --security-opt label=disable --ipc=host -p 8000:8000 -v “$HOME/.cache/vllm:/root/.cache/vllm” -e HSA_OVERRIDE_GFX_VERSION=11.0.0 vllm-mxfp4-navi vllm serve eousphoros/kappa-20b-131k-mxfp4 --host 0.0.0.0 --port 8000 --served-model-name kappa-20b-131k --tensor-parallel-size 4 --max-model-len 131072 --trust-remote-code

And that has kicked off a long compile step… should take 3 to 4 hours:

Capturing CUDA graphs (mixed prefill-decode, PIECEWISE):  33%|██████████████████████████████████▍                                                                       | 27/83 [1:11:05<2:28:11, 158.78s/it](EngineCore_DP0 pid=429) INFO 03-22 20:41:22 [shm_broadcast.py:542] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
(EngineCore_DP0 pid=429) INFO 03-22 20:42:22 [shm_broadcast.py:542] No available shared memory broadcast block found in 60 seconds. This typically happens when some processes are hanging or doing some time-consuming work (e.g. compilation, weight/kv cache quantization).
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE):  34%|███████████████████████████████████▊                                                                      | 28/83 [1:13:44<2:25:23, 158.61s/it]

I’ll follow-up later

1 Like