I’m attempting to run Kappa‑20B‑131K‑MXFP4 on an AMD RX 7900 XT (gfx1100) using vLLM nightly with ROCm.
The model appears to load normally and the engine begins initialization, but startup consistently fails once Triton attempts to compile the MXFP4 fused‑MoE kernels for the AMD backend.
This failure occurs across multiple environments (nightly containers, ROCm dev containers, and local source builds), and appears to happen specifically during Triton’s AMD compilation pipeline.
The failure occurs during the Triton AMD lowering stage:
ConvertTritonAMDGPUToLLVM → RuntimeError: PassManager::run failed
This reproduces consistently when:
- running vLLM nightly containers
- running ROCm dev containers
- building vLLM from source
The failure appears specific to the MXFP4 fused‑MoE Triton kernels on the ROCm backend.
Model
https://huggingface.co/eousphoros/kappa-20b-131k-mxfp4
The model card states:
MXFP4 is natively supported in vLLM (nightly)
Example command from the model card:
vllm serve /path/to/kappa_20b_131k_mxfp4 --trust-remote-code
Docker example from the model card:
docker run --gpus all -it --rm \
-v /path/to/kappa_20b_131k_mxfp4:/model \
--ipc=host -p 8000:8000 \
vllm/vllm-openai:nightly --model /model \
--trust-remote-code \
--chat-template /model/chat_template.jinja \
--served-model-name "kappa_20b_131k_mxfp4"
Minimal Reproduction
Running inside AMD ROCm container:
docker run --rm -it \
--network=host \
--ipc=host \
--device /dev/kfd \
--device /dev/dri \
--group-add video \
-e HSA_OVERRIDE_GFX_VERSION=11.0.0 \
-v /home/joe/models/active/vllm:/model \
rocm/vllm-dev:nightly \
vllm serve /model \
--trust-remote-code \
--chat-template /model/chat_template.jinja \
--served-model-name kappa_20b_131k_mxfp4
Expected vs Actual Behavior
Expected
vLLM should successfully initialize the engine and compile the MXFP4 fused‑MoE kernels for the gfx1100 (RX 7900 XT) target.
After compilation, the server should start normally and expose the OpenAI‑compatible API endpoint.
Actual
Engine initialization fails during Triton compilation of the MXFP4 kernels.
The failure occurs in the Triton AMD backend pipeline during:
ConvertTritonAMDGPUToLLVM
The resulting error is:
RuntimeError: PassManager::run failed
Startup begins correctly:
Automatically detected platform rocm
Resolved architecture: GptOssForCausalLM
Using max model len 131072
The failure occurs once Triton begins compiling kernels.
Trimmed Error (Root Cause)
WARNING Using legacy triton_kernels on ROCm
ConvertTritonAMDGPUToLLVM
RuntimeError: PassManager::run failed
The stack trace shows failure inside the MXFP4 MoE Triton kernels:
vllm/model_executor/layers/quantization/mxfp4.py
gpt_oss_triton_kernels_moe.py
matmul_ogs.py
triton/compiler/compiler.py
Specifically during the Triton AMD pipeline stage:
ConvertTritonAMDGPUToLLVM
Hardware
GPU
AMD Radeon RX 7900 XT
Architecture: gfx1100
VRAM: 20GB
CPU
AMD EPYC 7713
122 cores
RAM
377GB system RAM
Operating System
Ubuntu 24.04.3 LTS
Kernel 6.8.0-101-generic
ROCm Runtime
HIP version: 7.0.51831
LLVM: ROCm clang 20
GPU detection:
rocminfo | grep gfx
Name: gfx1100
amdgcn-amd-amdhsa--gfx1100
Containers Tested
ROCm nightly container
rocm/vllm-dev:nightly
Versions:
vLLM: 0.17.0rc1.dev
Triton: 3.4.0
ROCm DSFP4 container
rocm/vllm-dev:dsfp4_0215
Versions:
vLLM: 0.9.2rc2.dev
Triton: 3.4.0
Both containers fail in the same Triton compilation stage.
Observations
Key observations:
- The crash only occurs when the AMD GPU is exposed to the container.
- Without GPU device flags, the container starts normally.
- The model loads successfully before Triton begins kernel compilation.
- The failure consistently occurs inside Triton’s ROCm backend during kernel lowering.
Additionally Triton prints:
Using legacy triton_kernels on ROCm
which may indicate the newer kernels are not active.
Diagnostics
Triton / Torch environment
docker run --rm rocm/vllm-dev:nightly bash -lc "python - <<'PY'
import torch, triton, os, platform
print('torch:', torch.__version__)
print('triton:', triton.__version__)
print('hip runtime:', torch.version.hip)
print('cuda field:', torch.version.cuda)
print('python:', platform.python_version())
print('ROCM env:', {k:v for k,v in os.environ.items() if 'ROCM' in k or 'HSA' in k or 'HIP' in k})
PY"
GPU detection
docker run --rm rocm/vllm-dev:nightly bash -lc "rocminfo | grep -E 'Name:|gfx'"
Torch environment dump
python -c "import torch; print(torch.utils.collect_env.get_pretty_env_info())"
Suspected Root Cause (Triton AMD backend)
The failure appears to occur during Triton’s AMD compilation pipeline when lowering the MXFP4 fused‑MoE kernels.
The crash consistently happens inside the Triton AMD backend during:
ConvertTritonAMDGPUToLLVM → convertScaledMFMA → PassManager::run failed
The failing kernels originate from the MXFP4 fused‑MoE path used by the GPT‑OSS model family:
vllm/model_executor/layers/fused_moe/gpt_oss_triton_kernels_moe.py
triton_kernels/matmul_ogs.py
Given the assertion failure in MFMA.cpp referencing operand layout requirements, this may indicate an incompatibility between:
• Triton AMD backend
• gfx1100 (RDNA3 / Navi31) target
• MXFP4 scaled‑MFMA kernel generation
This reproduces consistently across:
• vLLM nightly ROCm containers
• AMD ROCm dev containers
• local vLLM source builds
If this is expected behavior for MXFP4 on gfx1100, clarification on current support status would also be helpful.
Questions
- Should Kappa‑20B‑131K‑MXFP4 currently work on RDNA3 (gfx1100)?
- Is the Triton ROCm backend missing support for this kernel path?
- Is a specific Triton / ROCm combination required for MXFP4 models?
Closing
The model itself is extremely interesting — particularly the 131k context length and persona system.
If this failure is a Triton ROCm issue rather than a model compatibility problem, I’m happy to test patches, alternate Triton builds, or debug configurations to help narrow it down.