No. For AMD, MXFP4 is only supported in hardware starting with CDNA-4 (MI350 series) GPUs.
gfx1100 - RDNA 3 - int-8, bf16, fp16.
gfx115x - RDNA 3.5 - int-8, bf16, fp16.
gfx12xx - RDNA 4 - fp8, int-8, bf16, fp16.
If something is able to load the model on gfx1100, it’s likely upscaling it in software. I think you’re seeing that. There are no 4-bit options on gfx1100. The only 8-bit option is int-8.
My GPUs are gfx1100. So I’d be thrilled to find out I’m wrong about this. All these new models working in 4-bits… it’s frustrating. My current strategy is to create my own Q8_0 from the original bf16 model.
I’m going to try to follow through what you did. Even if it does need to dequantize to a larger type, so long as it still fits and runs, that’s a win.