For a different model, on a different backend running directly on host OS, I think I’ve seen 4-5GB/s from btop on my gen4x4 setup when loading the model. That was with one of the R1 quants, and ik_llama.cpp. Sorry about not being able to retest; I do not have ready access to that machine anymore.
The R1 megathread on this forum had info from people trying to run very large models off the SSD, though that use case is more bottlenecked by memory management, and not directly comparable to initial model load on a different backend fully into VRAM.
I expect that different backends handle model loads differently.