Intel B70 Launch - Unboxed and Tested

Skimmed the benchmarks a bit, you tried any legacy quants? Q4_0/Q8_0 quants run much faster on some older GPUs (TensorRT1.0 cards, older AMD etc) than the Qx_K counterparts. I’ve heard the same thing might be true for Arc, but haven’t really tested it.

IIRC the Q8_0 models ran faster than the Q4_K models on the Tesla P40 at least, might be worth a test if you’re running iGPU/NPU inference with GGUF quants.