Suggestions for getting started with local AI

Not as much as you might think - the A6000 has 768GB/s VRAM bandwidth, compared with the 3090’s 936GB/s. Basically, 18% slower, which doesn’t really result in 18% slower inference.

In real-world single-card use, the A6000 is slightly slower than the R9700, but it has 48GB. It also CUDA on its side…which may or may not change the picture for multi-card usage.

See what I mean by “It’s complicated”? :smiley:

1 Like

Definitely is complicated!
The price of a used A6000 though seems to be around $3400/3600 which is almost the price of 3x R9700… and that would be 96GB…. not sure which setup would work better though. ROCm has come a long way, and if you are just doing inference, it should be fine, but I haven’t done anything with multi-card yet.

I can tell you that right now I am not impressed by NVFP4 models, other than Nemotron…. I am running ollama cloud and they have a bunch of NVFP4 models and while the speed is faaast; the output often has to be checked with programming. I can usually fix the output by running it through my LOCAL Qwen27B UD-Q8_K_XL or my Q6 40B Finetune version from DavidAU and it will fix all the bugs…
So I think some quality is lost with NVFP4, but its hard to say because ollama is using native FP8 on some of the models as well and its hard to grade the output.

See, this is where it gets interesting.

In my experience, ROCm isn’t as fast as Vulkan (even with tensor split, which Vulkan lacks under llama.cpp) for MoEs, only dense models. Also, out of the models available today, there isn’t really anything in the consumer/workstation world that can run dense models that would use a full 96GB at acceptable speeds for agentic workloads - even the Blackwell RTX 6000 Pro struggles once context fills up.

At least for my uses, it seems like there’s a genuine usefulness gap between 64GB VRAM and 200GB+. There isn’t much that can compete with Qwen 27B or Gemma 31B at Q8 (which fit comfortably within 64GB at max context) until you get to the massive models that currently run at the “It runs, but I can’t do any useful work with it” speeds that YouTubers seem to love to demonstrate (similar to the obsession with massive parallelism in vLLM, where they marvel at the 2000t/s for 512-client concurrency, completely ignoring the fact that each client’s getting about 4 tokens per second and has lost interest long before the answer is complete).

Same here. NVP4 is a marketing gimmick, as far as I’m concerned - most GPUs capable of running decent models at NVFP4 are also capable of running much better quality quants of those models.

It doesn’t help that Nemotron - the poster-child for NVFP4 - was largely just benchmaxxed hallucinatosh :smiley:

1 Like

many nvfp4 quants are trash for 2 reasons.

  1. they just chop down to w4a4.

you CAN in some cases go to w4 on certain layers with the right scale factors. but never quantize activations or kv cache. even on qat’d models its still bleh. just bite the bullet and leave those unquantized. same for any ‘important’ layers in the model.

  1. they do not re-train the model after scooping out 75% of the weights.

using GOOD training datasets to ‘calibrate’ the post-quantized weights is critical. if your dataset sucks and isnt representative of the workload, or you dont do it at all, the precision is in the toilet.

Check out what ngreedia did for this release: nvidia/Qwen3.6-27B-NVFP4 · Hugging Face to make it somewhat tolerable.

Qwen3.6-27B NVFP4: actual precision and layer geometry

Compared checkpoints:

  • Full checkpoint: /nvme/models/safetensors/Qwen3.6-27B
  • Quantized checkpoint: /nvme/models/nvfp/Qwen3.6-27B-NVFP4

Summary

This is the same Qwen3.6-27B multimodal architecture. No layers have been
removed and no logical matrix dimensions change. The NVFP4 checkpoint changes
the storage/compute representation of selected language-model weights. they are going all over the place instead of JUST w4a4.

Property Full checkpoint NVFP4 checkpoint
Tensor payload 51.747 GiB 20.416 GiB
Payload reduction — 2.53x smaller
Tensor dtypes 1,199 BF16 tensors 798 BF16, 401 FP8 E4M3, 193 U8, 802 FP32 scalar tensors
Text decoder 64 layers: 48 linear-attention, 16 full-attention Same
Vision encoder 27 blocks Same, all BF16
Token embedding BF16 [248320, 5120] Same BF16 geometry
Final norm BF16 [5120] Same BF16 geometry
LM head BF16 [248320, 5120] NVFP4 packed U8 [248320, 2560], FP8 scales [248320, 320]

The NVFP4 configuration targets 193 W4A16_NVFP4 matrices: all 192 decoder MLP
matrices (64 layers × gate/up/down) plus lm_head. It also targets 208 FP8
attention/linear-attention projection matrices.

Important: the U8 shapes are packed four-bit storage, not half-width model
matrices. Each byte holds two NVFP4 values, so an original last dimension of
5120 serializes as 2560 U8 values. NVFP4 uses groups of 16 logical weights;
the FP8 scale shape therefore divides that logical axis by 16. Each NVFP4
matrix additionally has scalar FP32 input_scale and weight_scale_2 tensors.

Storage by component

Component Full BF16 payload NVFP4 payload NVFP4 treatment
Text MLPs 32.373 GiB 9.463 GiB W4A16_NVFP4 packed weights plus FP8 group scales
Text attention / linear-attention 13.680 GiB 6.962 GiB Selected projections FP8 E4M3; linear-state tensors BF16
Token embedding 2.368 GiB 2.368 GiB BF16
LM head 2.368 GiB 0.666 GiB W4A16_NVFP4
Vision encoder 0.858 GiB 0.858 GiB BF16
Other text norms 0.099 GiB 0.099 GiB BF16

Common geometry

Every decoder block has BF16 input_layernorm.weight [5120] and
post_attention_layernorm.weight [5120], plus the three MLP projections below.

MLP projection Full logical BF16 shape NVFP4 packed U8 shape FP8 weight_scale shape
gate_proj [17408, 5120] [17408, 2560] [17408, 320]
up_proj [17408, 5120] [17408, 2560] [17408, 320]
down_proj [5120, 17408] [5120, 8704] [5120, 1088]

Linear-attention block geometry

Each linear-attention block has these FP8 E4M3 projections (with scalar FP32
input and weight scales):

Projection Logical shape
linear_attn.in_proj_qkv.weight [10240, 5120]
linear_attn.in_proj_z.weight [6144, 5120]
linear_attn.out_proj.weight [5120, 6144]

Its remaining state/SSM-side tensors remain BF16: A_log [48], dt_bias [48],
conv1d.weight [10240, 1, 4], in_proj_a.weight [48, 5120],
in_proj_b.weight [48, 5120], and norm.weight [128].

Full-attention block geometry

Each full-attention block has these FP8 E4M3 projections (again with scalar
FP32 input and weight scales):

Projection Logical shape
self_attn.q_proj.weight [12288, 5120]
self_attn.k_proj.weight [1024, 5120]
self_attn.v_proj.weight [1024, 5120]
self_attn.o_proj.weight [5120, 6144]

q_norm.weight [256] and k_norm.weight [256] remain BF16.

Layer-by-layer decoder map

Notation: LA = linear-attention block; FA = full-attention block.
Every row includes the shared NVFP4 MLP geometry in the common-geometry table.
LA rows use the FP8 projection geometry in the linear-attention table; FA rows
use the FP8 projection geometry in the full-attention table. All residual
norms are BF16.

Layer Type FP8 portion BF16 portion beyond norms NVFP4 portion
0 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
1 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
2 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
3 FA q, k, v, o q_norm, k_norm gate, up, down
4 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
5 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
6 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
7 FA q, k, v, o q_norm, k_norm gate, up, down
8 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
9 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
10 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
11 FA q, k, v, o q_norm, k_norm gate, up, down
12 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
13 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
14 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
15 FA q, k, v, o q_norm, k_norm gate, up, down
16 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
17 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
18 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
19 FA q, k, v, o q_norm, k_norm gate, up, down
20 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
21 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
22 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
23 FA q, k, v, o q_norm, k_norm gate, up, down
24 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
25 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
26 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
27 FA q, k, v, o q_norm, k_norm gate, up, down
28 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
29 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
30 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
31 FA q, k, v, o q_norm, k_norm gate, up, down
32 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
33 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
34 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
35 FA q, k, v, o q_norm, k_norm gate, up, down
36 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
37 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
38 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
39 FA q, k, v, o q_norm, k_norm gate, up, down
40 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
41 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
42 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
43 FA q, k, v, o q_norm, k_norm gate, up, down
44 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
45 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
46 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
47 FA q, k, v, o q_norm, k_norm gate, up, down
48 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
49 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
50 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
51 FA q, k, v, o q_norm, k_norm gate, up, down
52 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
53 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
54 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
55 FA q, k, v, o q_norm, k_norm gate, up, down
56 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
57 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
58 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
59 FA q, k, v, o q_norm, k_norm gate, up, down
60 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
61 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
62 LA qkv, z, out A_log, dt_bias, conv1d, a/b projections, state norm gate, up, down
63 FA q, k, v, o q_norm, k_norm gate, up, down

Non-decoder portions

The vision tower is not quantized: it has 27 BF16 vision-transformer blocks,
hidden size 1152, 16 heads, MLP width 4304, and a BF16 merger to the text
hidden size 5120. Its major matrix shapes, including the patch embedding,
QKV/projection, vision MLP, and merger, are identical in the two checkpoints.

The NVFP4 index reports 18,164,649,200 stored parameters because its count
tracks serialized packed tensors/scales. It must not be read as a pruned or
18B logical architecture: the logical layer count and dense dimensions match
the full 27B-class checkpoint above.

This image is getting dated but the story is worth repeating:

2 Likes

I will have to test the newest Vulkan again, right now Unsloth Studio seems to be the fastest while using llama.cpp directly; they are doing their own thing with args for each model and all automatically and seem to get the best performance numbers on models vs LM Studio.
I haven’t tried setting LM Studio back to Vulkan in a few months now, but the ROCm that came out in January was the first time I saw ROCm performance that is better than Vulkan at bigger contexts.
Vulkan was still faster by a few tok/sec in small tests but it would totally fall apart after like 75k tokens… which was the opposite originally with ROCm that would often crash past 32k tokens last year…

Which model(s) were you using?

Mostly Qwen3.5 (now 3.6) 27B and some fine Tunes, Qwen 3.5 (now 3.6) 35B-A3B, a bit of Gemma4, LFM2.5, and before those were out; Qwen3-14B-BF16 and GLM4.7 & GLM4.6v-Flash

Interesting. I haven’t found any situation where ROCm is faster than Vulkan for MoE models (at least with RDNA4 and sub-A20B models).

1 Like

Are you on Windows or Linux?

I haven’t tested this for RDNA 4 because I don’t have my R9700 setup yet and I haven’t really used my RX9070 for any inference outside of testing for Unsloth Studio.

Linux (Ubuntu 24.04, for the sake of compatibility).

Ah, I am using Windows on my workstation. I have servers with Linux but none of them have cards yet. That is part of the process that i want to complete this week, but i have been working with the guys from Unsloth to build Studio to where I can shard it out onto multiple systems while retaining the tool functionality on the main endpoint.

If you are thinking of a genuinely useful option, go with an used RTC 3090. It costs somewhere around $500 max. Install llama.cpp and then run the Qwen3.6-35B-A3B.

I would actually skip B70 for now. In local LLM tooling, the Intel Arc support is still inconsistent. In the early stages avoid dual GPUs/ NVLink. It just complicates things.

Where do you see RTX 3090s for $500 max? They are mostly over $1000. The only higher VRAM GPU that is cheaper is a AMD V620.

1 Like