Poor(er) man's local AI

a man in a suit and tie is applauding .

Excellent work.

There are 3 different places to quantize, with 3 different impacts. Weights, activations, and KV cache. Weights you can quantize down quite nicely to 8 bit, sometimes even mixed precision models where some layers are 4, some are 8, etc.

But do not quantize activations, or your kv cache. Those are the two biggest killers of precision. Measuring the impact of this is tricky and requires real datasets and benchmarking, but I have completely abandoned trying to quantize kv cache and activations as they often turn out to be far more detrimental in the long run.

Also, I look forward to your next insane build where you go for a TP16 setup:

3 Likes