Yeah, I am probably going to strike a balance between context and quantization. Right now for testing on the CPU I have it in 4bit in order to speed things up, but once the GPU arrives, I will likely run it at 6bit or 8bit. BF16 seems like it might be overkill, and just about indistinguishable from the 8bit, and only slightly different from 6bit.
At least that is the impression I get.
Thanks for the suggestion. I will have to look into that.
It is just a case of experimenting but the general idea is that the smaller the model is in parameters, the less you want it compressed with quantisation. But there is also no point to go BF16 if it leaves you with a KV cache that can’t fit anything of use.
calling gemma tiny is pretty funny, it’s not wrong but even the 31B at FP8 is 32gb of weights and once you add 256k of context, some multi-modal limits, context caching you’re suddenly around 80-96gb of vram, more with bf16 weights
kind of the beauty of it, you can run gemma on actually responsible hardware or scale it up to almost silly numbers for a 31B.
I’ve been playing with Gemma4:26b for the last couple of days, and I have to agree. It is both quite a bit better and significantly faster than Llama 3.3 70B.
It is a little chatty though, which drives token use quite high, but I kind of like it, as I can get a glimpse into how it reasons.
The biggest downside thus far is that it thinks it knows better than me.
I have been adding system prompts with increasingly forceful verbiage requiring it to trigger web search and use the information found to augment its answers, but it just keeps ignoring these instructions and then giving me either false or outdated answers based on internal knowledge instead.
Then when called out on it, it acknowledges it, apologizes, and then just continues doing the same thing. It feels like arguing with a teenager
I may have to just set things up so the web search and scraper force feed information into the prompt, which partially defeats the agentic approach Gemma4 uses, unless anyone has some suggestions how I can overcome this and make Gemma4 more… obedient?
But who knows. I’m still running it at 4bit quant to make it usable on the CPU. Maybe once I get the GPU installed and up it to Q6 or Q8 it will do better.
Anything LLM can use multiple different LLMs within the same chat to complete requests, depending on whatever rules you make above, or by manually changing the model mid-chat.
Turns out this was actually the case. I tested gemma4:26b-a4b-it-q8_0 today, and moving up to 8 bit made the model much better at following instructions.
And surprisingly it is still pretty quick. Prompt evaluation/pre-fill is actually faster than with the 4_K_M version, presumably due to the quantization processing being easier, and it only lost ~18% token/s performance, probably due to some peculiarity having to do with running on a CPU.
This is reasonably usable with a little bit pf patience, but when the GPU gets here it will be very fast.
I have little idea what I’m doing, but I’ve been tinkering for the past couple of months and scratching the surface with some LLMs. My linux box is running a Blackwell GPU with 96GB of VRAM and I’ve got 125GB of CPU RAM. Figured I’d wince and pick it up because the trend of us commoners not being allowed to purchase things seemed not great, and so far it looks like I was right although I’m still paying off that credit card from back when it was “affordable-ish.”
I’m just running Ollama models and mostly through OpenWebUI although I think I will be transitioning away from OpenWebUI (installed it without docker and it always needs to phone home over the Internet before it will boot up the server interface). Anyway, for coding my favorite model has been gpt-oss:120b. It does a great job when I post multiple file code into it with hundreds of lines. And it puts it out quicker than I can read it. Sometimes it will get stuck and I’m not sure how much of that is from a long chat and deterioration, but most of the time it figures out issues. Recently it did have an issue with a PHP reading a datetime from a JSON file and the PHP would apply some time zone to it for some reason, and it couldn’t figure out how to fix that even though it was far less lines than I usually pump into it. So I put it to Qwen3.8:27b and it fixed it no problem. But normally it’s my go to, but I’m going to do some more work with Qwen3.8 and also the Qwen3.5:122b and they may take the top spot.
For vision stuff I have my IP cameras just using YOLO for object detection, but during sleep hours and at certain locations on the property I will have it then send the image to Llama4:108b to get more info before I have the script take action (like alert us and yell at bears or moose with a networked loud speaker, or for funsies see if a person is holding an invisible bat before it plays the “What I Believe” rant from Bull Durham). It works pretty well, although when I ask it ten questions in a single prompt and ask it to respond with A, B, C etc for processing whether to play “YMCA” or “Gangham Style” or the Bull Durham rant based on what a person is doing, it doesn’t do as well. But I’m just dipping my toe into the vision stuff.
I think I’m about to try Anything LLM and start doing some RAG stuff next.