How much context tdo you consider to be decent?
30b class models usually fit in my VRAM even with quite large context. They just don’t evaluate as fast as I’d like. 20-25 tokens/s is a tough sell once you’ve gotten used to the 80 from an MoE model.
How much context tdo you consider to be decent?
30b class models usually fit in my VRAM even with quite large context. They just don’t evaluate as fast as I’d like. 20-25 tokens/s is a tough sell once you’ve gotten used to the 80 from an MoE model.
It was never about cost to me.
I was not about to leak all of my private data to any cloud model.
I utterly refused to use AI at all - even free ones - until I could run it locally without any external connections outside of my LAN.
I even firewall everything off on a private subnet, and disable all huggingface/ollama/etc external connections to make sure nothing escapes.
This is a hill I will die on.
If you’re writing code with a cloud hosted llm you aren’t leaking personal data? Especially if youre open sourcing it?
And yeah it’s not just about cost for me either. But there’s just no comparison between the output I get out of a frontier model and something that will run at home.
That depends. I personally have been using Olama Cloud that is ZDR and no token inspection or monitoring. However, you can’t 100% know if you send anything out to the Internet. Therefore, the critical stuff I run locally and I use the cloud for things that I still don’t want necessarily to be shared, but they are not intellectual property related, for example, or contain personal information, passwords, or financial records.