The bundled LM studio llama.cpp for whatever reason is way slower than fresh off GitHub through the CLI. I have no explanation. I get DigitalScream’s numbers with two 9700 pros after copying the general setup, Linux with a fresh Vulkan llama.cpp. Nothing else gives these numbers for me.
Try Unsloth Studio; they are doing additional optimizations every day atleast in terms of commands and such.
Reading the llama.cpp git log, have the distinct impression there are frequent performance improvements. Been running a compiled version pulled direct from llama.cpp github.
Might be because I have old hardware, but I feel its getting slower over time despite all the optimisations. Didnt do any benchmarking though…
Unsloth Studio updates it outside of the updates to the system itself and sometimes there are multiple llama.cpp updates per day, which pop up in your chat and you can just press to perform these updates.
Just pulled llama.cpp, and there are 10 commits since I pulled (and built) yesterday! I have built and use llama.cpp for Vulkan.
Installed Unsloth, but it seems to not like my GPU.
Am somewhat interested in comparing Vulkan performance to ROCm, so got 7.14 installed. Guess I wiil also need to compile llama.cpp for ROCm.
Which GPU do you have that it’s not playing nice with it because there’s been a lot of work in the last month to make it play nice out of the box on almost every GPU from Nvidia and AMD on the common operating systems.
Watched Wendell’s post on using a model-router, so tried to define a project to deploy a model-router on my local subnet. This went sideways. Seems LLMs are cranky. ![]()
Dropped back a level and came up with a project-skeleton for working with LLMs:
Got a clean project skeleton, using DeepSeek Harness and DeepSeek V4, created a “homelab” project with all my planned projects, and now need to find a new topic.
Just slightly unsettling when you point a Chinese LLM at your router, and it says:
Key works — I'm in.
I wouldn’t worry about where an LLM is from, considering that they don’t automatically connect to any kinds of servers. They’re actually one of the safer pieces of software, I guess you are running, because when you open an LLM, it’s not its own software. You are actually mounting it with another software. So I really wouldn’t worry about it in that context.
Well, I tend to agree, but that did not go well. Ended up completely breaking my network. Admittedly, a slightly complex setup, with a router, two configurable switches, and at least three subnets. Still, managed to bork the set. Had to manually recover.
And now about let it try again. Definition of not-sanity… ![]()
How was running an llm on the v540/ BC-160? Were you able to use both gpus on the V540?
The BC-160 works okay. For 45 USD (what I bought it for), the ~24 Token/s is not bad. The v540 was giving me issues because it seems that it requires bifrication support in order to use both GPUs. I could get it to use one (at random) if I tried to run both, it would crash or lock uo my PC. The motherboard that I haf it in does not support PCIe bifrication. The v340 on the other hand has a PCIe switch built in. No need for bifrication
Possibly the last on this topic. Working through the learning curve in using an LLM for development. As an example exercise, used the prior fan-control script as a base, and wrote a small C++ program, but with exact knowledge of the Linux AMD GPU driver.
- Article: Fan control service for AMD MI25 GPU
- Github: GitHub - pbannister/amd-mi25-fan-service · GitHub
Frankly, ran out of ideas for improvement. This was meant as a learning exercise. ![]()
Yes, there definitely is a learning curve when you’re using an LLM for development. I’ve done quite a bit of it now, but not really with desktop applications. It seems that LLMs struggle in C++, although the newer ones are definitely getting better at it. The code that it generates I am usually not satisfied with.
A few years ago, I would have (and did) say exactly the same.
Only recently have I had chance to really focus in this area. In the past few months, I have learned quite a lot. Took a bit to find what worked and what did not. (Made a bunch of dumb mistakes. Learned.)
Took a while before I closed the loop. An LLM has in effect read all of human knowledge (and slightly understood) … ya’ know … maybe I should ask what the world knows … about using LLMs. Turns out, a recently-trained LLM is pretty up-to-date, and if you ask the right questions, comes up with some pretty good answers.
Iterated on that a few times. (Still iterating.) Compiled a project-skeleton with a starting set of rules that are in effect compiled from all of human knowledge. Folded in what I have learned in my experience. With each exercise completed, ask the LLM to do a retrospective, and fold back in lessons-learned.
This seems to work - rather more than I expected.
Now I am looking at LLM-generated code that looks like something that I might have written.
Bit spooky. Working project skeleton is here:
My recent dumb questions have started pulling up recent research papers. Might be catching up with the bleeding-edge. Not quite there.
To pull this back to the original topic here, should note that there is constant progress all through that stack for running LLMs. Some of that very recent progress looks to much improve what you can do with a modest local GPU.
Might need an update here, eventually.