If I only have a 5800X3D + RX 9070 XT at the moment.
If i were to dedicate my PC to a linux local LLM OpenCode, Ollama, openclaw box, how far could i actually get having it write projects mostly on it’s own?
I don’t know anything about coding, but ive been getting decently far on some stuff using nothing but GPT, i decompiled a Gameboy Color game(allegedly) after about a month of work. Working on a Cura mod and rust game engine otherwise.
What if i were crazy and went for the R7000 ? At $100 a month(only doing it once) for premium codex, it’d actually pay for itself easily.
But how good are the local models that just came out? 50% of GPT 5.5? Or higher? Or it probably depends?
I’m mostly focusing on my game engine. Still have a long way to go, and i plan to try using some reinforcement learning to test my eventual games automatically.
My take is that local models are viable for supervised work, but without a large amount of vRAM, context and model size limits will make it unlikely to be successful with fully autonomous work. Not to say it’s not possible, but it will take work to engineer how to make it work, probably configuring various small models for highly specialized and narrow tasks. Context constraints will limit the number of tools you can give each model or the size of the content you can have the model read for context.
Recently on my work laptop, I’ve been trying Qwen3.6-35b with Codex. It works pretty well but it’s fairly slow. I did try using Qwen3.5-27b (pretty certain that was the size) with Claude Code to explore and document a large codebase, and it did reasonably well; it took a several days to complete tho. For both in using ~64k context window, but a larger context would likely help a lot.
Personally I use local models to workshop ideas interactively, like a sounding board. I’ve also connected my local models to SearXNG to help automate web searches. I’m planning to use them to help with document organization as well, but haven’t made moves there yet.
Overall it’s definitely possible with a small model and small context, but you’ll need to temper expectations and be okay with slow progress. If you want fully autonomous agents, quick responses, and large context, then you’ll either have to buy the hardware to support it or pay for a subscription.
I have a couple of them, and…well, I’ve spent a lot of time with Qwen3-Coder-Next-80B (at Q4_K_M) and Qwen3.6-35B-A3B (at Q8_0), and I have to say that Qwen 3.6 takes the crown, but both require over 32GB VRAM to run well (~95t/s for the former, ~110t/s for the latter). It’s a great balance of speed and accuracy.
Qwen 3.6 at Q8_0 feels (subjectively) about Opus 4.5 level - it doesn’t make dumb mistakes like Qwen3-Coder-Next, and with a good agent it’s pretty good at figuring its way around a codebase; it’s nearly there in terms of “trust it for autonomous stuff”. I’d imagine that 3.6 at Q4_K_XL would work almost as well, and with the R9700 you could easily get the full 256k context if you don’t mind running the KV cache at q8_0.
However, as @Burhan says, I’d probably avoid just setting it loose.
Well, to my mind Qwen 3.6 is a step-change in capability even compared with Qwen 3.5. All I’ll say is…when you’re looking at benchmarks, treat them as useful information but not concrete indicators of capability. Gemma 4 and the Nemotron models both do well on benchmarks, but Gemma 4 has a massive problem with its sliding attention window which causes it to forget context, and Nemotron seems to be benchmaxxed (ie effectively trained on the benchmarks themselves).
Also, don’t pay any attention to the one-shot demos you’ll find all over YouTube; they’re not really indicative of the real-world usage of these models - they don’t do well at building everything in one go (particularly given the 256k context limit), but they excel when building iteratively.
Things are certainly moving at an insane pace. I was struggling to get GPT 5.3 codex to do what i wanted, needed a lot more steps and help.
Where as 5.5 handles pretty much everything in one shot, and almost never has major compile errors or warnings even. Only rarely do i need to go back and clean up.
That’s about the level i’d need a local model to work at, at my current skill level, might be a little while longer. Maybe with the R9700 sucessor and a few more local advancements.
I think that it’s a matter of priorities and ROI, looking at the responses in the thread.
If you’re not doing critical work I think having a more limited model might even be good because it forces you to overcome the model and your own knowledge gap to reach the goal. With the added bonus that your hardware is yours forever, it’s private and all that.
On the other hand if you’re getting more work done, the one that pays your bills, having a faster and smarter model that’s constantly updated can be a smart investment or at least worth trying and do some math on before deciding to stick with it.
Last but not least (going on a bit of tangent here, but it’s related sorta) stocking on hardware capable of running AI of any kind locally feels to me like a good investment regardless. You have the means to be autonomus and that’s always a good move in my opinion.
This is my thinking. I strongly suspect that, pretty soon, the good cloud models are going to be stuck behind an economic wall - which is to say that everyone using (for example) Claude to any great extent is still costing Anthropic more than they pay them, and that can’t go on forever; GitHub have gone to per-token billing, and most of the others will probably follow suit as soon as “investors’ pockets” stops being a line item on their revenue sheet.
Ergo, getting used to the idea of running your own models locally, and incorporating “trying new models” into your workflow, seems like the most stable long-term strategy.
So, I don’t have an immodest setup (RTX Pro 6000 x2, on a threadripper machine, all the storage, etc.). I started this build about a year ago, and though I could run slightly bigger models than my last setup (3090 + RTX 4000 ADA x2,) I couldn’t find a great workflow or model that was even okay at tool use for coding.
The situation is much different today, I stopped using gpt-oss:120b which seemed like the best at the time, and switched to Gemma4:31b (bf16.) I need both gpus for 256k context, but I think a 5090 could do quite well with it. Gemma gets gummed up with a small context, but if you use another LLM with a coding harness to spawn subagents that take on small tasks, I think that a smaller model might even suffice.
Smaller models are getting better! I was getting depressed about the utility of local models even a year ago, but decent tool use + a good coding harness and you can accomplish a lot.
I’ve been using Claude code to orchestrate local models via opencode, which is probably not necessary, but I’m going to try some custom agents for opencode soon that I think can do the orchestration job handily with smaller models. I can offload about 50% of my tokens to Gemma at least right now, curious to experiment with smaller models for mechanical tasks. I really want to use only opencode with local agents to accomplish a similar output that I can get with Claude orchestrating local models…
Bit of a rant, but even if it’s slow, I think that you can improve the process more via the harness than you could before, and not rely so much on the model.
Good luck and I’m curious to see what you can come up with!
@aatchison - I’m curious as to why you prefer Gemma 4 to Qwen 3.6? In all situations I’ve tried, the equivalent Qwen 3.6 models always get the job done more accurately.
Obviously I don’t have a baller rig like yours, so the difference could be accounted for by the fact that I’m using quants (maybe the Qwen 3.6 models aren’t so sensitive to quantisation?), but I’ve found that the Gemma 4 models are just bad at tool use. At least, they are when using Cline as the agent.
I’m interested to try that model, but I’ve just had such success with Gemma 4 that I haven’t tried others recently. I’ll take that as a recommendation!
Quants and MoE of Gemma 4 have definitely performed worse than the full fat bf16.
I’m working on some projects that would make testing multiple models possible, even on the same task. This is good research, but I’m finding that harness and skill work is where my gains are right now. It’s kinda like switching Linux distros, there’s always another model, what what is the process you are feeding it?
I wrote an integration for Gemma4 when it came out. The small models had a very tough time constructing valid JSON reliably. Never tried it with coding directly, but adapting it was a hassle bc of the tool use issues for sure
If RAM prices weren’t insane right now I’d be so tempted to get the Minisforum 9955HX PC and pair it with the R9700
Too bad the framework 16 is the only strix halo product with a X8 PCI-e connection. It’s also wildly expensive
@wendell any experience with anyone trying to run a system like that? Minisforum with an R9700?
I hope framework is willing to send you a maxxed out Framework 16 for testing, would love to see strix halo + R9700 run local models and see them stressed to the limit.
@Streetguru - what would be the point? If you’re using the R9700, it’ll run just as fast on a Ryzen 5600 as it would the latest-gen CPUs. If you’re using a Strix-Halo GPU for LLMs, the R9700 will be idle. And if you tried to use both, it would be slower than either.
It’s fine, there’s only about 220W going through the 12VHP connector. We’re not in 4090/5090 territory here.
On another note…anybody who’s looking at Qwen 3.6 should take a look at this PR, which implements proper MTP support:
I’ve just tried it on my R9700, now that he’s added the Vulkan path, and it’s basically doubled my TG from 32t/s to 55-65t/s using the 27B dense model. Even the 35B MoE model has seen a ~50% increase, from 105t/s to 150-160t/s.
The only problem at the moment is that it crushes PP performance by ~60%, but he’s aware of it. Hopefully he can fix that and get it merged soon, because that kind of performance boost is phenomenal.
This is something I’m interested in, as working with code that has little to no (or bad!) documentation and strange syntax has been driving me crazy.
From some googling, Qwen 3.6 35b should run on a 24GB crad? Though I have an RTX Titan (2000 series) which doesn’t support FP16, so would that actually fit since the older cards straight up just use fp32 when processing this stuff?
I have tested a LOT of 12vhpwr cables. I believe that the latest Superflower PSU- the Leadex VIII, when used with a Wireview Pro2, is basically a solved issue. I haven’t seen any voltage spikes, temp raises, or current imbalances in 100’s of hours of use. I saw some issues in other ‘premium’ power supplies like Seasonic, Corsair, and Thermaltake, as well as in the earlier Leadex models, but this latest one- the 1200w specifically has some sort of cable design and construction that is really solid.
Ironically, instead of googling you could paste these questions into an LLM:) but since you already asked I’ll try to answer: yes it can run Qwen 3.6 35BA3 as it will definitely fit 24GB VRAM at Q4_K_M quants (also I’d recommenced Q8 for KV cache) with decent context size. Especially since its MoE model you can offload some (I’d guess 20-30% ) of model’s layers to system RAM without much performance penalty(of course depending on RAM and PCI-E bandwidth). This would allow you to use a good Q5_K_XL quant or more context size. You could offload even more(e.g. 50%), but expect way more significant performance drop.
And one correction that GPT 5.4 found : TITAN RTX does support FP16. TITAN RTX are listed as supporting FP16 and INT8, and NVIDIA’s datasheet lists 576 Tensor Cores.
So computing performance shouldn’t be that bad vs something like 3090 or 4090.