Open source models for coding?

I am looking for good open source LLM options that can be used for coding in something like opencode. I tried Gemma 4 and Qwen 3.6 and find them to be underwhelming and perform poorly (in terms of output quality, not in terms of execution speed).

I am trying to evaluate options that can be run on a 128GB Ryzen AI Max system and intend to purchase hardware eventually, but until then I will be trying out using cloud providers.

What models should I look at?

1 Like

Reality check: you’re not going to be one-shotting apps on local LLMs like you would in the cloud. That’s not the use case they’re designed for; doing actual work through an agent is a very different experience with them.

Run by an agent, like Cline or OpenCode (for example), Qwen 3.6 35B or 27B can be excellent - if configured correctly. That means setting up the full context, full GPU offload etc.

I use 35B on a 1.2 million line Ruby codebase with Cline, and not only does it work well, it often spots stuff that Opus 4.7 misses (that’s not to say that it’s better than Opus 4.7 in general, mind). With my R9700, it’s also noticeably faster and recovers better from tool call failures.

So…the question is: what are you planning to use it for, and how are you planning to use it?

3 Likes

I have these small models often failing after multiple iterations for simple tasks as “rewrite this 30 line batch script into bash” while claude sonnet can one shot it without problems.

Which models? And which quants?

What models are you using that are failing? Qwen3.6.-27B definitely shouldn’t be failing, if you are using some low tier quant; definitely check out Unsloth because even UD_Q4_K_XL shouldn’t fail, but you need more than 16GB VRAM… 24GB minimum.
I Have heard that Qwen3.5-9b is decent, and some people have made 3.6-9B now but you need again to do UD_Q8_K_XL if you really want the most from it.

1 Like

One of my systems is a 128gb strix halo, and there are some great coding models that can run on it. First for the models you were having programs with.

  1. Make sure you have a good quant, don’t go lower then 4 but 8 is better.
  2. Inference and loading settings. On a stix halo always turn mmap off. For inference settings for a model like qwen 3.5 35b or the 122b that a strix halo can run in q6 with full context length I set temp to .65, top k 20, repeat 1.03, top p .95, bottom p .03. You can play with the number of active experts aswell I run the 122b with 20 active.

For running models on strix halo you should definitely look at these strix halo toolboxs Strix Halo AI Toolboxes

For running models on any system everyone should check out unsloth studio. You can run and train models in it. There is auto model loading and inference settings if you want, tool calls correction, speed increase and more. Introducing Unsloth Studio | Unsloth Documentation

Models I’m running qwen 3.5 122b or three of the 3.6, deepseek v4 flash, mistral 3.5 medium (it is slow but great I may make a MOE version to fix the speed), step 3.7. Let me know if you need more.

2 Likes

Qwen just announced that 3.8 will be released soon, and there will be 3.8 open weights models.

3.6 didn’t have a large MOE model, and 3.7 didn’t have open weights, so there are many people hoping 3.8 will have a ~100B MOE model.

1 Like

I’m actually doing a lot of testing for them and putting in tons of bug reports continuously. And together we are improving the platform rather quickly. I’ve been basically using it every single day. So almost every bug I am finding it and writing detailed reports on it, or at least as detailed as it needs to be. I definitely recommend that anybody going into local AI, check out Unsloth Studio because I think that it is the best mixture of everything that people are likely to use. From a local LLM outside of a direct agent for programming or some other specific task.

The software also does great for training, which I am not currently into, but have only done testing for, and does this well on AMD, even in VRAM limited scenarios.

Well, the 3.5 had the 122 A10B version, but honestly, that one is actually worse than the 27B model when it comes to deep thinking.

The most dense of the entire QWEN family was actually the 27B. Because even the 397B only has 22B active.

I would welcome a 50-80B dense like what we got with Qwen3.

1 Like

Point of order - there was no 50-80B dense model in the Qwen3 family, as far as I know. There was Qwen3-Coder-Next 80B, but that was MoE with 3B active.

Qwen3 Coder 80B was MoE? My mistake then, I thought it was Dense.
With only 3B active, no wonder it wasn’t very good when I tried it… I had the best luck with Qwen3-14B-BF16 originally.

The Qwen3.6 27B dense model while slower than the 35B MoE, is one of the top models that you can run on a machine with 128 GB (and 64 GB with Q8).

I use both of them depending on the task, if it is coding or architecture discussions. My coding is usually long horizon agentic coding tasks, using opencode.

Try BF16 as test, it “should” at least solve your tasks.

1 Like

I like to use the Unsloth Dynamic Quantization versions in Q8 XL tensor size. So essentially it has BF16 tensors for the important ones and quantizes ones that are less important.

This.

a $20 codex sub will get you access to excellent models that will 1 shot pretty complex stuff.

If this is for hobby or privacy purposes go nuts but if you want to actually grind out code, you’re probably better off paying for a subscription and writing it off on tax, or just accept that its cheaper than buying hardware for it.

Just be aware the difference between something like Qwen in “normal person” size and even luna or terra is huge and far less rooting around fighting the model to make it do things properly

Just for completeness, things have changed a bit since Qwen 3.8 27B came out. It’s totally possible to one-shot pretty complex stuff with it now.

Even without that, I’ve more than made back the money I put into my pair of R9700s since I bought them relative to subscriptions - I work on complex mostly human-written systems, and because I actively don’t want the AI doing my entire job for me and therefore use it as an assist rather than a crutch, local AI is far more suited to my needs.

Apart from anything, I don’t have to worry about optimising my workflow for token costs…I can just dump that time into thinking about things that matter.

1 Like

I have had very good luck with coding on Poolside’s Laguna S 2.1. It’s a bit big though (118B) so you’ll need to have a lot of RAM.

It’s an MoE model though, so you don’t necessarily need super high memory bandwidth, as it only runs ~8B parameters at once.

I have this one running as a secondary model on my CPU (EPYC Milan with 8x channels of DDR4-3200, so if you are patient with pre-fill, eval stage is actually surprisingly high for a CPU) and it has proven to be quite solid with coding.

1 Like

This is not my experience.

3.8 is good but you need hefty hardware to get decent context window and its results are nowhere near sol (or even sonnet 3.7 from months ago).

Now I’m not saying local is pointless. But if you actually care about the end result and don’t already have hardware it’s likely that for most single dev projects you’ll get better results cheaper with a frontier sub.

Local is imho more suited to other things: triage mostly.

As I said, it kind of depends on your use cases, and to a certain degree whether you’re holding it correctly.

Cards (heh) on the table - I bought my R9700s in Jan this year, cost £2450. Before that, I had a pair of 3090s for roughly a year - those cost £800 for the pair…which I sold for £1350 to get the R9700s. So…total outlay = £1900 (I totally get that this can’t be done now), bearing in mind that I already had the machine I put them in.

In that time, a single Claude Max subscription (which doesn’t actually provide the amount of usage I’ve had from the local instances) would’ve cost me £1800.

Compared with the guys who were using Claude at work, where I refused and used my own in a more limited and focused fashion (essentially forced, because of the more limited models at the time), there was much less rework required (and that rework was costly for the business - including losing a strategic partner, thanks to the edict from on high that we should “trust Claude” because “velocity is everything”) and fewer bugs recorded. In a couple of cases, the fix was actually to use my local instances for long-running analysis tasks that couldn’t be achieved with Opus due to usage limits, in order to fix the financial data that Claude broke. For context, that’s very much a non-trivial system with 1.5M+ lines of code and running around £15 million in transactions monthly, but I can honestly say that I’m yet to find a situation where a Qwen MoE couldn’t cope with it given careful instruction. I never got to try 3.8 27B on that codebase, because I quit before it was released.

At the same time, my GPUs have been running overnight batch analysis to prevent my forum being shut down as a result of Online Safety Act bullshit, as well as fulfilling the documentation and reporting requirements of that same bullshit.

To do all of that, I’d have needed at least two Claude subscriptions as well as API costs on top.

The difference between local and frontier, for me, essentially came down to two things: it forced me to be more focused and specific when assigning it a task than I could’ve been with Claude, and there was no temptation to trust the output (which is where most people come unstuck with frontier models), while also freeing me from token allowance anxiety. Is there a lot more human attention required? Absolutely, I’m not denying that, and this is the “Are you holding it wrong?” test I mentioned up at the top. But I would also say that’s not necessarily a bad thing.

All this comes with the caveat that I haven’t tried 3.8 27B in Q4-level quants, which is where most folk with high-end gaming GPUs are going to land (a reasonable new spec 27B would be a pair of 16GB 5060 Ti GPUs) - thanks to buying relatively early, I’m blessed in that I can run it in INT8 on one rig and FP8 on the other. What I can say is that…if I’d had that model 18 months ago, it would’ve genuinely felt like cheating.

1 Like

OK… update.

Haven’t used this for code yet, but glm5.3-flash just dropped recently.

Holy shit.

Its supposedly 1-10% the cost of models like opus 4.8 and ballpark same intelligence with very good instruction following. Multi-modal, 1m context.

I got it to give me a decent network documentation plan (PDFs attached). This is comparable to what i got out of 5.6-sol.

Cost?

0.9 CENTS via openrouter

This is a game changer.

NetBox_Evaluation_and_runZero_Integration_Plan.pdf (68.3 KB)
Network_Documentation_Platform_Proposal.pdf (63.2 KB)

Previously known as ox alpha

It does make buying hardware to run local, or not running cloud due to cost a hard sell.

Yes, its chinese. Yes its running on chinese hardware. So maybe consider what you put into it.

But holy shit.

1 Like

I have been using GLM 5.3 Flash quite a lot through my O lama subscription and I have to say it’s quite good and it can do a lot of things while not costing so much. I think that more specialized models that are capable of doing the task that most people are doing are where the future of AI will go. And this is also why open source AI is so important and why local AI is important because you can do things without worrying about token limits or usage limits and run huge tasks that can take an entire day without having to care about it.

1 Like