Local one-shots, you say? Yep, we're there with Qwen 3.8 27B

Honestly, I can’t believe how good this little dense model is. 5000t/s PP, 100t/s TG on a hacky unlocked CMP 170HX, and yet it can one-shot games like this with OpenCode:

https://www.thefretboard.co.uk/games/station-defender/

Yes, I woke up this morning and wanted a tower defence game. My usual go-to 3.6 35B failed miserably, even after I gave it 3.8 27B’s plan. The 27B, on the other hand, hit it in one go.

It even took it upon itself (without prompting) to playtest the whole game from beginning to end, to make sure it was possible for a human to beat it.

Previously, I’ve only managed this kind of quality with Opus.

Perhaps this is how the bubble bursts?

3 Likes

I’d agree that it’s a milestone in achievable local LLM coding performance on less than extreme hardware. I’ve been running a 3-bit quant on a 5070Ti and it’s the first local model that persuaded me to properly wire up a local agent into my dev flow.

However, it’s not enough to get me to drop my Claude (or other frontier model) sub and I don’t see it bursting the bubble*. It does certainly shift the landscape in terms of how much I’d be willing to pay and what I expect in terms of quality and performance for my sub though.

My experience will be different because “one-shot” tests are entirely meaningless to me for what I do. I have large projects with legacy codebases. Qwen 3.8 needs for more direction to deal with those than the like of Opus or Fable, especially when running on my constrained hardware.

*I’m really just using that term to reflect the post. I’m not someone who believes the dev side of LLMs is a bubble at this point, although datacenter hyperscaling and OpenAI in particular may well be.

1 Like

I also use my local LLMs for large real-world projects (but a description of those would be far too long to be relatable in a post on a forum, hence the one-shot).

The issue here is most likely that you’re running the model smashed down to 3 bits, whereas I’m running it INT8 with 512k context, and that’s two orders of magnitude in weight range. That’s not to flex, rather to point out that there’s a genuinely huge difference between the two; I haven’t been able to find anything I need to do that has noticeably different results between Opus and 3.8 27B INT8.

sorry, what? on a single 10gb card?

1 Like

An 8GB card, actually, unlocked to 64GB. The 10GB cards (mostly) only unlock to 40GB.

EDIT: Correction above, you need to win the silicon lottery to unlock 10GB → 80GB.

1 Like

The 8gb card the little card with hynix memory.The 10gb card the full with Samsung memory and unlocked for 80gb.

Sat Aug 22 22:59:09 2026
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 610.43.02 KMD Version: 610.43.02 CUDA UMD Version: 13.3 |
±----------------------------------------±-----------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 3050 Off | 00000000:16:00.0 On | N/A |
| 34% 37C P8 10W / 70W | 403MiB / 6144MiB | 0% Default |
| | | N/A |
±----------------------------------------±-----------------------±---------------------+
| 1 NVIDIA Graphics Device Off | 00000001:16:00.0 Off | N/A |
| N/A 46C P0 35W / 250W | 16MiB / 81920MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+
| 2 NVIDIA Graphics Device Off | 00000001:40:00.0 Off | N/A |
| N/A 44C P0 35W / 250W | 16MiB / 81920MiB | 0% Default |
| | | Disabled |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 38936 G /usr/bin/gnome-shell 332MiB |
| 0 N/A N/A 52630 G /usr/bin/Xwayland 2MiB |
| 0 N/A N/A 52774 C+G /usr/bin/ptyxis 19MiB |
| 1 N/A N/A 38936 G /usr/bin/gnome-shell 5MiB |
| 2 N/A N/A 38936 G /usr/bin/gnome-shell 5MiB |
±----------------------------------------------------------------------------------------+

Nice - looks like you won the lottery, twice :slight_smile:

1 Like

Yes i buy these cards on marc because i no this card hawe lot of memory.

And nao i told you all the 96gb A100 card with hynix memory coming.

I dont no why only 4 memory stack opened your card.

Okay, but that’s also the gap between realistic local hardware for even enthusiasts and esoteric builds or spending car money for hardware that doesn’t stand up as a financial cost against the subscriptions to frontier models or cloud rates.

Except…no it’s not. As noted, I’m using a CMP 170HX, which cost me £900 (£300 less than the 5070 Ti that you’re using, in fact). The rest of the machine was about £500.

That’s less than the average gaming PC, even two or three years ago.

Are you trying to claim this isn’t an esoteric build?

I’m claiming that there are lots of enthusiasts doing exactly what I’ve done. Head over to r/localllama and have a look, if you like.

It’s not exactly an esoteric build, it’s just a build with the cheapest 64GB card that money can buy for folk willing to accept a bit of risk and limitation, and are capable of typing ./install.sh.

Even forgetting about that, you can run it on any 32GB card - plenty of those are available at retail for the same price as a 5070 TI or lower - with far higher quality than the 3-bit version you’re using. 16GB cards simply aren’t suitable for LLMs unless you have several of them…and this very forum shows that plenty of enthusiasts have 32GB cards (primarily B70s and R9700s).

This is silly. The average user is buying one GPU and it serves for all their needs. They aren’t hunting down some obscure mining card in the hope that it will be a key to hypothetical amazing LLM performance from models that may or may not exist when they buy it.

The card you’re talking about isn’t even available now for the prices you’re talking about. Good for you for doing the homework and taking the risk and all that but it’s really not an option for anyone.

“realistic local hardware for even enthusiasts”

Hence talking about 32GB cards, and the fact that your assertion that a neutered model with insufficient context for the job of working on complex codebases running on a 16GB card with frontier models is not relevant to the majority of people running local LLMs. Remember? That’s where we started - you trying to say that we’re not there yet because of your personal experience.

It’s pretty simple. I’m not suggesting that everybody buy a 170HX, I’m saying that in the consumer/workstation class (which includes dual 3060s all the way up to R9700s) there are a whole ton of people putting at least 32GB VRAM to use for this, for the cost of a single 5070 Ti or less - which is apparently your high water mark.

That includes a significant portion of this forum’s active users. Unless you think that the 6 million people who’ve downloaded Unsloth’s 3.8 27B model alone were all doing so with the 2-bit and 3-bit quants?

Nobody’s talking about the average user here. The average user doesn’t even have a discrete GPU in their system.

Unless, of course, you want to talk about Mac users who’ve bought a 64GB laptop in the last couple of years. They can run Q4-Q8 quants too, and last I checked there were quite a lot of people in that position.

I bought it about two weeks ago. There are plenty of them on eBay for £950-1050 and accepting offers, which is exactly what I did.

The problem here is your extrapolation. Just because it’s not an option for you, doesn’t mean it’s not an option for anyone. If somebody has a need for 64GB VRAM, even if the price of these goes to £1500, it’s still the cheapest way to do it. Perhaps not the best, but definitely the cheapest.

2 Likes

Do you really feel like you can properly judge a model by running it at a 3bit quant?

Logic tends to completely fall apart at less than 4bit quants. This is even more true on MoE models, where I don’t like to judge them at below 6bit quants. When quantized this low - in my experience they - just don’t resemble their full models enough to really be able to assess them.

When I was testing Gemma4:26B going from Q4_K_M to Q8_0 was absolutely transformative.

Unslothed have published the relative accuracy of their quants and plenty of other people have run comparisons so people can make their own mind up about what sort of marginal improvement in performance they can expect from upping their equipment to support a higher quant.

However, I wasn’t “judging the model” based on the 3-bit quant as if I’m saying I’m getting the full capabilities. I was responding to the idea that this is a “bubble burster” by stating my experience as someone with hardware that reflects somewhere near the upper bound on what most people would be running before you plunge into throwing money or time purely at running a local LLM. I still use Qwen 3.8 and it’s a nice thing to have to offload the simpler tasks from my cloud sub but it doesn’t replace it.

ok. dammit. I have been REALLY tempted by these 170HXen. did yours unlock to the full 64GB and run stably like that?

we have A30s at work which are also a flavor of “theoretically an A100 but disabled in various ways” and since I know how damn fast they are I am VERY tempted to try my luck with one of these cards. I’m just very skittish about the whole “maybe you really DO get a lemon” thing.

Yep, mine unlocked straight to 64GB and runs 100% stable. I haven’t heard of any 8GB → 64GB cards which didn’t; I think the binning strategy for these was different to the 10GB cards.

The problem I have is cooling; I just can’t find a fan optimised for high static pressure which is also quiet enough to reasonably run in the house. The result is that it hits 85C with about 3 minutes of constant use, and starts throttling down from ~1400MHz to (eventually) ~800-900MHz.

I’ve currently got an Arctic P9 Pro, which is a 92mm 3krpm fan with 6.3mmH20, attached to a duct designed specifically to push air through this card.

I’ve got some 120mm fans which do around 5.2mmH20, so I’m considering mounting that at the other end of the card (with a gasket, since it’s outside the case) to pull the air through. An alternative would be to get some kind of blower fan on the case, but they’re usually pretty loud (also, I don’t have any designs for a mount for a big one).

1 Like

I read your post last week and it didn’t really mean much to me then. After installing 3.8 27B last night I’m really interested to know what your prompt was to one shot a game like that?

I forget exactly what the prompts I used were, but basically I built up this plan file:

https://www.thefretboard.co.uk/games/prompts/TOWER_DEFENCE_PROMPT.md

I’ve run it a few times, and it always ends up with largely the same game. Give it a go :slight_smile:

3 Likes