As I said, it kind of depends on your use cases, and to a certain degree whether you’re holding it correctly.
Cards (heh) on the table - I bought my R9700s in Jan this year, cost £2450. Before that, I had a pair of 3090s for roughly a year - those cost £800 for the pair…which I sold for £1350 to get the R9700s. So…total outlay = £1900 (I totally get that this can’t be done now), bearing in mind that I already had the machine I put them in.
In that time, a single Claude Max subscription (which doesn’t actually provide the amount of usage I’ve had from the local instances) would’ve cost me £1800.
Compared with the guys who were using Claude at work, where I refused and used my own in a more limited and focused fashion (essentially forced, because of the more limited models at the time), there was much less rework required (and that rework was costly for the business - including losing a strategic partner, thanks to the edict from on high that we should “trust Claude” because “velocity is everything”) and fewer bugs recorded. In a couple of cases, the fix was actually to use my local instances for long-running analysis tasks that couldn’t be achieved with Opus due to usage limits, in order to fix the financial data that Claude broke. For context, that’s very much a non-trivial system with 1.5M+ lines of code and running around £15 million in transactions monthly, but I can honestly say that I’m yet to find a situation where a Qwen MoE couldn’t cope with it given careful instruction. I never got to try 3.8 27B on that codebase, because I quit before it was released.
At the same time, my GPUs have been running overnight batch analysis to prevent my forum being shut down as a result of Online Safety Act bullshit, as well as fulfilling the documentation and reporting requirements of that same bullshit.
To do all of that, I’d have needed at least two Claude subscriptions as well as API costs on top.
The difference between local and frontier, for me, essentially came down to two things: it forced me to be more focused and specific when assigning it a task than I could’ve been with Claude, and there was no temptation to trust the output (which is where most people come unstuck with frontier models), while also freeing me from token allowance anxiety. Is there a lot more human attention required? Absolutely, I’m not denying that, and this is the “Are you holding it wrong?” test I mentioned up at the top. But I would also say that’s not necessarily a bad thing.
All this comes with the caveat that I haven’t tried 3.8 27B in Q4-level quants, which is where most folk with high-end gaming GPUs are going to land (a reasonable new spec 27B would be a pair of 16GB 5060 Ti GPUs) - thanks to buying relatively early, I’m blessed in that I can run it in INT8 on one rig and FP8 on the other. What I can say is that…if I’d had that model 18 months ago, it would’ve genuinely felt like cheating.