How far are local models from Opus 5.5? Even on expensive hardware?

I’m just trying to figure out what kind of current hardware would be required to match or somehow exceed Opus 5.5? As it seems to be the new peak of solid general AI coding performance?

I know matching the software capabilities of GPT / Claude for long running goals is another matter, or has that been figured out yet? I haven’t tried running linux/open code on a while on my RX 9070 XT system, but previous it was quite a pain and barely worked, didn’t try a lemonade server though.

Like Might Gorgon Halo with 196GB of memory be alright? Or would someone need that new Threadripper Halo system with 4 MI350ps? Or the NVidia equivalent.

Kimi K3 is probably the best you can do with an open model.

If you quantize it as hard as possible, it still needs 610 GB VRAM + RAM. The full-precision model would need 1600 GB.

I don’t know how much you can give on these benchmarks, but if you trust them and got close to one TB of VRAM GLM 5.3 is decent:

I was hoping someone had some real world uses, but holy crap 1TB of RAM only gives you decent???

I’m running a very large local setup with latest ds4, glm5, etc.

Different models are good for different use cases. Your specific one matters. Test it extensively on the hosted version of the model before considering a local instance.

From a cost perspective for most people it makes no sense. For hobbyists, its a really hard swing unless you are dropping 100k.

1 Like

He said VRAM, aka GPU’s dedicated RAM. That’s A LOT of GPUs or AI-specific hardware like those modules with HBM.

1TB of registered RAM is probably cheaper.