Bonsai 2 27B (ternary) on a 12 GB Arc B580: 128K context, ~80-90 t/s — looking for testers on other Arc cards

I’ve spent the last few weeks getting PrismML’s ternary Bonsai 2 27B running properly on my Arc B580, and it’s turned into an open-source llama.cpp branch that I’d love more people to try.

On the B580 at 128K context, all in the card’s 12 GB: about 80-90 t/s writing new code, 250-370 t/s when editing code you’ve pasted in (speculative decoding does a lot of the work there), and still around 44 t/s with 115K tokens of history. It draws about 120 W.

The interesting part is how much the XMX matrix units matter. PrismML’s own fork added basic Intel support last week, and on the same card, same model and same settings it does about 38 t/s on new code against 88 here, and 8 t/s against 36 once there’s 32K of context. Most of the gain comes from:

  • repacking the ternary weights to 2-bit at load so Intel’s TernSYCL int8 x int2 kernels can use them directly,
  • reading the 4-bit KV cache straight into the XMX units for attention (using an int8 x int4 instruction Xe2 has but that isn’t in Intel’s public extension list),
  • trimming memory in the speculative-decoding draft context so 128K fits in 12 GB.

It also runs Google’s Gemma 4 12B about 2.5x faster than stock, and works as a local model for Claude Code.

Linux build steps, a Windows zip and the full guide are here:

What I’d really like is results from other cards. The fast paths are for Xe2 only (B570/B580, Lunar Lake). One A770 owner found it runs on the fallback kernels at about 9 t/s but needed -b 256 -ub 256. If you try it, please post your card, OS and t/s either here or in the GitHub discussion: