Shapelearn Qwen 3.8 27B (13.1 GB VRAM)
62 points by syntaxing 7 hours ago | 7 comments

_ache_ 4 hours ago
From my own test. It's not faster than the unsloth model.

Disclarer: I'm unsing Vulkan on an AMD GC.

reply
Systemerror7A69 26 minutes ago
AMD 7900 XTX with Vulkan here as well, wasn't faster on my test either. Might be much different on Nvidia though.

I assume the limit for me is memory bandwith, as the 7900 XTX has the same bandwith as the 3090 from what I can gather and I already reached ~60 t/s with Unsloth. Those would fit with the numbers Byteshape has for their cards.

4090 and 5090 have much higher bandwith apparently, so on those cards you can probably get much more out of the kinds of performance improvements they are doing.

reply
DiabloD3 50 minutes ago
Surprised its not meaningfully slower.

Vulkan and ROCm paths are missing a few optimized versions of the quants they're using.

reply
sheo 4 hours ago
reply
Mashimo 3 hours ago
In the comments it reads like bonsei falls apart on longer running tasks.
reply
txrx0000 3 hours ago
Not really. The largest IQ4_XS quant here is still worth it because Bonsai doesn't offer larger quants. They could beat it if they made a quaternary variant though, I don't know why they're stopping at ternary.
reply
rguiscard 2 hours ago
I wonder the same thing for Bonsai 2. ByteShape offers 5 models from IQ2_XXS-2.56bpw (8.8GB), IQ3_XXS-2.88bpw (9.9GB), IQ3_XS-3.01bpw (10.4GB), IQ3_S-3.23bpw (11.0GB) to IQ4_XS-3.84bpw (13.1GB). Their benchmarks show gradual improvement with size and users can pick one to fit theirs need. Bonsai-2-27B now is about 8.6GB. It might be good to have a quaternary version around 10-11GB to fit a computer with 16-24GB RAM.
reply