Breaking the 1.58-bit Barrier for Ternary LLMs
102 points by matt_d 3 hours ago | 9 comments

infogulch 23 minutes ago
So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

reply
kadushka 16 minutes ago
By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
reply
om8 60 minutes ago
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
reply
janalsncm 28 minutes ago
PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

reply
mitxela 5 minutes ago
which is important though since sending it across the wire over and over and over is actually the main bottleneck.
reply
om8 58 minutes ago
If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
reply
plqbfbv 2 hours ago
Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
reply
NooneAtAll3 2 hours ago
This is the only time "1.58 bit" phrase makes more sense than "1 trit"

Who knew that if you actually look at information entropy you can pack stuff better!

reply
Kevcmk 2 hours ago
Woah. Good science.
reply
kadushka 41 minutes ago
[dead]
reply
kadushka 2 hours ago
[flagged]
reply