Breaking the 1.58-bit Barrier for Ternary LLMs

(arxiv.org)

81 points | by matt_d 1 hour ago

5 comments

  • om8 14 minutes ago
    Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    • om8 12 minutes ago
      If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
  • plqbfbv 15 minutes ago
    Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
  • NooneAtAll3 16 minutes ago
    This is the only time "1.58 bit" phrase makes more sense than "1 trit"

    Who knew that if you actually look at information entropy you can pack stuff better!

  • Kevcmk 38 minutes ago
    Woah. Good science.
  • kadushka 32 minutes ago
    [flagged]