Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Thank you for sharing this. I like to test out running LLM's on edge computing with limited RAM and GPU/CPU so this research will have practical implications on my activities. I also appreciated how the authors formulated 1.58 (it's log_2(3)) because that was embarrassingly confusing for me when I was first introduced to ternary LLM's.
  • So this compression is only pertinent to the LLM file format? In memory it'd have to be expanded into the 1.58-bit form - 5 trits per byte.
  • Only a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
    by wgd
  • sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.

    And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...

    0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

  • This is the only time "1.58 bit" phrase makes more sense than "1 trit"

    Who knew that if you actually look at information entropy you can pack stuff better!

  • Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
    by om8
  • So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

    If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

  • > We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout

    I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?

    by c7b

Explore Birbla archives