

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- The title of the paper is correct. The paper does not seem to actually get to the point in a generic way, but hyperfocuses on, effectively, one type of error compensation.
Highly quantized models, especially with highly quantized KV caches, will, effectively, attend to the wrong tokens and be unable to easily discern highly similar tokens. The bastardized way of explaining this is gradient descent techniques get stuck in localized minimum and global maximums, so what happens when you turn the slopes into hard stair steps?
We need to move to smaller models and smaller caches and better samplers, not new quant methods (although I'm willing to also take those too).
by DiabloD3