Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- When I first built features with GLM, there were lots of bugs, and it took me ages to fix them manually. Now the features implemented with GLM 2.0 have almost no critical bugs after testing. I can’t even imagine how capable K3 will be. It may well be on a level that ordinary people cannot access.by L-patpat
- > k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k.by wxw
- This isn't quantized, right? Just a smaller context?by dgritsko
- I was so excited that it is open source until i realised the model required 1.5To of VRAM. Unsloth has compressed it in 1bit at about 570gb VRAM with 75% accuracy, that's almost mac studio territory...by fumber
- This is just an API level change right? The model itself should be the same I think.by HawtAds
- Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
- I make a point of never going beyond about 220k, unless absolutely necessary (and it's almost never necessary), anyway, even with models that degrade more slowly, so this is just a discount.by SwellJoe
- This seems functionally similar to OpenAI having a step in pricing once you exceed a certain context length (also at 272k aka 2^18 aka 256k).
Having a lot of active context increases the per-token cost (flops issued and bytes read per token out) so it makes sense to pass that cost on to users. I'm actually surprised it's implemented as a hard cutoff instead of a smooth gradient.
by wren6991