Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Removed
  • 404 on the blog page? https://zartbot.github.io/blog/
  • I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.

    TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems

  • I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.

Explore Birbla archives

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression · Birbla