Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I've been building latency-sensitive LLM systems for a while, and I've come to rely heavily on pre-fill-considerate mechanics like ping-pong overlapped async context construction. For interactive mechanics, the worst case, even if rare, is problematic.
A toy/simplified version lives here: https://github.com/chaboud/goulash
Consideration of mutation rate (a sort of temporal Shannon-ish coding/ordering) lives in there (with some RoPE-friendly structuring). Note: That was a vacation project, not the day job, but similar principles apply even with larger models.
by chaboud - The settings may change, but the two major problems in CS remain the same: cache invalidation, naming things, and off by one errors.by epistasis
- >Cross-request KV prefix caching is the largest practical lever in agentic LLM serving. It's why your coding agent's fiftieth turn costs a fraction of its first.
This is incorrect. If you are generating the 51th turn, then prefix caching makes the turns 1 to 50 cost a fraction of what they normally cost, to generate the 51th turn. Generating the 51th turn is more costly because it is not in the cache yet.
Unfortunately there is no way to steelman the statement, since quadratic scaling of attention means that the cost goes up even if you assume perfect caching of past inputs.
What the author actually means is something more mundane. The initial prompt is massive in a standard coding harness and this means prompt processing takes much longer than expected, creating the misleading impression that costs go down.
Edit: I noticed the word cross request too late. In that case the harness prompt is expected to be cached from other users, in which case the first message actually has a massive unfair advantage.
by imtringued - I work on inference at a neocloud, but opinions are my own .
The economics and thus tools you can throw at inference change at various scales . As a “blunt” contrived example , on a gb300 the GPUs communicate super fast over nvlink, and the cards can offload kv cache to dram and then disk, “fast enough “ for these tool heavy agentic workloads.
Which come together to mean that at high enough scale and in the right scenario, we can work with wild ttls on the kv cache and still comfortably hit SLAs and tokenomics.
by talolard - LRU seems like the ideal strategy for most things LLM-related. Everything in this realm is about recency bias. I think it is a feature in this context, not a problem.
When I give an agent a piece of corrected information regarding a long running task, the last thing I want it to do is try and statistically compensate for the fact that it is new information. I want this new information to dominate the old information.
by bob1029 - This is almost unreadable. The "papers" are never referenced anywhere, so the claims being refuted cannot be evaluated. The whole fact of the TTL doesn't seem relevant at all. There are many, many well-researched admission and eviction policies that this readme doesn't mention. I just don't get why we are reading this.by jeffbee
- > It didn't work, and why it didn't work turned out to be more interesting than the policy would have been.
Spoken like a true Claude.
Snarking aside, I am glad that our AI agents make it cheap enough to do these experiments and publish these write-ups that people finally bother to publish null findings. Very useful!
by eru - What I feel a bit annoyed by, and what I feel obviously LLM-run ablations like this fail to capture, is any kind of reflection around previous research or any kind of proof that this is the best you can do. You don't know, you pulled the lever and you got something, is the best? Can you do better? What is the constraint?
As an individual researcher, you do not have 22M$ to run a massive brute-force search for your problem. You are constrained to your little subscription and you will barely dip your toe in the sea of possible solutions to a problem. So letting Claude run an autoresearch loop on your problem and then having it summarize it for you brings 0 value because you dont know what the downsides and trade-offs of LRU caches were, and how you would possible solve it.
by augment_me