

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via OR.by mordae
- I was interested in reading this until my slop detector went off at the paragraph starting with “It's worth being precise about where the benefit comes from, because it isn't raw speed.”
I love AI, but I really hate reading it.
by brokenodo - > If squeezing the best open models onto GPUs and serving them to millions of developers sounds like your kind of problem, come work with us.
What is the typical job title and/or skillset for this?
- Why int4? There are a lot of superior 4 bit formats like nf4 from bitsandbytes.by om8
- So they quantize models, only tell about it in the blog post (instead of a warning on the model page), and even in the blog post pretend there's no difference by benchmarking on small context tasks many of which are saturated. Coding agents will probably be severely negatively affected by KV quantization.
I'd say serving quantized models without saying so on the "store" page is fraud.
by lostmsu - > View pricing in the Cloudflare dashboard ↗
Why… I wanted to see if it’s worth it to use cloudflare’s endpoint but I can’t even see the pricing
by syntaxing - I think Cloudflare not providing ZDR on their inference is the biggest public indicator that Cloudlare glows.
We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.
by HDBaseT - Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation.
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
by scrlk