

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Absolutely LOVE the name of this. It’s cute and I immediately envision the behavior. Solid choice
- This seems like something that would be great to open source; I can see it being a good building block for other globally distributed services.by rzz3
- Will they need a Zookeeper for all the Meerkats I wonder?by sailfast
- This is really interesting. I wonder if etcd can benefit from this as it uses Raft for consensus and split-brain can be a problem depending on how its setup (usually when a Kubernetes control plane is initialized).by nunez
- Yeah, for real. etcd is one of the biggest limitations to Kubernetes in my opinion. I would love to have a region-agnostic Kubernetes cluster.
- dumb guesses
1. I'd bet they either already do log-routed config, or simulation testing finds this race before production does.
2. The first part of Meerkat that will need clocks for correctness will probably be stale local reads, not the core consensus protocol.
by crank42069 - I imagine something akin to a bat signal that alerts Aphyr to these types of posts. Looking forward to the Jepsen write-up.by _cairn
- TBH I’m doubtful of most people building their own crypto libs and distributed consensus implementations. But maybe cloudflare can pull it off. Good to see them pushing the state of distributed consensus.
Take aways:
* it’s not in prod yet. I suspect those many round trips are going to get expensive on median aka typical redistributed deployments. Curious to see how it goes once in the wild.
* they say it isn’t likely suitable for eg databases.
* they talk about formal verification, which is good and feels appropriate.
Looking forward to seeing more!
by pstoll - Cloudflare did not develop QuePaxa, although I'm sure they did a lot of work to operationalize it. The original protocol is from a SOSP 2023 paper.by michaelmior
- > I’m doubtful of most people building their own crypto libs and distributed consensus implementations. But maybe cloudflare can pull it off.
Who else than Cloudflare (or similar company in expertise and size) would be a better fit to implement distributed consensus?
by williamdclt - if youve ever fought a raft cluster on a bad network with leaders flapping, elections storming and latency spiking this genuinely doesnt seem that bad. i believe this will be very useful to those dealing with messy networksby ebeirne
- Looking at the implementation sketches, this algorithm looks even trickier to implement than Paxos (already a notoriously tricky algorithm to implement) and on top of that, I think the failure case in this algorithm is subtle and different -- very long tail latencies. In Paxos/Raft the latencies are more likely bounded by the timeouts (not eliminated of course) so you can build other systems to expect certain delays, but in this case, you may write something, wait for an ack, then abandon and retry, then realize the old write succeeded, etc. ad infinitum.by aabhay
- I feel like it would be much better if the article focused on QuePaxa because IMHO it's an algorithm that finally brings some novel ideas to concensus (e.g. not relying on timeouts) by kinda coming at it from a gossip protocol angle and is not getting the attention it deserves. The post shouldn't have focused and introduced Meerkat which hasn't been fully developed and tried in production. If they clearly presented the pros and cons vs not just Raft (which is popular but doesn't even play in the same league because it is relies on a leader) but other leaderless or multi-leader concensus protocols that would have been of greater value. The Paxos family of algorithms are a much closer fit here and there's a reason why some serious large planet scale systems choose it over Raft.
E.g. 1. Intro about issues with concensus 2. Intro to QuePaxa 3. Comparison to other algos that are close to it 4. Mentioning active work on implementation via Meerkat and intent to bring to production with followup posts.
As always when it comes to concensus it's all about trade-offs. And with QuePaxa that might be the increase in messages (note: I don't mean message round-trips). We'll see how it goes but it will definitely be interesting.
by eis - This sounds like a very direct approach to linearizability - just put everything in a linear order.
But this includes READ operations too! You have to get global consensus for every read! Most distributed systems have additional complexity and latency on write paths so that reads can be completely local. If you can accept slow read operations this seems like a great trade off, but I think that is going to relegate this to niece usage.
by advisedwang - The article says that if your read hits a master there’s no consensus needed. Perhaps I read wrongby aabhay
- What's interesting here is that this would be the first production implementation of an asynchronous consensus algorithm (QuePaxa). Paxos, Raft, etc. are all partially synchronous, meaning they rely on timeouts and only make progress if message delay is sufficiently small compared to timeout durations. QuePaxa doesn't rely on timeouts and makes progress even under wild fluctuations in message delay. The question is whether performance is competitive enough in the normal case, when message delay is small and doesn't vary much, and historically the answer has been "no" and that's why asynchronous protocols weren't used.by nano_o
- Wasn't there an impossibility proof for consensus without timeouts? At the boundary between a consensus failure and success, there must be a certain message that, if you delay it enough, causes a consensus failure, and that implies either you wait forever for that message (deadlock) or you eventually give up waiting (timeout).by inigyou
- This article is a bit hard for me to grasp the main ideas of because, given Cloudflare's requirements (e.g. no strong leaders), it immediately seems like they should be comparing to leaderless protocols like Paxos-class algorithms. Comparing to Raft and saying it's better because Meerkat is leaderless is confusing, because Raft is an adjustment to Paxos to specifically have strong leaders. So I'm 3/4 the way into the article and I don't see what's unique here.
I think the unique idea here is supposed to be QuePaxa's idea of avoiding timeouts for ensuring liveness. The actual discussion of QuePaxa is limited to one paragraph at the end, and tbh only a couple sentences of that paragraph.
I feel like the article could've been titled "Consensus protocols and linearizability: a brief explainer", or "Paxos vs Raft", or similar. It just doesn't feel like it communicates what it claims to communicate, and is a bit confused on who its audience is, just IMO.
by m11a - Agreed. When they describe the advantages over Raft, I can't help but read 'this is Paxos'.by mrkeen
- Agreed on - it took a long time to get to the “so what is new here” vs the broader topic of distributed consensus. Stylistically would prefer more upfront “here is what is novel here”by pstoll
- It feels like somebody prompted an AI agent at CF "why Meerkat is better than Raft" when drafting this blog post.by buremba
- Worth noting if it wasn't obvious from the article that Cloudflare did not develop QuePaxa. It's from an SOSP paper back in 2023[0]. The article is discussing what is the first known large-scale public deployment of the protocol.by michaelmior