Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • For me the killer feature of Kafka was the ability to set the offset independently for each consumer.

    In my company most of our topics need to be consumed by more than one application/team, so this feature is a must have. Also, the ability to move the offset backwards or forwards programmatically has been a life saver many times.

    Does Postgres support this functionality for their queues?

  • The camps are wrong.

    There's poles.

    1. Is folks constantly adopting the new tech, whatever the motivation, and 2. I learned a thing and shall never learn anything else, ever.

    Of course nobody exists actually on either pole, but the closer you are to either, the less pragmatic you are likely to be.

  • Has this person actually benchmarked kafka? The results they get with their 96 vcpu setup could be achieved with kafka on the 4 vcpu setup. Their results with PG are absurdly slow.

    If you don't need what kafka offers, don't use it. But don't pretend you're on to something with your custom 5k msg/s PG setup.

  • I really believe this is the way: Event log tables in SQL. I have been doing it a lot.

    A downside is the lack of tooling client side. For many using Kafka is worth it simply for the tooling in libraries consumer side.

    If you just want to write an event handler function there is a lot of boilerplate to manage around it. (Persisting read cursors etc)

    We introduced a company standard for one service pulling events from another service that fit well together with events stored in SQL.

    https://github.com/vippsas/feedapi-spec

    Nowhere close to Kafka's maturity in client side tooling but it is an approach for how a library stack could be built on top making this convenient and have the same library toolset support many storage engines. (On the server/storage side, Postgres is of course as mature as Kafka...)

  • How do you implement "unique monotonically-increasing offset number"?

    Naive approach with sequence (or serial type which uses sequence automatically) does not work. Transaction "one" gets number "123", transaction "two" gets number "124". Transaction "two" commits, now table contains "122", "124" rows and readers can start to process it. Then transaction "one" commits with its "123" number, but readers already past "124". And transaction "one" might never commit for various reasons (e.g. client just got power cut), so just waiting for "123" forever does not cut it.

    Notifications can help with this approach, but then you can't restart old readers (and you don't need monotonic numbers at all).

  • You have to be careful with the approach of using Postgres for everything. The way it locks tables and rows and the serialization levels it guarantees are not immediately obvious to a lot of folks and can become a serious bottle-neck for performance-sensitive workloads.

    I've been a happy Postgres user for several decades. Postgres can do a lot! But like anything, don't rely on maxims to do your engineering for you.

  • >The claim is that it handles 80%+ of their use cases with 20% of the development effort. (Pareto Principle)

    The Pareto principle is not some guarantee applicable to everything and anything saying that any X will handle 80% of some other thing's use cases with 20% the effort.

    One can see how irrelevant its invocation is if we reverse: does Kafka also handle 80% of what Postgres does with 20% the effort? If not, what makes Postgres especially the "Pareto 80%" one in this comparison? Did Vilfredo Pareto had Postgres specifically in mind when forming the principle?

    Pareto principle concerns situations where power-law distributions emerge. Not arbitrary server software comparisons.

    Just say Postgres covers a lot of use cases people mindlessly go to shiny new software for that they don't really need, and is more battled tested, mature, and widely supported.

    The Pareto principle is a red herring.

  • My general opinion, off the cuff, from having worked at both small (hundreds of events per hour) and large (trillions of events per hour) scales for these sorts of problems:

    1. Do you really need a queue? (Alternative: periodic polling of a DB)

    2. What's your event volume and can it fit on one node for the foreseeable future, or even serverless compute (if not too expensive)? (Alternative: lightweight single-process web service, or several instances, on one node.)

    3. If it can't fit on one node, do you really need a distributed queue? (Alternative: good ol' load balancing and REST API's, maybe with async semantics and retry semantics)

    4. If you really do need a distributed queue, then you may as well use a distributed queue, such as Kafka. Even if you take on the complexity of managing a Kafka cluster, the programming and performance semantics are simpler to reason about than trying to shoehorn a distributed queue onto a SQL DB.

Explore Birbla archives