Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Meta is most certainly addicted to bad data.

    I believe this because Facebook continues to send spam to an e-mail address that was only used by me for my cat, and only once; and the cat has been dead for 15 years. The dead cat address received three spams from Facebook just yesterday.

    There is also bad data out there about another cat that died ten years ago. He keeps getting snail mail from political candidates trying to convince him that they're deeply interested in the cares and concerns of people like him. A dead cat.

  • The biggest spenders in the marketing business are doing so primarily to justify their own size. That is to say, if you're a trillion dollar business, you are going to spend a billion dollars on advertising purely to say that you spent a billion dollars on advertising. The quality of the data - or the clicks - matters less than the fact that you've bought your brand a sense of ubiquity. It's fractally nested corporate classism.

    Direct response marketing - i.e. when you are buying ads specifically to get someone to buy a book - is a far smaller fraction of the market. It had a moment in the early 2000s when Google and later Facebook figured out how to extract lots of information about their customers and sell it on to other direct response marketers. But even then, the really lucrative marketers aren't legitimate businesses, they're scammers using the ability to micro-target ads to find their biggest rubes as cheaply and silently as possible.

    If you're a regular person trying to buy ads for a legitimate business, you're probably going to get swamped by all of this and taken for multiple rides by several different kinds of scam.

  • Various dodgy bots and crawlers hammer my website continually. Consequently, my web logs and analytics are garbage. So bad data is all we have, but it's better no data at all.
  • > A policeman sees a drunk man searching for something under a streetlight and asks what the drunk has lost. He says he lost his keys and they both look under the streetlight together. After a few minutes the policeman asks if he is sure he lost them here, and the drunk replies, no, and that he lost them in the park. The policeman asks why he is searching here, and the drunk replies, "this is where the light is".

    [1] https://en.wikipedia.org/wiki/Streetlight_effect

  • Everyone is addicted to bad data.

    The Red-Green-Refactor pattern is a coping mechanism to deal with the fact that if we want a change to work and the build process tells us we didn't break anything, we bowl right past any subtle hints that we are in the wrong, and our whole code change is a house of cards standing on a bad assumption that will immediately collapse when breathed on.

    I have a love-hate relationship with negative tests because of this, and I wonder if there's some way with static analysis or maybe AI to validate that the test that is green because nothing happened isn't green now because I broke the API and the test is now testing nothing in, nothing out instead of something in, nothing out.

    Sooner or later in some refactor someone finds a way to break the code without CI catching it.

    But we are just people. And if you squint you can see how our relationship to green builds is the same drive that management, sales, and marketing, and scientists get with charts that Make the Numbers Go Up even when the data is just correlated and the proximate cause they were looking for is a hallucination.

    Mark Twain knew. Lies, Damned Lies, and Statistics.

  • Yeah, but with one more piece of data they'll have solved personalized mattress sales forever. Nevermind that the entity collecting the marketing data is a convenience store chain and doesn't sell their data, allegedly.
  • I dunno. There is a lot of valid and interesting criticism to write about digital marketing. Lots of people have attempted to study the efficacy of digital advertising, and I'd love to see a deep dive into the "MarTech" ecosystem and how shady much of it is. Anyone who has had to work on the go-to-market side of a company knows the general feeling of frustration the author is experiencing.

    But this post reads like an aspiring "thought leader" posting a hot take to LinkedIn. It feels like lazy pandering to the "dumb marketers don't math good" crowd.

    Here's the same author with a post titled "Marketing and The Modern Data Stack" where he gets very excited about the marketing automation and big data, kicking off the piece with line "There is a huge transformation happening in the data space." and really sells the data-driven future with "In short, you need data, lots of it and it needs to be tightly integrated across the entire customer lifecycle.": https://www.jacquescorbytuech.com/writing/marketing-modern-d...

  • > Marketers are addicted to bad data

    This article cites a random Statista page for the "36% percent of people in the UK use an adblocker" stat.

  • If it gets you paid, it's good data.

    If it makes your clients happy, it's good data.

    Business data is for business purposes, not for science purposes.

  • > Business data is for business purposes

    I mean isn't the business purpose for the company running the ad campaign to actually increase revenue? Sure for an outside ad company just making the client happy is sufficient but for internal teams and the customer themselves the data is still bad.

  • sure, but generally most businesses do wish to pay for fraudulent activity.

    It seems we have reached a point in time where many "business people" seem unable to discern a difference between fraud and not-fraud, but whichever businesses those people are responsible for will not be sustainable enterprises in the long term.

  • Some people are in business to get paid by creating actual real value all the way down the chain to the ultimate customer. Those people are the ones you'd rather work for, and those are the ones who will care about "science purposes".
  • A lot of the time you are running experiments and seeing how the data changes, not looking at the data in a vacuum.

    Many of the example critiques here don't apply so much when looking at the changes in data:

    - if I got 50% more hits on my site this week vs last week, that's meaningful despite 36% of people blocking ads

    - if my open rate doubled when I changed my email subject, also meaningful

    The other examples are hard to pick holes in as they simply say "Z is a lie", but I can be looking at multiple data sources to decide how much of a lie Z is.

  • Following this kind of process blindly and optimising for it leads to a terrible product. However it can lead to more revenue in the short term.
  • The problem is that number optimization easily leads you down the path of optimizing the wrong thing.

    Need more people to click on your email? Easy. New subject line: “you’re gonna die soon”

    Need more people to click on your ad? Again, easy. Have you considered boobs and butts?

    These aren’t just made-up either. We’ve all seen those weird ass mobile game ads. Vague, maybe a little bit of fetish, sexual, violent, ominous. Those ads work. 100% they settled on those ads by optimizing their numbers.

  • Having been at smaller companies without the data, tooling, discipline, and resourcing to conduct viable experiments, and then being at a company that is actually one of the best in the world at it and building solutions for these problems at scale for marketing teams, so many problems come back to the human element in how data is tracked, how teams collaborate or fail to which can lead to pollution and dirty data, and how decisions are made as the moment trade-offs need to be considered it becomes personal and enters messy human relationship territory.

    But ultimately from a pure "did this work or not" standpoint you are right. Incrementality experiments are the gold standard.

  • This has been known a long time. Freakanomics did a two part podcast on does advertising actually work. It is long but well worth a listen if you want to hear some science about advertising.

    https://freakonomics.com/podcast/does-advertising-actually-w...

    The second part goes into internet advertising.

    https://freakonomics.com/podcast/does-advertising-actually-w...

    There are transcripts of the episodes on the page if you want to read instead of listen.

  • > This has been known a long time.

    This article predates that first Freakonomics episode by 11 days.

  • I work in adtech. We actually had to mangle some of our analytics data because it didn't match the broken data a client was used to. Our data was more accurate, but because the numbers didn't match the other analytics suite the client used, so we had to make it worse.

    I wish I was joking.

  • I work in HR. We had to mangle our salary analytics stuff because it kept suggesting that clients should pay more to attract talent. Like most employees in the same sector and position were being paid more that what they wanted to pay and that meant our analytics were wrong.
  • > I wish I was joking.

    I worked for a company that had a gui that ran on customers' desktops aka not our hardware. We were discussing a new feature but we weren't sure how many people used the particular module in the GUI that the feature would impact. There was a discussion about how we could find out which users had that module in their layout and one of the devs says "Let's ask Tom. He's the product manager and he talks to the clients. He'll know!"

    I then asked: "If you don't mind my asking, where are the individual GUI layouts saved?". I asked this b/c at a previous job I also worked with a desktop GUI but the layouts were stored on the user's desktop which made it tough to access.

    The dev replies to me "On a server", I respond: "On a server we own and have access to?", dev: "yes"

    To which I replied: "Why don't we go look at the gui layouts on the server instead of talking to Tom?"

  • The average time on site and bounce rates I have got from advertising my software inside ChatGPT and Reddit are so terrible that I can only assume that most of the clicks I am paying for are fraudulent, although it is not clear who is doing the fraud.

    https://successfulsoftware.net/2026/08/13/my-experience-buyi...

    https://successfulsoftware.net/2025/08/11/what-i-learned-spe...

  • I mean, paid traffic will almost always be worse than people actively seeking you out or finding you organically.

    But I think there are definitely some opportunities for boosting your marketing/product strategy. Maybe ask ChatGPT for guidance on optimizing your advertising/marketing efforts?

    I think there's enough low-hanging fruit (landing page that offers a free trial behind an email sign-up) that it'd probably help.

  • I have a hypothesis that that majority of ad clicks are misclicks where someone closes the new page within <1 second.
  • The answer is at every level, and each level is separated into multiple sublevels of fraud.