Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This paper focuses on the use of the providers' web and phone tools, and the data sharing arrangements they built through ad networks and tracking services like Data Dog. It doesn't talk about how they handle data from API calls. For one thing that can be pretty opaque.

    I have a friend/client who understandably doesn't trust the existing privacy policies of the major providers. They have the same problem many of us do: We want the most powerful models, we're willing to pay for them, but we see over and over how much of a frontier the frontier actually is. Frontiers are ugly if you don't have guns.

    So for now maybe platform tools like Open WebUI and TypingMind are a good workaround since the big boys don't (apparently) train on API data (for now) or (probably) send that data to advertisers. It would be interesting to confirm that.

  • Prior to this new AI age, your data was calculated and what was inferred about you was "shallow", but as of today. The sort of profile and things that can be known about you is scary especially if you are constantly engaged with cloud AI. IMO, the number one risk of using cloud AI is loss of privacy and loss of freedom. With AI and capabilities, more controls can be placed on people and the more you put yourself out there, the more you are going to lose.

    For example, we now have self driving cars, we have cameras everywhere. Based on your chat with a cloud AI, you can automatically trigger an automatic monitoring event that follows and tracks you in the real world with the fleet of cameras, cars, GPU, cell signal. Your tracking due to AI has moved into the real world and eventually, a self driving car will take you in to be "processed" against your will, not even for what you posted in a public forum, but for your private ribbing and chatting with some cloud AI.

    So definitely put local AI into the mix and keep personal stuff and thoughts local only.

  • A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you. They don't care about the agreements you've signed. You're just a pile of cash to them
    by 20k
  • This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!

    My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.

    Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).

  • In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.

    We have all become Milhouse now.

  • My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.

    Perplexity does this. Visiting a past perplexity search url exposes your full conversation.

  • It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't. Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.
  • Tangential:

    I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.

    This partial prompt data might potentially be used to "pre-warm" some kind of cache.

    But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.

Explore Birbla archives

AI companies leak data to advertisers [pdf] · Birbla