Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I often get a "verify you are not a bot" these days, which just needlessly wastes my time (because I need to click on something, and this in turn takes away seconds; multiply this like x50 per day and that's a time waster, now I need an extension to disable this crap check). So I am biased here.

    Most people will say "yay, it is great you waste the time of AI bots via ads!". Well, I already think ads should not exist in the first place, nor bots, but both exist - but the real issue is when websites now steal my time. That was different in the 1990s. I think mankind made several missteps here.

  • I get this so often too. And on so many websites, the verify your are a bot comes back after what feels like only a few minutes.

    Combined with how slow modern website are, I feel a sense of dread opening any website. Asking an LLM feels lower friction... but at what cost...

  • This is great, I think everyone wins here. Normal users like us get a standard, ad-free web experience, while those sending the most annoying web traffic not only get a smaller, simpler response (which should take less processing on the server side I imagine), but they are benefited by the content already being in a minimal simple format for LLM ingestion, PLUS they get to slip ads in that stream without affecting us normal readers.

    I'm also completely unbothered by the precedent of drip feeding product ads to someone's LLM chat history, and having that influence future conversations, because if you're going to outsource your buying decisions to an LLM, frankly I don't really care if you buy stupid products at that point - you brought that on yourself.

  • > A couple of the bots did not even get that. GPTBot and ChatGPT-User, the agents OpenAI uses for training and live fetches, came back 406

    I find that odd, If I were doing that, that's what I would want to target the bots used for training in order to get the ad content in the training corpus.

  • The advertiser might but Time might not. Time probably wants a content licensing deal before being used for training.
  • Maybe for ad attribution, when they inject campaign tracking information into the results and whatever LLM frontend presents them to an actual user and it results in a click, they could claim that the ad spend caused this click. If they inject the ad into the training data, it's harder to convince the ad campaign manager to spend more, because of missing ad attribution.
  • This is a neat idea. Might even make sense in cases where you actually want the AI bots hitting your site (we want it for our ecom platform, for example) - just serve Markdown content optimized for AI.

    Be careful not to do this with Googlebot, though. Google would consider this "cloaking" and could ban your entire domain for it.

  • The same things your bot reads are saved as material for training in the future by your provider, so this seems as an attempt to poison the training data. Just this weekend had a dinner with someone who insisted that part of his business is to seo promote business in chatbots and this looks as part of the infrastructure behind such efforts.
  • Next step the ads become increasingly sophisticated prompt injection to ensure they make it to the user, but the prices plummet in the meantime because the ads have low efficacy. They get purchased mostly by the scammers and by the time they start showing up in user chats they’ll be full on LLM assisted interactive user manipulation campaigns with far worse outcomes than the worst of YouTube and social media ads.
  • That's actually really neat, even without advertising I can serve humans my full page and AI bots a version with some parts removed. Or added.
  • I wonder if this also means that a normal user with a privacy-focused browser also gets this version. I can't see the point (except for click fraud?) of showing ads to AI, so it feels like this is more for the ad-blocking and privacy crowd.
  • why would that “crowd” present themselves as bots by modifying their user agent string? the point is to poison someone’s agent with sponsored “knowledge” over time.
  • The issue with that is it’s still a markdown page. I’d suspect it’s more about making sure people don’t get around ads, or possibly even cross-user memory/training.
  • I think this is less traditional ads, and more "content focused on poisoning the AI with very promotional statements"
  • I was on a NYT page and used Firefox mobile's "Summarize" button... it spit out some sort of unrelated cooking recipe. I'm guessing this type of bot interference will become more and more common.
  • IMO it's safer to skim the article than relying on AI summaries. You'll get mostly the same end result without the risk of disinformation.
  • Is the intention that some type of long-term context would be seeded with "ideas" for the AI to serve up if it is ever asked for bank recommendations? It seems kinda ad-hoc and untargeted, but perhaps for high cost services it might be worth it.
    by WJW
  • I'm wondering if it's also a way for this marketing agency to artifically (ahem) inflate the impression numbers they report back to Ally Bank, or whoever the customer is.
  • They might be trying to get into the context of long-running chat sessions.
  • sounds like plain old prompt injection. These days chatgpt might look at 50 webpages when i ask it to research a topic. Seems quite possible that the final answer is influenced by such ads.
  • This feels like the early days of SEO over again. There are no agreed-upon metrics yet, so you can sell all manner of snake oil. Hell, some of it might even work!
  • The best way for a politician to lie is to convince someone else of the truth of the lie and then put that someone in front of the cameras. That way, there's no hint of body language or anything else that indicates it's a lie. Both the denotation of the lie and the human context of the lie will be in harmony.

    This sort of reminds me of that. LLMs are by their nature credulous. They can be trained to not give in easily to some things, like the capital of the US, but in general they constitutionally have a tendency to believe what they read. What they read is basically their universe. There's only so much room and so much training data to really strongly pin raw facts in their weights. The only way they can not believe some marginal fact presented to them in their input is to possibly have read something that contradicts it in the same session... and the vast, vast majority of the world is those marginal facts, not really objective things like capital names.

    So if you can work a confident statement in to an LLM's input about some semi-relevant topic, it's truth to the LLM. And, being truth, the LLM will then happily and confidently elaborate on it quite a bit.

    Of course, if it's irrelevant to the current query, it probably won't have much effect. Ads have always been a game of numbers, anyhow. Even a query about a science topic has some probability of eventually turning to a question about banking in the same session. It's probably a good idea to rather strictly partition your conversations to stick to a single topic, not to defend against this but just to maximize the effectiveness of what is in the context window by keeping it focused, but I have to imagine there's plenty of people out there who reuse conversations all the time and end up with single conversations covering a huge array of topics.

    The good news, and the bad news, all at once, is that Google isn't going to take this one sitting down. If they're going to replace the search engine box with an LLM, well, they're using the same LLMs we're all using, if not in fact a bit cheaper one for the work they do, and by golly, that bot should be serving up Google's ads, not Time's ads! Who do these uppity content creators think they are, anyhow?! So there is definitely going to be work done in the field of ad-blocking content served to LLMs.

    by jerf
  • I would also like to read the stripped down markdown copy and not the original. All the time.
  • NetNewsWire "Reader View" is a truly wonderful feature: https://netnewswire.com/help/mac/5.1/en/reader-view.html
  • When the GDPR came into force https://npr.org started redirecting to https://text.npr.org with an explanation that sounded like that fact was supposed to annoy you, but I honestly find the latter superior especially now that they’ve added a few more lines of CSS to it.
  • The only website that I know has that has that is Daring Fireball (add ".text" to the end of the url), but the author of that blog invented markdown
  • Safari has a very pleasant reader mode that, while not pure markdown, does capture quite a bit of the experience by standardizing presentation and stripping out distractions. You can set Safari to automatically engage reader mode on websites you specify, when it detects an article.
  • There was a golden age where lots of websites had WAP (https://en.wikipedia.org/wiki/Wireless_Application_Protocol) versions available while also running their "normal website", and the WAP version was always like 1/100 of the size of the normal website. Still, "normal websites" were minimal compared to now, but when you were on modem, even those websites loaded slow. For some time, most of my browsing were via the WAP versions of the websites I visited during the bi-daily hour of allowed internet access.

    Maybe now we'll get something similar, just happens to be for LLMs, but for us who like less bloat, it can be a better viewing/reading alternative. Hope it spreads :)