Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Is anyone using something like this? I'm skeptical of scrapers paying and, if paying, paying enough to make it worthwhile.
  • I do appreciate that cloudflare is one of the few platforms offering options to web sites on how they want to engage with AI and even monetize traffic coming from it.
  • This is fine so long as it’s easy for me to turn off. I just don’t want to accidentally lose all AI traffic one day.
  • Do you actually get any signifcant AI traffic today?
  • Most of the internet unfortunatly ploinks fown a 'free' service and never even looks at what it defaults to. E.g., quite a lott of cloudflare "protected" sites block their rss feeds from being read by machine.
  • I tried this, blocking AI training blocked the Google search bots and cut my traffic in half. I would not recommend.
    by guyn
  • I wonder if this has anything to do with the cf bug that stripped all POST data from requests to a SPA I manage for 4-5 hours last week. That was a real good time, figuring out that it wasn't trying to show challenges or anything. Default setting for any web app protection from cloudflare should always be "off" unless you're under attack, and then who knows what settings will or won't break your configuration.
  • Has there been any update on the pay per crawl program?
  • So is it possible to say "No bots except Google, OpenAI, Grok, Claude and Perplexity"?

    As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.

    Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.

    By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?

  • Google is actively working on not sending you visitors anymore.
  • What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl.

    What would force their hand?

    It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.

    That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.

  • *AGOD
  • Pay to crawl is already here.

    The usage patterns of how people pay and use AI is basically the same model the web should be using: you pay a small bit of money to access monetized pages, just how you pay a small bit of money to get AI responses.

    It just needs people and browsers to get onboard with protocols. Crawlers will have no choice but to pay for content behind these 402 gateways.

  • "I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl."

    You seem to misunderstand. You pay or you die. There's nothing in between. Cloudflare will happily collect the tax. As does Apple (collecting 20B$ yearly from Google for the "tax"). Cloudflare's users also won't mind about how the company handles ADOG as long as they get a chunk of the cake by getting freebies and cheap services.

  • Universal tax collector of the internet. A penny for every page access. ADOG will not mind as it cements their incumbent status and pulls up the drawbridge by erecting a huge financial barrier for any new entrant.
  • I get a mixed feeling about all this. Cloudflare is unilaterally making all these decisions which impact the whole internet traffic flow. Taking the lead is one thing, however decisions like this should have the direct involvement of Internet Engineering Task Force (IETF) to account for all stakeholders, otherwise we run into the situation of a fragmented internet
  • I think the best answer is, nobody knows. The previous equilibrium for content scraping for search engines on the internet was already at times an uncomfortable one. But I agree that from a game theory perspective, "the AI bots take and give nothing back in return" is not just hyperbole, it's the actual situation. If Google is successful in what seems to be its plans and it becomes a box where you type a question and Google gives you an answer and only a vanishing fraction of the users click through to any underlying website, that instantly eliminates the entire value proposition for vast swathes of the web to actually be on the web.

    Something has to happen or Google will end up starved and locked out of everything, by means both technical and legal. Then nobody gets anything.

    I don't have the answer as to what happens next, and I doubt anyone else who proclaims one super confidently. But we can do some constraints analysis. There is no world where everyone works for free so Google and other AI engines can get all the value from the content, so we can eliminate those possibilities. I think we can safely discard the world(s) in which all content production just stops. However, off the top of my head, it's hard to get much tighter than that, and that definitely leaves a world where effectively everything everywhere ends up going pay-to-access.

    Microtransactions have, to date, failed comprehensively, though, so the constraints on what "everything is pay-to-access" gets weird without them.

    And there is never guarantee that there is any solution to any set of constraints. Things can end up overconstrained in reality as easily as a math problem. I don't actually think it'll go that way, but when analyzing this question I think it's important to not let "but $SOMETHING just has to have some way to work, because... uh... it has to!" Let the constraints do the talking. You could end up with a scenario where all content of any value is locked down, and it's fundamentally difficult and expensive to ever access or discover it, and consequently the entire content production industry radically contracts compared to its current size, if there is no pragmatic solution to microtransactions that is low-enough friction to get over the psychological and economic hurdles that have killed it to date. If everything is locked behind "macrotransactions" that's a much smaller commercial web. Probably a much higher quality one, too, but at a pretty stiff cost.

    by jerf
  • I find it unsettling that we are willingly outsourcing the decision on who can access our sites to an increasingly dominant corporate entity.

    The reasoning behind this is also flawed: blocking "bots" and "AI" means that our AI agents working for us are unable to do their work for us, because of knee-jerk bot-blocks.

    by jwr
  • Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go this route.
  • Adding the link to GitHub here if anyone is curious:

    https://github.com/techaroHQ/anubis

    by neya
  • Not sure why Anubis is getting so much hype on HN, but honestly, it is not the solution. A real solution would use behavioral modeling. Most browser fingerprinting issues are already largely solved anyway.
  • > Please consider installing one of the many PoW schemes such as anubis

    Why not go all the way and mine monero instead of just completely wasting the work?

  • Please don't use Anubis, it makes visiting websites very difficult (often multiple minutes wait times) on old hardware and low-end smartphones.
  • >I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites.

    ???

    Unless you're browsing around with the googlebot user agent string, you should be getting turnstile challanges at most, not blocks. And if you're getting a turnstile challenge it's unclear how it's different than an anubis challenge. If you're outright blocked, it's probably a site decision (eg. block all VPNs or block everyone not from a given country) rather than cloudflare's.

  • PoW schemes like Anubis don't work. Increasingly bots are using headless browsers and are basically able to solve captchas, proof-of-work(s) and basically bypass all any any attempts to block them. It's becoming impossible to stop bots from hammering your sites/services for unwanted traffic.
  • > For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

    It's kind of exhausting seeing Cloudflare playing both sides of the arms race.

    I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.

    > This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.

    And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.

  • I see Cloudflare as trying to forge the appropriate path ahead. Neither allowing the free-for-all nor trying to block everything isn't playing both sides, it's the path straight down the middle. Providing tools for producers and consumers to do things with permission and compensation.
  • How are they playing both sides? I thought their scraping products were also about having it behave and not take down systems
  • The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:

    > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).

  • Don't that feel like a threat to businesses who dare to avoid their content being stolen?