Comments

Hacker News

I wonder if this has anything to do with the cf bug that stripped all POST data from requests to a SPA I manage for 4-5 hours last week. That was a real good time, figuring out that it wasn't trying to show challenges or anything. Default setting for any web app protection from cloudflare should always be "off" unless you're under attack, and then who knows what settings will or won't break your configuration.

by noduerme

Has there been any update on the pay per crawl program?

by graeme

So is it possible to say "No bots except Google, OpenAI, Grok, Claude and Perplexity"?

As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.

Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.

By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?

by TekMol

What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl.

What would force their hand?

It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.

That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.

by holografix

I find it unsettling that we are willingly outsourcing the decision on who can access our sites to an increasingly dominant corporate entity.

The reasoning behind this is also flawed: blocking "bots" and "AI" means that our AI agents working for us are unable to do their work for us, because of knee-jerk bot-blocks.

by jwr

Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go this route.

by fc417fc802

> For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

It's kind of exhausting seeing Cloudflare playing both sides of the arms race.

I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.

> This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.

And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.

by tekacs

The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:

> Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).

by simonw

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • I wonder if this has anything to do with the cf bug that stripped all POST data from requests to a SPA I manage for 4-5 hours last week. That was a real good time, figuring out that it wasn't trying to show challenges or anything. Default setting for any web app protection from cloudflare should always be "off" unless you're under attack, and then who knows what settings will or won't break your configuration.
  • Has there been any update on the pay per crawl program?
  • So is it possible to say "No bots except Google, OpenAI, Grok, Claude and Perplexity"?

    As far as I can tell, Google is the only one sending me visitors. And the other big AI players might do so in the future.

    Another option would be "No anonymous bots". So at least if a bot would want to crawl my site, they would have to identify themselves. Since the rise of the AI bots, I am getting hurt badly with insane amounts of requests from residential IPs that mimic real humans. The only difference being they don't make me any money. Only produce costs.

    By the way, how is the situation over at Amazon's Cloudfront? Do they offer something that helps? Anyone here with them?

  • What’s the end goal for Cloudflare and the web here? I don’t think ADOG (anthropic, deepmind, openai, google) is going to pay to crawl.

    What would force their hand?

    It’s more likely they’ll strike undisclosed agreements with major sources of discussion like reddit etc.

    That’s not to say getting new information as a way of context-providing is not going to happen but that’s not scraping.

  • I find it unsettling that we are willingly outsourcing the decision on who can access our sites to an increasingly dominant corporate entity.

    The reasoning behind this is also flawed: blocking "bots" and "AI" means that our AI agents working for us are unable to do their work for us, because of knee-jerk bot-blocks.

    by jwr
  • Please consider installing one of the many PoW schemes such as anubis rather than use these cloudflare "features". I increasingly encounter outright blocks rather than any sort of captcha when visiting cloudflare "protected" sites. Each individual site isn't particularly important to me but it's depressing to watch the process unfold like this. You really are choosing to erode the core basis of the internet if you go this route.
  • > For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.

    It's kind of exhausting seeing Cloudflare playing both sides of the arms race.

    I just can't imagine bringing myself to use their technology to build agents and build AI products when they're also doing things like this.

    > This also lines up the incentive model we want to foster. Losing trusted status across the more than 20% of web domains that sit behind Cloudflare is a deterrent with teeth. Trust becomes something you can carry with you, and something you can lose.

    And even more so, LLM language aside, fun and fascinating to see them flagrantly calling out their position here as if it's a positive.

  • The big news here is that Googlebot will be blocked from September 15th onwards by one the "block training" policies, because Google use the same crawler infrastructure for their search index AND for training Gemini:

    > Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).