

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I wouldn’t trust an AI company to honor this as far as I could throw themby gdiamos
- In a world where people's searches are already mostly answered by AI, is there any point to disallow the training?
I could totally see a future where Google just stops showing you links to actual web pages altogether and just gives you their chatbot.
by OroPla - > Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.
> We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.
I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.
by AnonC - I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots.
[1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth
by mskalski - I'm playing around with it for my MCP hiring protocol ojcp[1] and it seems to work very well for signing attestations at the header level.by fraywing
- What does this really mean though? You can use an LLM to search.by nullbio
- I maintain a cloud IP ranges database, and I'm going to test this out.
I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:
"Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."
by kinduff - The irony is that search engines are AI companies now. Telling them 'index me for search but don't train your models' is asking them to split a brain that’s already fully merged.by qsbuilder
- Exactly. I don’t get this. You may decide to not train but your content will still show up in search.by tchalla
- Even looking for companies which supply services is now far better on AI chats than Google. For me being visible in AI training is going to be more important than search in the next year or two.
If I was a big AI company I'd certainly be tempted to make sure that anyone who excluded themselves from "AI training" also got themselves excluded from AI results.
by VBprogrammer - "Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.
- ...and the only way to stop[1] that is by effectively DRM'ing everything, which is a level of dystopia that I don't think even Stallman ever anticipated, nor do I want to happen.
[1] Analog hole and other workarounds aside, naturally.
by userbinator - I didn't know websites could opt out of providing data to Google's AI training. Looks Google added support for this via 'Google-Extended' in robots.txt back in 2023:
https://blog.google/innovation-and-ai/products/an-update-on-...
by skybrian - "Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."
Is that really true
CF classifies anyone not using a popular browser with Javascript enabled as a "bot"
CF fingerprints www users
As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused
IA = Internet Archive
- As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.by devmor
- CF classifies anyone not using a popular browser with Javascript enabled as a "bot"
This is my biggest complaint about CF. They are implicitly supporting user-agent discrimination in favour of Big Browser, instead of discriminating on actual behaviour.
...and of course there are already companies running tons of VMs with "officially sanctioned" browser + OS stacks, that can get past all these "protections", for a fee.
"AI bots" is the newest boogeyman they came up with to take away freedom.
by userbinator - I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway.
Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page".
If you don't want your content to end up in some database/archive don't publish it for the whole world to see.
by DharmaPolice