Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This makes sense. Why would everyone be coted for the same answer
  • Nice study :) Keep updating!

    The title is very misleading - authors openly state the low and slewed sample pool.

  • > 94.8%

    This seems consistent with Sturgeon's Law.

  • We don't have a page rank mechanism. Most likely people are paying or threatening AI companies to boost certain sources as authoritative
  • Makes me wonder if paid Medium and paid news are input to AI training. Surely, they are.
  • This makes sense. A handful of websites hold most of the “trusted” info because they’re massive. I don’t expect you to quote my blog with only three entries. The real trick would be getting AI companies to stop hammering sites that don’t show up in answers, but even if they don’t use a source for an answer, crawling still provides value.
  • It's also basically impossible to block most AI crawlers. Most completely ignore robots.txt. Some claim to respect it but I've found the majority that claim to also ignore it. You have to block their useragent and then their IPs if they spoof it and use botnets like some China-based AI crawlers.
  • Odd results for me. Last month I tested five AI chat sites, with web search turned off, and four of them had a shadow of information about me as a person (what kind of books I write, what tech I use, and a random bit of other information). The linked site gave me a zero score because it was testing if the AI models recommended my site for business or sales queries.
  • Blocking known bot identifiers via robots.txt does nothing by the way. Too many labs are running sneaky crawlers that do not respect robots.txt. You will need to take extreme measures: blocking basically all datacenter IP ranges, VPN IPs, aggressive rate limiting, etc.

    Blocking LLM crawlers has become the number one use case for our IP database customers at https://focsec.com/

  • This is why we need the HTTP 402 standard to become common.

    If websites charge pennies per AI crawl, they will make more money than ever being reference in that 6.2% of websites that get cited (of which even another small percent get any follow through that leads to a sale or ad click)

    HTTP 402 also basically extends the pay per token model people have gotten used to with AI model providers, except applied to the whole web, with the added privacy benefit in that there is no need for sellers of content to “know their customer”, and indeed it may even be impossible to do so because of how the payment gateways operate.

  • Seems like the opposite if the problem you’re trying to solve is “my site isn’t cited over someone else’s”. HTTP 467 - pls cite me, I’ll pay you.
  • The problems with micropayments are

    1. The market for lemons. In fact sites trying to "monetize their content" are the most likely to be lemons, so just asking is a signal that your "content" is not worth anything.

    2. If your information does have value, it competes in a market with other sites full of high quality information that aren't trying to monetize it, meaning your specific information has to be specifically very valuable and not available yet for free elsewhere. If it is valuable as information (i.e. not something like creative writing), it will quickly spread and become freely available.

    Lots of "content" just isn't worth anything (or has negative worth: it wastes your time). e.g. consider youtubers begging to get viewers to like/subscribe to increase their reach, and people generally don't despite it costing them nothing. Because it's not even worth a click to them.

  • Are we saying that it's now a problem that we're not getting scraped?

    This appears to be a new generation of "SEO", marking itself as a service for getting into AI results?

    This is not the future I want to be a part of.

    Perhaps it's inevitable that after a break from everything being driven by money that LLMs will now be ruined by people spending $X to get into AI to make back $X+1, leading to an arms race of ever increasing X, to the detriment of users.

  • I get that pushback. My goal with this index/report wasn't to write a new 'AI SEO' playbook or encourage people to start gaming the system. The goal was simply to shine a bit of light on what's going on. Right now, there is a massive information asymmetry: AI companies are turning their assistants into primary search engines, but webmasters have zero visibility into whether their sites are actually being cited in those live answers.

    You're absolutely right to be concerned about an arms race. If the data showed that doing X, Y, and Z guaranteed a citation, we'd be right back to the worst days of keyword stuffing.

    But what this data actually shows is the opposite: even if you do everything 'right' (allow the retrieval bots, provide perfect machine-readable schema, don't block anything), you still have a 94.8% chance of never being cited. How LLMs cite is entirely opaque.

    I built this to give site owners a baseline measurement of what is actually happening to their content today (as a free, anonymous view), so they can make an informed decision on whether keeping their doors open to these crawlers is actually worth it.

  • > Are we saying that it's now a problem that we're not getting scraped?

    No, it is saying that even the ones that are being scraped are not being included in results

  • LLM's were obviously going be ruined by two factors:

    1- LLM Crawler Optimization: the new SEO

    2- Weightings-For-Pay: for a fee have your product or service come up more frequently in associated answers

  • This is a weird complaint.

    Let’s say I ask “who created Linux?” Claude correctly tells me Linus Torvalds, and links to Wikipedia.

    There are probably thousands, maybe hundreds of thousands of other sites that have that some piece of information. Are LLMs supposed to link to every site?

    Most sites do not have unique information at all, and even sites that do rarely contain only unique information. Implying that most sites deserve links because they were crawled seems like a statistical fallacy. You could say the same thing about the percent of sites crawled by Google versus ever showing up on first page of results.

  • I just asked ChatGPT-5.6 that question, no source was given.

    Not even Wikipedia.

  • I wouldn't even consider it a complaint. For a builder, I would consider it an opportunity...
  • This is not a complaint.

    I think your assessment of this result being as expected... but this is about the LLM equivalent of SEO becoming an area of interest

  • Its a marketing post for a tool that offers visibility to site owners. Honestly these type of "neutral data reports" masquerades should be flagged or adequately disclosed.
  • The question is very important. What if you ask/tell an AI "I need a 2 bedroom vacation rental in Park City for a trip this fall". If you only get Airbnb results you are missing a lot of the true answer.
  • It isn't a complaint. It's an ad. This is literally an ad for some sort of "get cited by AI" service.
  • If the goal of the web was just to transmit objective facts like 'who created Linux', this wouldn't be a problem at all. Wikipedia handles that pretty dece.

    The issue arises when we move away from objective trivia and into subjective, localized buying intent, which is where the web monetizes itself.

    If I ask an LLM 'who created Linux?', there is one right answer. But if I ask an LLM 'who are the best commercial roofers in Houston?', there isn't one right answer. There are dozens of highly qualified local businesses that do possess unique value, unique pricing, and unique availability.

    When an LLM answers that roofing question, it typically cites 3 to 5 businesses. The other 40 legitimate roofing companies in the area are left out. My study isn't arguing that every single one of those 40 companies deserves to be in the answer; it's pointing out that those 40 companies currently have no idea they are being left out.

    In the Google era, if you weren't on page 1, you could look at Search Console, see your ranking, check your backlinks, and understand why. In the AI Search era, businesses are being scraped to build these answers, but they have zero telemetry on whether they are actually making the cut. This index is just an attempt to provide that missing telemetry.