

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- For reference, Semrush shows some statistics on these domains & how much traffic they are estimated to be receiving from organic search:
- wifitalents.com, peaked 15 July with 18k visits & declining
- worldmetrics.org, peaked 27 Jul with 8k visits & declining
- gitnux.org, peaked 20 Aug with 8k visits & declining
by arlattimore - I've been vary of using ai to search considering all the spam out there. I think I'd rather, perhaps naively, whitelist wikipedia, reddit, arxiv, some news sources, etc than include everything.
Is there nothing out there that does this? I'm paying for kagi and I can see that it has an api, is that maybe sufficient if configured properly?
by CapsAdmin - Reddit is full of ai accounts now thoby kingkawn
- I direct them to search for discussion on fora. There are still legacy sites on niche topics that the bots don't post in. This is particulaly useful for reaching into the past because traditional search doesn't surface anything but new content.
- If you use Kagi Assistant, you can pick one of your lenses (i.e. lists of domains to restrict searches to) in chats. Not sure if their API has that as well or some other way to restrict searches. Also not sure if the Assistant (or API) respects blocked domains when searching.by 1313ed01
- What protection do LLM search engines have against training off content generated by other LLMs?
Will we get to a point where AI-generated sites make up a majority of the internet, and LLMs are training upon their own regurgitations, with exponential amplification of all their lies and flaws?
Or will the pre-2022 corpus human knowledge be considered the low-background steel standard, and anything after that less and less reliable unless certified that it has been created by a human mind and untainted by hallucinations?
by sph - Will we get to a point where AI-generated sites make up a majority of the internet
I dunno if they'll be the majority (I suspect we're alredy close to 'yes, and it's already happened'), but I feel very, very confident that they will be the majority, if not the totality, of sites that the vast majority of people see.
by kjs3 - They’ll train on prompts and anything else you send in. Many LLM responses are sorta finger printable: I assume this is intentional
- I've mostly stopped using the Internet to learn new things and have gone back to books from the library. The majority of technical books at the library were published pre-2020s and hopefully, publishing slop physically won't be profitable enough to flood that market, too. Now that the Internet has largely been destroyed by slop manufacturers, whether or not the words are(/were) worth putting on paper becomes a useful discriminator.by coldpie
- I have a feeling we're already there.
- > What protection do LLM search engines have against training off content generated by other LLMs?
You're talking about a scenario that won't blow itself up in the next few quarters, so it's of no interest to them.
by gdulli - I’m so happy to see the negative Perplexity posts today. I felt like I was taking crazy pills hearing people think this service was at all useful/trustworthy.by jkahrs595
- "Sharing a nameserver pair is strong circumstantial evidence of a common Cloudflare account rather than proof of ownership"
-when you read one statement that let's you know to believe no other assertions in the article....
by rcar1046 - We should check out the ownership of this site, plus the others that the poster posted in the last 12 days...by ljf
- Perplexity is about to learn that Google is an anti-spam company first, search engine secondby alangibson
- Eh.. Using page rank in 2026 is like using a bow and arrow in 1918
- Well, was. They did a bad job of it these last few years which allowed any AI that could crawl the web to seem amazing for search because it could pick the best posts from Reddit or whatever other forum had the best context for your question, but now we're watching the AI snake eat its own tail.by threetonesun
- Interesting that this remains on the front page @dang - this and the other 3 sites the poster submitted (two others yesterday and one 12 days ago) all appear to follow the same AI generated pattern, all newly registered, AI written and with vague 'about us' pages. I'm not sure if Jakob Greenfeld registered/owns them all, or if it is linked to his marketing/sales business - but it is rather fishy.by ljf
- @mentions aren’t a thing on HN. If you want to contact Dan and Tom, use the “Contact” at the bottom of the page. They are very responsive.by latexr
- It's difficult to read more than a few sentences, when this itself is clearly a Claude artifact.by jpimbert
- AI;DRby MasanskY01
- I had the opposite reaction: maybe it was not written by AI, but it would have been better if it was.by anigbrowl
- Indeed. A google/brave search on that founder generated nada.by samuell
- Weird, there are 2 negative articles on the HN front page right now about Perplexity, both from “research” sites that are clearly LLM slop themselves.by iamacyborg
- Agreed, I thought the subject matter was interesting enough to try and labour through the tedious prose, but once I got to "Their scale is the point." I just had to stop and just skim the rest.by shrikant
- There's also the irony that this AI written piece criticizes how low the domains are on the tranco list when trellner.com doesn't even make the list, haha.by lukeinator42
- I do think models currently don't have enough source skepticism.
If you look at agent traces when asked to compare two options to help inform a decision, many of the comparison pages cited in research are often hosted by one of the companies being compared; nearly all are AI-generated AEO plays. Not deeply considering the motive of published information is currently a glitch that can be exploited, but the window will close.
I'm sure model providers will set up some crappy pay for play verification system for "trusted" product information, comparisons, and reviews.
by toddmorey - Or reddit - lies writen by other LLMs or by paid shills.by rvba
- Sounds like a tricky problem that will get a low-tech solution like a blacklist, whitelist, or chatGPT-approved vendor list.by alansaber
- There are a lot of ways to make this a lot better easily. First of all, they could use a blacklist of sites that sell guest posts on adsy/etc.Also, if an article only links to one of the products listed, or only one is a dofollow link, they should also be excluded.
That'd probably cut down on a huge portion of spam by itself.
by jrhizor