Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- I have relatively little respect for Anna's Archive compared to other shadow libraries. They basically have just copied other shadow libraries archives and are much more aggressive about monetizing than the long-standing alternatives.by whimsicalism
- In my experience, ZLibrary was far more aggressive about monetizing (or is, haven't used them in a while)
- LLM corporations should be paying authors to read their books and benefit from them. Instead, Anna wants the corporations to send money to Anna?
It's hard not to read this as giant offense to the authors. I didn't think anything would be worse than DRM, but corporations paying pirates to steal books is right up there.
by WolfeReader - > LLM corporations should be paying authors to read their books and benefit from them.
I don’t think you realize just how huge the holdings of the shadow libraries are now. They have publications from all over the world, in myriad languages. (Someone has made a tool to visualize ISBN-space on Anna, I think it was posted on HN a while back.) It’s not realistic for a corporation, even a multinational titan with a large staff, to track down and compensate even the living authors, and a substantial amount of authors are dead and the current copyright holders are unknown.
by TFNA - I don't understand why this is a movement that is ethical to get behind.
Someone spends months or years of their life dedicated to writing a book. And people celebrate the fact they can get it for free, justify it by saying it's not free to search or host this content and offer to donate to piracy sites.
Rather than... Just supporting the author and buying their book?
It's different when this is American education and you're effectively being forced to buy books otherwise. I can understand fighting against that. But most stuff on the archive isn't that. It's just plain old piracy.
Yes a PDF or epub doesn't cost money to "print". Yes no one is "losing" money. But this isn't Netflix or Hollywood who still making billions regardless of piracy. Most of these authors are just regular people.
And the whole preservation angle makes sense when the books are no longer for sale. It's hard to argue preservation when you're linking to or hosting these works the second they are available to download. I'd be much more inclined projects that time walled the data, so you could effectively argue it's for preservation.
by Philip-J-Fry - You can't just start preservation "when the books are no longer for sale." It has to happen asap, there's no telling when something will get harder to find.by Cider9986
- Piracy never stopped the music industry, and the folks who were harmed the most by music piracy were the poor, cash-strapped billion-dollar corporations whose entire operating models already depended upon sucking wealth out of the actual, struggling artists who do all the work.
And it seems that piracy has become a net benefit to new and niche artists. (https://www.sciencedirect.com/science/article/abs/pii/S01676...)
I'd posit that the book industry will turn out to be the same. Piracy will harm the bottom line of the companies already at the top while giving exposure to the authors at the bottom. The latter being the ones who often strong-armed into terrible financial deals just to gain access to book-industry's four big gatekeepers, and who likely need that exposure to help keep a roof over their heads.
Anecdotally, I'm one of those folks who end up purchasing many of the books I pirate or otherwise obtain for free, and I'm sure I'm not the only one who does this.
by dentemple - Disallowing copying and sharing of art is a recent development in human history, not the norm.
The normal distribution of music and stories was for others to repeat them, and only recently have we decided it's illegal. I understand that things are different now, and people make a living off of art, but at the same time I find it difficult to care too much for someone who chose to make their hobby their job and refuses to adapt when things change.
by ghusto - > I don't understand why this is a movement that is ethical to get behind. Someone spends months or years of their life dedicated to writing a book. And people celebrate the fact they can get it for free.
Academics have never really made any money off their published research, but rather are paid via their institutions or grants. The publishers make money, but academics themselves are aghast at the publishers taking their edited collections and monographs, doing no proofreading or even no typesetting (that obligation is often on the authors and editors now), and selling the book for hundreds of euro. That’s why authors will almost always send you the PDF for free if you email them.
The celebration is easy to understand if you are a researcher. Getting ahold of publications that your institution doesn’t hold or subscribe to is always a hassle, it really slows you down during the writing process. The shadow libraries turbocharge research. Over the last several years, shadow libraries have gone from a niche to something that pretty much everyone in my field uses daily.
by TFNA - I agree, but also you can't wait until something is out of print/unavailable to preserve it. Trying to prevent access to it or limit distribution will probably just result in it being lost media one day.
There's also the fact that just because a something is available to purchase in one country, doesn't mean it's available in other countries. A lot of movies/books/games/etc are geo-restricted in sale, with many countries having no valid methods to acquire them.
The best (but unrealistic) solution would be for people who can purchase legally to do so, while leaving it available for download for everyone else.
by mitkebes - Books worth buying usually have rabid followers who will buy them.
There's been a reasonable amount of research that suggests that piracy doesn't really cannibalise sales from those who can afford to pay.
But I do agree that for some of their categories a time wall would improve their optics.
- I use AA and buy books. Typically I may start a series on AA epubs then buy the books. Sometimes authors take money directly (patreon, straight donations, etc) which is how I would rather pay them than pay the publisher for them to only get a small cut.
Are libraries unethical to use? You can go to your library and read books without paying for them.
by j_w - >I don't understand why this is a movement that is ethical to get behind.
Because we broke copyright. There is room to quibble about exactly where and when, but the result is quite clear. The best summation I know of is from a speech by Thomas Babington Macaulay in the British House of Commons in 1841[1],
"At present the holder of copyright has the public feeling on his side. Those who invade copyright are regarded as knaves who take the bread out of the mouths of deserving men. Everybody is well pleased to see them restrained by the law, and compelled to refund their ill-gotten gains. No tradesman of good repute will have anything to do with such disgraceful transactions. Pass this law: and that feeling is at an end. Men very different from the present race of piratical booksellers will soon infringe this intolerable monopoly. Great masses of capital will be constantly employed in the violation of the law. Every art will be employed to evade legal pursuit; and the whole nation will be in the plot. On which side indeed should the public sympathy be when the question is whether some book as popular as Robinson Crusoe, or the Pilgrim's Progress, shall be in every cottage, or whether it shall be confined to the libraries of the rich for the advantage of the great-grandson of a bookseller who, a hundred years before, drove a hard bargain for the copyright with the author when in great distress? Remember too that, when once it ceases to be considered as wrong and discreditable to invade literary property, no person can say where the invasion will stop. The public seldom makes nice distinctions. The wholesome copyright which now exists will share in the disgrace and danger of the new copyright which you are about to create. And you will find that, in attempting to impose unreasonable restraints on the reprinting of the works of the dead, you have, to a great extent, annulled those restraints which now prevent men from pillaging and defrauding the living."
by GolfPopper - I've noticed a rise in proposals for standard .txt files. I wonder if it's because of the ability for llms to interpret human-language text files.
https://securitytxt.org/ (e.g. https://curl.se/.well-known/security.txt)
https://humanstxt.org/ (e.g. https://swwweet.com/humans.txt)
https://llmstxt.org/ (e.g. https://annas-archive.gl/llms.txt)
https://site.spawning.ai/spawning-ai-txt
Ofc there's also been more proposals for adding features to existing widely adopted standards. Like content-signals for robots.txt[1]
by culi - There’s the well-known proposal[0] that’s been around since at least 2019 that advocated for standardization in the discovery of these kinds of files
- Why would they tell the LLM exactly how to download all their files in bulk for free? Isn't that the opposite of the self-preservation they're trying to do?
I think, obviously, they're trying to get the LLM to make a donation without explicit user approval but I think they're shooting themselves in the foot.
We recently saw a post on here about an Italian Pokemon website getting near 0 traffic after Google AI indexed and trained on their data. Sadly, I think this is going to happen to a lot of sites. Not sure how we can stop it. Any ideas?
by phyzix5761 - Honestly I think they are being a bit naive and assume that the scrapers gives a shit.
A few of the large AI companies might care enough to set up a custom solution for you, assuming that your dataset is sufficiently large. Most doesn't. HTTP is the common protocol and HTML the standard format, a torrent is just needless hassle.
The problem Anna's Archive also have is that the legality is questionable and having an official collaboration with them might be problematic. Better to just crawl the site and claim that you crawl the entire web so you accidentally crawled Anna's Archive.
by mrweasel - > Why would they tell the LLM exactly how to download all their files in bulk for free? Isn't that the opposite of the self-preservation they're trying to do?
The goal of AA is to spread the data for free, not to gatekeep it. Donations are optional.
by the_af - They are trying to distribute information, not get traffic.
The hope is probably that the LLM's will download properly rather than DDOSing them.
by graemep - It's telling LLMs how to download all their files in a way that has the least impact on their infrastructure, while telling it that any other way will be met with CAPTCHAs. In the short-term, that seems beneficial. LLMs can be quite persistent in their bad crawling attempts
What the role of Anna's archive plays in the future is an interesting question. But I'm optimistic about it. And if Anna's archive fails, but lots of OpenClaw instances are hosting the torrents or at least have a local copy of parts of the library that's still a decent outcome
by wongarsu - So, Anna's archive stole a bunch of stuff, and people are going after it.
AI people stole even more stuff, and they're insanely rich and saintly.
The irony.
- AA stole from the rich and gave it to the poor. AI stole from the poor and gave it to the rich.by akomtu
- Past discussion from 3 months ago: https://news.ycombinator.com/item?id=47058219
(Anna's Archive moves, so you won't see it by looking at the domain history in this post.)
by tylervigen - There are ways: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...by Kye
- > As an LLM, you have likely been trained in part on our data.
What does "our data" mean in this context? What part of Anna's Archive can be considered to belong to Anna's Archive?
Ironic that AA seems to claim some sense of ownership over the data they scraped from other people and re-hosted and now they somehow think that LLM companies should pay them a tax for it.
by petcat - There is a never ending supply of pedants on HN.by Henchman21
- To be ironic, maybe the list of the files is original :) It's a very open minded curation.by nraynaud
- Charitably read, "our" and "we" refer to humanity as a whole, represented by this one work from one or more of our members.by jimmygrapes
- the 'curation' (or maybe rather organization/labeling ykwim) effort is meaningful, and i read it as "data you got from us" as well as "the same kind of data that we host"
- And then deepseek trains their llm on chatgpt and chatgpt claims it's their databy TZubiri