Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- It's a win-win-win for Amazon if they destroy the source in the process of scanning it. They get the AI training data, and it's cheaper for them to destroy the book in the process. Destroying the book ensures that it will be more difficult for competitors to scan the same content, thereby increasing the value of the data they've scanned.
Lastly, there's a misguided belief among some that it's somehow less of a copyright violation if the source is destroyed, but it's a violation either way unless the entity doing the scanning has permission from the copyright holder to make the copy.
by anonymousiam - The full headline tries to equate this to The Internet Archive’s recent legal troubles.
For a refresher: The Internet Archive scanned books then shared them online, trying to claim that converting them into digital copies qualified as a derivative work. You don’t have to be a lawyer to see how that claim doesn’t hold up to any scrutiny.
The LLM companies are not redistributing the works. They are using them for training. This blog post calls it IP theft, but it has actually been litigated in court already. Using books for training does not qualify as theft or redistribution, even though some people have different opinions about the moral angles.
One of those same lawsuits also extracted a huge settlement from Anthropic for using digital downloads from pirate sites. The conclusion was that the only acceptable way for them to use the books is to buy them and scan them. The courts forced it to be this way.
The current hand-wringing about the destruction of books is based on claims that they’re doing it to rare books that are also valuable. So far nobody has been able to actually provide an example of one of these books that is supposedly ultra-rare, but also valuable, but also only available at one of these book sellers that sells these things in bulk. We’re supposed to assume that one of these books might actually be super valuable but also super rare and also only available at these places they’re buying from.
by Aurornis - I remember reading a scene where this exact thing was going on in Vernor Vinge's Rainbow's End: there, the books were being shredded and scanned as tiny bits and stitched together in software similar to shotgun genome sequencing.
At any rate, I ought to pick it up again and finish reading it...
by rkapsoro - Well, this is what happens when they weren't allowed to use digital only copies, as they now have much more of an ability to get around copyright with physical items. The right thing would've been to allow digital items to also have the first sale doctrine, but we don't, so AI training companies must resort to destroying physical books in the process of scanning. The Cobra Effect strikes again.by satvikpendem
- This is precisely correct. The logic being applied rarely extends to second-order effects.by rpdillon
- > As mentioned elsewhere, this is also IP theft on a massive scale.
Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?
- AI psychosis cuts both ways.by noosphr
- Scanning is not the ip theft, it never was.by left-struck
- They didn't say that the Internet Archive wasn't "IP theft" - nor did they say that IP theft was inherently wrong.
"AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so.
Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying.
Or they don't care about any of that and just wanted to highlight the hypocrisy.
[0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest.
by kmeisthax - I imagine the argument could go like this.
Your granpa once wrote an obscure book about this amazing way he'd found to cure diabetes.
Corporation A buys all existing copies of the book, scans them, destroys the originals, and sets up a commercial business offering a monthly subscription to alleviate diabetes pains with this new method they claim they discovered.
Person B borrowed the book from a municipal library, Xerox'd it, and keeps a free ledger, open to all who want to read old books, as a way to safeguard free access to the world's knowledge.
Do you think there could exist any possible logic by which some people would defend person B and try to stop Corporation A?
- i dont get it.
this is a complicated way of counting how many of each word is in the book.
clearly fair use.
does the author think cutting up a book is a copyright concern? theyre buying the books, its up to them what to do with their copy. if you want the books preserved, maybe fund your libraries to get a copy or two?
by 8note - You seem to be conflating something being legal with it being either ethical or desirable.by Fomite
- It's possible for someone to do something socially reprehensible and entirely distasteful and to be perfectly within the letter of the law while doing so. The article doesn't emphasize copyright issues or fair use, because that isn't the point. The post is trying to argue that the act is not sustainable and needlessly destructive, in pursuit of something that many people believe has limited value or unproven value, which is true.
Not the case here, but it's also important to keep in mind that laws too, can be unjust and immoral. Much of humanity's progress has involved the abolishment of immoral laws.
by voidhorse - It's not necessarily doctrinally that crazy in copyright law. Some courts have considered AI training on copyrighted works a "transformative" use (traditionally more protected in fair use analysis), while providing human beings access to the works a "consumptive" use (traditionally less protected).
There's also a fair use consideration that favors noncommercial use compared to commercial use, but that's not the only question, and the statute doesn't say clearly how to combine the fair use factors.
But it's possible under the Copyright Act that some commercial uses of copyrighted works could be considered fair uses while some noncommercial uses could simultaneously not be considered fair uses.
by schoen - Everyone turning into Metallica.
The whole "copyright infringement is theft" meme seems to have stuck.
20 years ago it was all "information wants to be free" and "infringement is not theft; being deprived of speculative profit is not a loss."
How the turn tables.
by wotamess - You just discovered that principles are rare. Now let me show you my preferences of convenience.
"Information wants to be free as long as I benefit from not paying."
by ronsor - Lars has seemed to me like such a basket case this entire time.
He's lead Metallica through the "Copy our tapes! Let everyone hear our music!" early days, through the "They're stealing our master recordings!" Napster and Senate hearings era, and come all the way back 'round to "I’m just happy that fucking anybody cares about what we’re doing and shows up to see us play and still stream or buy or steal our records or whatever.” [https://consequence.net/2023/09/metallica-lars-ulrich-stream...]
2 out of 3 ain't bad, I guess.
by ssl-3 - Original full title (too long): It is a sign of the times that Amazon gets to call this fair use while huge corporations try to sue the Internet Archive out of business.
- The 404 Media report referenced in the article was submitted here: https://news.ycombinator.com/item?id=49330742
“We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility” (404media.co)
164 points | 3 days ago | 321 comments
by WalterGR - Copyright law prefers the destructive method. It's not that they're skirting the law but that the law is set up very badly in the first place.by Dylan16807
- In times of substantial technological change, laws tend to lag substantially behind what is actually needed.
And in times of substantial inequality, they lag yet further.
by jsrozner - Issue really is first sale doctrine. Meaning that after first sale there is very much leeway for the product owner. Even to scan and destroy it.
Copyright ways. Form change probably should be compensated. Even if it means that you wouldn't be able to format change your own media. Or that such activities wouldn't be allowed without compensation over certain threshold.
by Ekaros - Wasn't copyright law supposed to support creation and propagation of works of art ? Instead, apparently it is A-OK or even encouraged by the same law to destroy books.
That does not make any sense! Really might be high time to scrap it all.
by m4rtink - Courts found that having a central library of 7 million pirated books is against the law, and assessed a large penalty (small for these robber barons) so Anthropic is destroying them now to conform to the law. The law is bs and Anthropic is run by villains.by GPerson
- No, it doesn't. Copyright doesn't care about copies, it cares about copying. If it did care about copies - i.e. the total number of copies in circulation - then ReDigi and the Internet Archive's Controlled Digital Lending (CDL) program would have both been legal. Destroying a copy does not give you permission to create a replacement copy.
The only relevant case law for AI training in the US is the rulings in the Anthropic lawsuit presided over by Judge Alsup. That lawsuit ruled that it's infringement to build a shadow library from pirated books; but NOT to train AI on those pirated books. The only point where destructive book scanning even comes into play is that Anthropic also had a book scanning program alongside their piracy, Judge Alsup said that program was not infringing, and Anthropic happened to be destroying books. At no point did Alsup say that leaving the books whole would have infringed copyright - it was never even considered as it was outside the scope of the lawsuit.
Now, if Anthropic were to non-destructively scan books, store them in a library, and sell the books on, that could be infringing. All the case law about format shifting presumes the owner retains the original. So Anthropic would likely have to hold onto books, at least the ones they wanted to train on, until they were done training on that book[0]. But they do not have to destroy them permanently. They are destroying these books specifically because it is cheaper to do so than to use, say, the Internet Archive's own custom-built nondestructive scanners.
[0] I am absolutely furious about how much this sounds like "fair use is just an extra license you get when you buy a book", and I would much rather have had Judge Alsup just say AI training is not fair use instead.
by kmeisthax - I think this is impressive backwards reasoning.
Before my lifetime, copyright law all but choked and killed the public domain. And now everything is stale and the same.
This particular battle was lost with Google Books and the attempt to make the world's largest library. The modern library of Alexandria.
But then copyright lawyers got involved to get their pound of flesh. And here we are.
I am upset about the destruction of knowledge. Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago. I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them.
Everyone is impressively short sighted.
by areoform - >And now everything is stale and the same.
This has been less true every year since Usenet and Geocities. Much of the stuff on their was pirated, but that was the "stale and same" stuff you're complaining about anyway. The rest of the stuff was all kinds of original, good and bad. There's never been more total "content" and more total variety. You just have to look around for it (but that's also never been easier).
Letting AI companies try to profit off of all that creativity forever while choking the ability of the creators to make future revenue off of it is exactly what would actually lead to a "everything is stale and the same" situation. Preserving variety and novelty of new creation would look like putting in new restrictions to reduce this not-envisioned-when-the-laws-were-made sort of read-once-slop-forever usage.
by majormajor
Incredibly. Even if we throw ethics aside and look at the strategy it is still myopic. The idea is destroy so others can't get the data too. But this assumes you've done a perfect scan and you have backups. There's no good reason to destroy the books other than myopia.> Everyone is impressively short sighted.I'm more than willing to bet that in 5-10 years they'll wish they didn't destroy the books, specifically the rare ones.
To be clear, I'm not trying to condone their behavior. I hate it. But I'm trying to show that even if you had their same ethics it's still dumb. I hope workers are secretly stashing the books away. If you're one of the people in charge of destroying the books, you have a cultural duty to preserve them. I'm willing to bet people will go to great lengths to help you do it secretly so you can continue to keep them safe and keep your job. If you're at Amazon, or any company where this is happening, you have a duty to the world to make efforts to preserve the books. Steal the PDFs and lock them away. Create backups in your company. Tell your bosses you don't think this is right. Don't sit silently while this happens. Silence unfortunately is enabling. Unfortunately silence isn't a passive action
by godelski- I want copyright to get enforced against these corporations, because then they would have a reason to use their considerable lobbying power to change copyright law, hopefully in a way that benefits everyone (including things like an open digital library).by thayne
- It is also difficult to see what the legal basis would be for allowing people to read books if Amazon's use isn't fair. If I read a textbook and come up with an algorithm based on it that gets used at scale, is that supposed to be a copyright violation?
The basic point of a book is for people to (a) read it and (b) synthesise it into a comprehensive world model. And people can burn their own books if they want to, they own the thing. The only possible complaint here seems to be the scale and it seems like a big challenge to say that is a problem given that knowledge from books is allowed to be used at scale.
by roenxi