Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why is the textual data bound up in these paper books worth the trouble?

    These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?

    For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.

  • Frontier researchers have found that dumping more and more data into training is effective at improving LLM capabilities. Nobody has a gears-level understanding of how training on some particular kind of data leads to some particular behaviors, so they generally take the attitude that more is better.
  • I have to agree; even if they are getting regular 20th century out-of-print books, is that going to add a significant percentage to their training data?

    I can only think of it being a 'low-background steel' situation where they want to locate original, non-digitized text for validation or knowledge bases.

  • I'm not an expert in LLM training, but I think we can all agree that the writing on the Internet is generally very low-quality and surface level compared to the depth of books. Most books in the past were even edited by a separate person from the writer!
  • I support physical media and doing whatever you want with said physical media. If you want to overpay for a copy of Sharepoint 2007 For Dummies and destroy it, knock yourself out. Just because media has been printed and is "rare" doesn't mean it has any practical value.

    You could easily write an article about how Goodwill and the Salvation Army dump millions of "rare" books that no one would ever buy for $0.50. And thrift stores will dump multiple copies of said books. AI companies would only ever need one!

    The only thing I disagree with is AI companies feeling entitled to freely use every piece of copywritten work without restriction, but that applies just as much to web scraping, pirated books, etc.

  • This is what the copyright laws dictate no ?
  • No. A copy is still a copy even if you destroy the original.
    by bena
  • i would assume that would come into play if they were uploading scans of the books? there must be some gray area where training like this doesnt apply to that. or ya know, just do it and face the consequences later because you already have the data and know that ai obsessed government will just shrug their shoulders.
  • Those laws are lobbied for by large corporations, these are not just laws that exist outside of that context. They can also be changed, or Amazon could just incur the fines.

    Large corporations will move fast and break things when it’s convenient; they don’t care much about the law - just about profit.

  • Not to my understanding. To begin with, it's far from given that "rare" books are all covered by copyright. But if they are, it's at best murky: whether you destroy the original doesn't really have anything to do with what you're doing with scanned contents. The scanned contents themselves may be inherently a copyright issue, regardless of destroying the original. The actual trained model has separate arguments more in its favor, so if no scanned contents exist - IE the data is read once for training and not stored or saved, they have a better argument. But in that case the destruction is totally disconnected from copyright, as they'd be totally okay to rescan the material.
  • Bias disclaimer: Amazon is my current employer, but I don't work on AI or anything else mentioned in the article.

    Yes, this is a result of copyright laws. The other commenters are wrong/uninformed.

    If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc. There are essentially zero advantages, other than it is what is required under US copyright law (or at least, it is what their highly paid lawyers believe is required under US copyright law).

  • So, if a corporation can suck up entire books to teach their machines how to think using the information from those books, can we (all humans) join a single corporation that provides all books to its employees? Just need one copy of each and we'll make that copy digitally available for our employees so they can learn from and utilize the knowledge from the books. It's not copyright infringement, they're employees.
  • Almost like… a library?
  • The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.

    So far from how I see how powerful organizations work, I'm not as certain as I would like to be.

  • > AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here)

    People say this all the time, but so far nothing has convinced me it's true.

    LLM development has more or less plateaued, and the current boundaries are very real - energy, resources, capital.

    At this point we're talking about marginal improvements against the same asymptotes of all technological innovations.

  • "rare" is used in these headlines/articles to incite and generate clicks
  • > We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak.

    Not even the title of one of those rare books?

  • It doesn't really matter. The article suggests that they are selecting and tracking books by ISBN, which means: books that have been published or at least reprinted in the last 50 years or so. And they are trying to get as much of that set as they can, regardless of quantity, quality, or any other consideration. Which means that older books printed before ISBNs became common may be relatively safe, at least from Amazon. And some books just don't come up on the used market very often.

    (small historical irony: when Amazon first started selling books, they used the Books in Print database, which included a lot of books not actually in print.)

  • As I understand it, "rare" in this context could mean anything—even a washing machine manual from the 1980s...
  • > The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.

    It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.

  • Rare as in your grandfather's John Deere manual from 1982, not rare as in a test print run of The Great Gatsby. The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.
  • > The number of books ingested by these AI companies is a drop in the bucket compared to old books destroyed every year through normal means.

    Citations? Also what exactly are these 'normal means'? As a bibliophile who loves scouring used book stores for out-of-print titles this is a topic I'm very interested in.

  • This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
  • > What difference does it make to me

    It's the scale that matters as first, and secondly, most people don't shred their books after reading them once or twice. This is just beyond words.

  • It's better value to me that someone is digitizing and training with books then they just languish somewhere in perpetuity.
  • If there’s only 3 copies in existence, and everyone on the frontier wants it in their corpus, what do you think will happen?
    by pohl
  • > What difference does it make to me what someone does with a book after they buy it?

    I think this is a bit of a myopic take. It's like saying it's my property I can do what I want, yet there exists designated historic homes or neighborhoods that are deemed to have cultural value where modifications do in fact need to be approved.

    I'm with you on the "rare" part. If people are thinking about 70+ year old documents or ancient manuscripts, I doubt that's what AI is being trained on and is being destroyed, but it's reasonable that people find _that_ idea distasteful.

    You can say it's manufactured but if these companies ignore this criticism, it's just another way AI companies are committed to losing the public.

  • What difference does it make to me if someone shoots the last bison? I wasn't getting to eat it either way.

    If people are really concerned about bison, they should lobby gamekeepers to release photos of them.

    https://en.wikipedia.org/wiki/American_bison#/media/File:Bis...

  • The Embassy of the Free Mind (https://www.embassyofthefreemind.com) is a rare book library in Amsterdam that is scanning books the old fashioned way… leading to https://SourceLibrary.org — a collection of over 5,000 books from the renaissance that have never been translated before. Consider donating, if this is a topic you care about!
  • The length of this article and the unnecessary visualizations felt like a waste after getting to the end and discovering they won’t reveal anything about the books that were scanned.

    The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”

    > As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.

    Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.

    There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.