Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."

    Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?

  • No but with 100% clean data you can easily train a model to filter ai generated content.
    by npn
  • Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content
  • > Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time?

    No. They also use lots of other methods to get training data.

    by eru
  • As a rule: all high quality text is useful.

    There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.

    At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.

    Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.

  • Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.

    So the solution is:

    - mandatory copyright registration and renewal with links to where the work can be acquired

    - a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.

    In this way, a hundred digital preservation societies can bloom.

  • > copyright

    This is what caused the problem in the first place. If people had unrestricted access to content then the world would be a better place. And works wouldn't be so rare that it's worth AI companies buying and destroying them to gain some edge, as well as remain in legal compliance.

  • You ask "Why destroy physical books?"

    I ask "Why save physical books?"

    If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

    Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.

    I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.

  • > If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

    The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.

    Same can be said about many books.

  • > If they are truly rare, then they are likely not valuable

    "likely" being the keyword here, what about heavily censored books?

  • What you personally find important is not what everyone else finds important or inspiring. Destroying something takes it away from every single future human being.
  • That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
  • Good point, why waste time deciphering old badly burnt scrolls when anything worthwhile should have been preserved.
  • You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).
  • So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
  • >If they are truly rare, then they are likely not valuable,

    Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.

  • Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.

    I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.

  • I think is the scanning for profit and destroying them in the process that angries people.

    If they we’re just kept where they were you can always say they’ll eventually be scanned or have that potential.

    I do believe the issue is a bit overblown but the core of it sounds reasonable to me.

    by guax
  • I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.

    That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.

    by akk0
  • The dissonance when a cause you support is loudly represented by disingenuous types.

    At some point they’ll hit on data center water consumption as yet another reason to support the cause.

  • "I imagine they are only buying one copy of each book…"

    I doubt that corporations of this scale do that extra kind of… book-keeping. They more than likely buy books by the pound.

  • Doubtful. A robust protocol would be to scan several of each edition (to ensure no scanning errors), and scan each edition. Then too, these books are being purchased in lots with accidental duplicates, and all the major labs are doing it. So we're likely talking about tens of each book. For rare books - anything over a couple of hundred years old or small print runs either, that could well be most or even all copies. This wouldn't be immediately noticed either, especially if the books aren't currently considered noteworthy or well known.
  • >I imagine they are only buying one copy of each book

    The previously struggling second hand bookseller in my town has upgraded their car from a 15 year old hatchback Renault to a brand new Range Rover. Some Canadian company has been buying any book he can provide them for the last year. Their quotes aren't by number of books, or even weight, but by volume. As in, they pay him by the shipping container, and he sends several of those a month.

    I think it's reasonable to assume the books in this supply chain which aren't destroyed in digitisation are just pulped and sold to Procter & Gamble for toilet paper manufacturing. I can't see any other fate for Anthropic's second and third copies of The twelfth edition of Vera Lynn's 1980's memoir "We'll Meet Again".

  • Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.

    Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.

    Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.

    Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.

    Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.

    Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?

  • > establishing goodwill with the public

    Not sure even rare books will dig these big AI companies out of the hole they're digging for themselves.

  • Anthropic wasn't explicitly told to destroy the physical copies, but it weighed heavily in their favor.

    "Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

    "For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.

    The third fair use factor favors fair use for the purchased library copies converted from print to digital."

    Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal. 2025). https://docs.justia.com/cases/federal/district-courts/califo...

  • From what I understand, to work with copyrighted books they need to essentially format shift (i.e., scan and destroy the physical book). So a book vault would not solve this issue.

    A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's not worth it.

    At this point my only hope is that in the long run these scans make it to the public somehow (leaks, copyright changes/expiration, whatever), where they can then be accessed and preserved by everybody. Then we could have our true digital library of Alexandria.

  • Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
  • I have been reading more about exactly what is happening. There is an indication that the scanning and destruction of books can destroy the last copy of books that are rare. This way the information only exists in digital form inside the large language model. This claim is increasingly appearing in news articles. It is an interesting observation that needs additional confirmation. This can alter the permanent record in unexpected ways. https://dallasexpress.com/national/the-vanishing-page-ai-fir...
  • Given all the bad PR around this issue - you'd think they'd do the 10 seconds of examination "is this rare" before putting it in the cutter.

    I'm genuinely surprised they don't.