Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • They should at least give the scanned copies to a non-profit or government entity to preserve and release them when the copyright expires.
  • Forever preserving the content of a rare book by consuming one (1) copy is the kind of thing we should always be doing more of.

    Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.

  • > Roughly everyone can now benefit from its content.

    You mean every Anthropic?

  • Roughly Anthropic owners benefit from the content, unless the scans are made available somehow.
  • It's going to be hard for you to do more of that, considering that AI forms are destroying the books. You know, the whole point of the article...
  • Where can I read any of these destroyed copies from anthropic?
  • > Roughly everyone can now benefit from its content.

    You mean, roughly everyone who uses that particular model provider’s products. For truly rare books, this has an anticompetitive flavor, since it ensures others can’t train models from the same knowledge.

  • I have a deep rooted doubt that "roughly everyone can now benefit from its content" If there's any value in the content it will immediately be monetized with an aggressive pay structure and had it sat in a library, that same value would be free.
  • One (1) company destroys and (illegally?) copies a rare book. This knowledge is now of that company, not of humanity. Other companies will feel the need to do the same and before you know it knowledge is walled off in the gardens of the AI companies, and the rare books are gone and destroyed.
  • There are plenty of scanners that do not destroy the book being scanned. Here is one, I am sure there are many others: https://www.youtube.com/watch?v=b9LcTZU-HHI
  • It should be more than obvious, that the problem are not the scanners.
  • I'm not really the most AI-friendly person around but this news cycle is just as equally annoying.

    "Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.

    From https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:

    > The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

    The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.

    Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.

  • It’s true that they’re destroying old books. The claim that they’re destroying rare books seems to be repeated without much in the way of evidence.
  • When Google scanned its books, it did not destroy any of the books. This was because of the idea of the "cultural object" where the book has value as a physical artifact. The experience of reading a physical book is different than reading an electronic text. There is an argument that reading a book in its original form is better because it is closer to the original experience. This can also be said of records and tapes. The experience is designed to work in the original format. A book is a designed object with cultural significance. The container for the words can be as important as the actual words. There should be clearer distinctions and policies between mass market production where items can be destroyed and there will still be many of them left and unique and original content with limited copies available. This is not hard to do. Not to do it shows carelessness.
  • Yes, since we complained about them violating copyright so much, it does not violate copyright if they destroy the original, so they are now destroying the original. We got our demands met.
  • That's not how copyright works, at all.
  • There is good discussion on the previous thread from 6 days ago: https://news.ycombinator.com/item?id=49068738
  • It’s been stunning how often this story is reposted from various news sources across Reddit and Hacker News.

    Each time it’s some different outlet - but when you dig in, the piece is just verbal framing around the original story written by 404 media:

    https://archive.is/9MQrK

    What’s even more stunning is that the original article doesn’t provide evidence that “rare” books are being destroyed. That doesn’t even appear in the original article title.

  • On the surface it seems bad, although many of these books were probably rotting in place rather than being read. If the companies are willing to make the digital version available this may actually prreserve the books.

    More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.

    This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.

  • What would happen if we created large volumes of books with AI but faked their dates before 2022 to induce model collapse?
  • As long as economics incentivizes next quarter thinking nothing will be done. We can't even handle human portion of global warming which is arguably more important than just some AI model collapse scenario.
  • Entropy is a thing.

    Ancient Rome stopped creating aqueducts because they had all the ones they needed. They failed to pass that knowledge on to the next generation and so they just forgot how to create aqueducts.

    Once AI starts making the majority of content, people will simply forget how to make content. Before we know it we're all fat slobs in floating chairs like in Wall-E.

    Best case scenario I think is similar to what we see in Ian M. Banks Culture books where the machines basically take care of us out of the goodness of their hearts and we just kinda fuck off into obscurity.

  • Yeah, I really don't know where this goes from here. Like certain sources, particular Reddit and Twitter, have to largely be considered spoiled at this point. It's hard to quanitfy the amount of bots but my intuition tells me it's really high.

    But changing how non-AI people write, that's an interesting angle. Because where do we go from here? In 100 years will we still be overvaluing pre-AI sources? That doesn't make sense. Of course a lot can (and will) change in 100 years.

    But i'm reminded of pre-atomic steel, which is steel made before the first atomic bombs were detonated and thus have really low background radiation. This is necessary for making MRIs and such. People will go and find it from shipwrecks and such (ironically, many of which are from WW2). It's also a finite resource. What happens when we run out?

    Now pre-2022 texts aren't consumed (other than destructively scanning books of course) but it is also finite. We can't make more of it.

  • Humanity hasn't caught up to the reality of the times; we should adjust copyright laws so anything over say 10 years old can be hosted for free on the Internet. We should have colossal databases of music, books, movies, etc. that can be freely traded and downloaded without breaking any laws. Maybe in the past the commerical protection of art and IP was necessary but with the Internet and AI I think we need to, as a species, get over it and focus on information archival and dissemination over protecting income streams.
  • I imagine you're in favor of the artists being compensated in some other way then? Or should they just suck it up?
  • 10 years may be too little but generally I agree.

    What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle down.