AI companies destroy physical books – let's scan rare books before it's too late
126 points by darccio 3 hours ago | 70 comments

cladopa 12 minutes ago
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.

By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.

If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.

reply
brightball 2 minutes ago
Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
reply
ziyadb 33 minutes ago
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

reply
mbeavitt 26 minutes ago
From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.
reply
silverwind 13 minutes ago
More importantly: Once Anthropic is gone, all knowlege is lost.
reply
psma_egeliaa 11 minutes ago
It will probably be actioned off in the bankruptcy proceedings.
reply
flatline 12 minutes ago
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying

> The print original was destroyed. One replaced the other.

So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.

reply
jdiff 6 minutes ago
Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.
reply
raptor99 10 minutes ago
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.

They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.

A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.

reply
smalltorch 30 minutes ago
Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.
reply
Filligree 17 minutes ago
Obviously. Copyright infringement is settled law.
reply
SkyBelow 7 minutes ago
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.

reply
alerighi 18 minutes ago
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.

Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).

reply
HeWhoLurksLate 17 minutes ago
we also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up
reply
yehat 13 minutes ago
Nobody will ask you, they'll be taken from you, in case you missed what happens around.
reply
thisisauserid 40 seconds ago
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
reply
shrubble 8 minutes ago
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.

I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...

reply
maxdo 3 minutes ago
Ok . How trustworthy is the claim . It maid by a resource that on its own has problems with copyright.

Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge

reply
pmoriarty 21 minutes ago
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.

After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.

reply
azatom 15 minutes ago
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
reply
ZoomZoomZoom 35 minutes ago
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
reply
embedding-shape 33 minutes ago
> The main question is why aren't they leaking it to AA themselves?

Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.

reply
Filligree 23 minutes ago
Because that’s illegal.
reply
CamelCaseName 19 minutes ago
You ask "Why destroy physical books?"

I ask "Why save physical books?"

If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.

I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.

reply
tele_ski 9 minutes ago
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
reply
gruez 2 minutes ago
Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.
reply
PartiallyTyped 3 minutes ago
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.

Same can be said about many books.

reply
ainiriand 10 minutes ago
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
reply
yehat 5 minutes ago
I ask "more clean air", you answer "why, did you deserve it, having clean air is not free, there's a real cost" Somebody else asks "We need more accessible energy", you answer "why we bother with your needs, you're not efficient, energy belongs to more efficient purposes, you can live without that much energy". I can continue with more, but hope you got another viewpoint. Btw, I'm disgusted there are "humans" like you in existence. We definitely don't share the same cultural ancestry, and I hope ours will prevail at the end, rather than cold blooded, mechanical "brains" like those of your kind.
reply
gravypod 5 minutes ago
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!

reply
vasco 14 minutes ago
I agree with you for the same reason I think McDonald's is the best restaurant in the world!
reply
ryandvm 17 minutes ago
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
reply
JsonDemWitOster 5 minutes ago
The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.

Which is really why the outrage cycle over Anthropic's actions is largely misplaced.

reply
ForHackernews 21 minutes ago
Project Unica is an initiative by the University of Illinois libraries to scan and preserve publications that exist as only a single known copy: https://news.illinois.edu/u-of-i-librarys-project-unica-pres...
reply
RcouF1uZ4gsC 22 minutes ago
What often gets missed is that they are buy one physical copy and turning it into a digital copy.

They have done zero to destroy the durability. In fact, it’s probably more durable.

If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.

reply
JohnFen 18 minutes ago
> turning it into a digital copy.

Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?

reply
Filligree 16 minutes ago
They are. They can’t release it; that would be copyright infringement. But they’re absolutely planning to make further use of the book later.
reply
SkyBelow 4 minutes ago
The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.

As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.

reply
carlosjobim 13 minutes ago
I would suppose that most of the rare books in the world aren't for sale, so run no risk of being digitized and destroyed by Anthropic? They are probably forgotten somewhere on a shelf or in a box, slowly rotting. The reason these books are rare being that nobody wants them.
reply
eulgro 41 minutes ago
We've been seeing that headline for a few weeks now and I really don't understand the problem.

Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.

Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.

So what's the problem here exactly?

Also from the article:

> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.

I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.

reply
squidbeak 30 minutes ago
> A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies.

If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).

https://www.bbc.com/news/articles/cp3rprx2wl4o

> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.

> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.

> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."

reply
geye1234 16 minutes ago
Presumably, being from the 18th Century, copyright law wouldn't apply?

Detestable if they're doing it anyway to prevent competitors getting hold of it.

reply
quietsegfault 20 minutes ago
Rare and out of print does not mean important or valuable.
reply
Filligree 15 minutes ago
Usually it means the opposite. Books that are old and valuable tend to be out of copyright, so they do see new printing runs.
reply
xandrius 36 minutes ago
The problem is that you probably do little research or read very few old books.

There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.

You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.

Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.

reply
brainwad 33 minutes ago
Surely the people granted a legal monopoly to print the books will have kept a copy of the masters in order to reprint any lost works. Surely copyright works as intended for the public good and isn't just rent seeking. Surely.
reply
ForHackernews 18 minutes ago
I can't tell if you're making a joke but many (most?) rare books predate the modern copyright regime and the original printing plates are somewhere in a 17th century midden heap.
reply
josem 38 minutes ago
I think the problem is precisely that for some books there are not too many copies around like you described and if they destroy them we might eventually lose access to them directly. It might sound too extreme, but I understand the fear behind this.
reply
ab71e5 23 minutes ago
It is also legal to buy potatoes and burn them, nobody would care if I do it. But if I buy up a food supply enough to feed a country and burn it it would be wrong.
reply
quietsegfault 20 minutes ago
I’m confused - are you intending to say Anthropic and Google are buying ALL the books?
reply
FartyMcFarter 39 minutes ago
The problem is we don't know what we're losing, due to lack of transparency.
reply
anon373839 17 minutes ago
> We've been seeing that headline for a few weeks now and I really don't understand the problem.

It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.

reply
JohnFen 12 minutes ago
> I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them.

Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.

But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.

reply
m00dy 28 minutes ago
Since when books have become a supply limited asset ?
reply
simmerup 25 minutes ago
Try and read a book that’s been burnt and find out
reply
warkdarrior 49 minutes ago
> Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.

Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.

reply
throwatdem12311 48 minutes ago
At least the knowledge will be available to everyone instead of mashed together and regurgitated poorly through proprietary LLMs.
reply
torh 48 minutes ago
At least there will be a copy left for us. The AI companies won't share these books in their original form.
reply
brainwad 44 minutes ago
Because it's illegal. That's the whole reason they are shredding books in the first place, because copyright law forces them to do stupid things.

Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.

reply
Paratoner 40 minutes ago
The poow AI execs being forced to commit acts of intewwectual tewwowist when all they wanted was to cynicawwy make the wowld a wowse place
reply
JohnFen 16 minutes ago
> because copyright law forces them to do stupid things.

Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.

reply
subscribed 15 minutes ago
Copyright law didn't force them to torrent terabytes of books what they did and got caught doing so.

I'm not convinced they destroy the books to obey the law, lol.

reply
embedding-shape 34 minutes ago
> because copyright law forces them to do stupid things

This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?

Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?

Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.

reply
brainwad 29 minutes ago
It is a shit idea. But that's copyright law for you - if you want to digitise the work for yourself, you according to latest precedents have to destroy the copy you digitised ¯ \ _ ( ツ ) _ / ¯

I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.

reply
xandrius 43 minutes ago
Google probably wanted to sell the whole of Google Books.
reply
naasking 40 minutes ago
Yes, government regulations are almost always behind commercial entities making seemingly irrational choices.
reply
voidhorse 43 minutes ago
Yeah, which is completely fine. There's a major difference between:

A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:

- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.

- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.

- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?

B: Company uses freely available scanned copy of the text:

None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.

I much prefer B.

reply
Filligree 21 minutes ago
B is illegal, and Anthropic ate a billion dollar fine for trying it, so you can’t even claim they don’t want to.
reply
voidhorse 11 minutes ago
Then give up training on antiquated books. Why does an LLM aimed at providing utility for people living in 2026 need to be trained on rare (thus probably obscure) texts of yore in the first place?

Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.

reply
spwa4 7 minutes ago
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.

Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.

If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?

But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.

Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.

To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?

But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.

I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!

The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.

Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:

a) EU companies making ML models have to self-sabotage against their competition.

b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.

Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.

[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...

[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client

reply
majke 51 minutes ago
[dead]
reply
grammarisking 19 minutes ago
[dead]
reply