Top
Best
New

Posted by darccio 6 hours ago

AI companies destroy physical books – let's scan rare books before it's too late(annas-archive.pk)
612 points | 373 comments
thread_id 2 hours ago|
I don't see any mention of Project Ocean - AKA Google books. Before AI they endevoured to digitize books in a massive online library. This inccluded rare and out of print books many which are archived at libraries. Because they had to preserve the books and return them in the condition they received then they created elaborate technology to accomplish this. The project was met with significant legal challenges from authors and publishers which was eventually overcome. The legal precedents that were established from Project Ocean laid the ground work for the process as it exists today. Books from libraries are still being preserved.

https://en.wikipedia.org/wiki/Google_Books

https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

probably_wrong 2 hours ago||
Whenever I need something from Google Books I inevitably reach the message that this is a limited preview and the part I need is not included.

I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.

ndiddy 1 hour ago|||
At one point Google Books was supposed to act as a clearinghouse for scans of out-of-print books. You could have purchased a scan of any book on the site for a reasonable price, and libraries could subscribe to a service where the full text of all books was available. This settlement then got shot down because some research libraries and authors argued that this was anti-competitive, and instead wanted Congress to pass a law to free up the rights to orphaned books so anyone could start a competing service. No progress on this was subsequently made because nobody in Congress cares enough about the rights to out-of-print books to get legislation passed. The whole reason why they're out of print when ebooks and print-on-demand exist is that they won't get enough sales to make it worth the time and money to figure out who the royalties should go to. It's not a flashy issue that would make a ton of people vote for you to get re-elected, and it won't create a ton of new jobs. The result is that now nobody outside of Google gets to see the full Google Books scans.
allturtles 50 minutes ago||
Yes, this was a great tragedy. I was very sad to see academics at the time arguing against Google providing what would have been one of the greatest storehouses of readily available knowledge in the world, in favor of an imaginary alternative that didn't exist and never would.
p0w3n3d 1 hour ago|||
Quod licet Iovi, non licet bovi

Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

misnome 1 hour ago||
> Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)

No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.

allturtles 55 minutes ago|||
There is so much misinformation/confusion about this... they go sued after lending "unlimited" copies, but they were sued (and lost) for lending exclusive copies (controlled digital lending):

> “At bottom, [the Internet Archive’s] fair use defense rests on the notion that lawfully acquiring a copyrighted print book entitles the recipient to make an unauthorized copy and distribute it in place of the print book, so long as it does not simultaneously lend the print book,” Judge John G. Koeltl of the U.S. District Court in Manhattan wrote. “But no case or legal principle supports that notion. Every authority points the other direction.” [0]

[0]: https://www.insidehighered.com/news/tech-innovation/teaching...

ndiddy 56 minutes ago|||
That's why they got sued, but the suit is mainly over whether controlled digital lending is legal at all rather than their "emergency library". Archive.org lost the case on summary judgment, meaning that they could not come up with a single fair use argument for CDL that the judge found compelling enough to let the case go to trial. The full judgment is here https://storage.courtlistener.com/recap/gov.uscourts.nysd.53... but here's a couple excerpts:

> The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers.

> In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."

titzer 2 hours ago|||
Like all things, Google will eventually realize they cannot make significant ad revenue and they will eventually give up and discontinue serving this, though I doubt it's more than a scratch in terms of disk space.

It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

dblohm7 1 minute ago|||
> It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.

They never should have been trusted in the first place.

ktm5j 8 minutes ago||||
I think you're missing the point. Maybe I'm wrong, but I'm pretty sure they're trying to point out that this book scanning doesn't need to be destructive. I'm not sure if AI companies are using a scanning method that damages the book or not, but they destroy the books after scanning to avoid copyright issues (ie they aren't duplicating the books). This Google project seems to demonstrate that this isn't actually necessary, and that the AI companies are just doing it out of laziness.
greyw 2 hours ago||||
It's data for their AI pipeline. Basically digital gold.
nazgulsenpai 2 hours ago|||
They are also an AI company now. Why would they stop?
JKCalhoun 2 hours ago|||
Also, why would they ever share their collection?

Book scans, secreted away, are worthless to the public.

al_borland 2 hours ago|||
They could use them as training data, without providing access the actual books.
jacekm 1 hour ago|||
> The project was met with significant legal challenges from authors and publishers which was eventually overcome

I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.

chungusamongus 43 minutes ago||
IIRC the courts ruled that because it would be implausibe for a person to use google books previews to read an entire work (you'd have to make a whole bunch of separate accounts to do so), it could not plausibly impact the market for that work.
wrathofquan 24 minutes ago|||
I've worked in academic libraries since 2010. I've always felt like it was a mistake for our digital library leadership to put so much trust in Google Books despite their promises to maintain the integrity of libraries (this mostly happened by the way).

At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.

toomuchtodo 1 hour ago||
Internet Archive version: https://openlibrary.org/

Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...

Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...

Web app: https://archive.org/want/?mode=donation_book

For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.

If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.

shagie 55 minutes ago||
As an aside... https://www.google.com/books/edition/The_Systems_Bible/mrOsb... (and I haven't hit any "you can't read this" limits yet).

While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle. https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-...

toomuchtodo 49 minutes ago||
> While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle.

I do not recommend Kindle books, as you don't own them and Amazon is closing any DRM loopholes, but I understand that if you must have access to a resource, they are a solution. I'll noodle on wiring up a Kindle so it can step through pages programmatically controlled for CCD capture and OCR using vision LLMs.

Kindle update toughens DRM and ends Libby loophole - https://news.ycombinator.com/item?id=49388279 - August 2026

NishanStepak 13 minutes ago||
Nondestructive scanning can cost 10x as much. This is about cost. It is not about preservation. Google never destroyed the books it scanned. Amazon and Anthropic are attempting to save money. They are not considering whether or not a book is rare. They are treating books as a commodity. Rare books are rare. It is easy enough to identify when there are a limited number of copies of a book. The issue is saving money on items which cannot be easily acquired. There are plenty of books where there are thousands of copies available. Destructive scanning of these books is not the issue. It is indiscriminate destruction of items that are unique and in limited supply. Rare books are more than their content, they are the typography, materials, design, smell, and physicality of the items which matter. They are often very different than mass market hard covers or paperbacks. Not every book initially was produced in massive quantities. This is incorrect. Many books before they became important were done in limited runs. The lists from what I am reading often include books which are limited in quantity. It seems to be an attempt to get everything possible, not just the massively produced items. The problem is making AI companies separate the truly rare and unique items from the commodity mass produced items. Nondestructively scan the rare ones, cut up the ones where there are thousands of copies.
cladopa 4 hours ago||
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.

By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.

If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.

TFNA 21 minutes ago||
> any important book has been duplicated by thousands, tens of thousands or even million of units.

Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.

The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.

nloomans 3 hours ago|||
> a digital copy with the ability of doing millions of copies is stored somewhere

somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”

the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.

> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.

archivists keep everything, because we don't know right now what will be important 100 years from now.

radu_floricica 3 hours ago|||
> somewhere were we can't access it

By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.

Plus having the info part of a LLM makes it immediately available to literally billions.

I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.

winterismute 14 minutes ago|||
> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.

I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).

jjulius 3 hours ago||||
>Plus having the info part of a LLM makes it immediately available to literally billions.

Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.

If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".

And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.

SkyBelow 2 hours ago||
Depends upon what you want.

For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.

The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.

But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.

That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.

chefandy 4 minutes ago||
> But, that little bit of data is a bit more data than existed before,

No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.

> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.

That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.

I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.

darkwater 3 hours ago|||
> Is it "locked"? Yes, by copyright laws,

> Plus having the info part of a LLM makes it immediately available to literally billions.

Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".

gnfargbl 3 hours ago||||
> permanently locking human knowledge inside private corporate servers

History tells us that very few "permanent" situations are truly permanent.

Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.

darkwater 3 hours ago|||
> History tells us that very few "permanent" situations are truly permanent.

If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.

mohamedkoubaa 19 minutes ago|||
Exactly. Calling anything on an SSD permanent is criminal.
mvlipwig 34 minutes ago|||
I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.
pibaker 11 minutes ago|||
> any important book has been duplicated by thousands, tens of thousands or even million of units.

It is common for academic books to have publication runs in the low three digits.

You may argue these books are not important. But how do we know if we fail to preserve it?

afpx 3 hours ago|||
I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
mannyv 2 hours ago|||
I have books, but they are just objects. They're nice objects, but just objects.

Fetishizing books isn't going to help.

In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.

brightball 3 hours ago|||
Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
sajithdilshan 3 hours ago|||
Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
shiandow 3 hours ago|||
Somehow I don't think they're looking for the books that have been copied over and over.
GreenLightGo 37 minutes ago|||
Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...
jll29 3 hours ago|||
Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
pshirshov 3 hours ago|||
Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
pfdietz 1 hour ago|||
Here we have another entry in the long list of "things described on the Internet that never happened".
dataflow 3 hours ago||||
Where did you see they tend to buy all the copies? This comment is the first time I've heard of this.
wmeredith 3 hours ago|||
I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?
dataflow 2 hours ago||
I could see buying multiple copies being useful to mitigate problems, like damage.

But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?

p-e-w 3 hours ago|||
It’s just another lie of the type these threads tend to be filled with nowadays.

Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.

I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?

diseasedyak 2 hours ago||
It really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.
wongarsu 2 hours ago||
I wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are
mistercow 3 hours ago||||
Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them.

But where did you hear that they’re buying “all copies”? And to what end?

dbspin 3 hours ago||
This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.

To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.

Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.

famouswaffles 2 hours ago||
>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.

I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.

dbspin 2 hours ago||
OK... I'm going to assume good faith even though your wording makes it somewhat unlikely.

Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.

Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.

Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.

famouswaffles 51 minutes ago|||
I think there are two issues here:

1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.

Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.

2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.

Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.

Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.

sajithdilshan 3 hours ago|||
what is your source?
rvz 3 hours ago||
First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.

Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.

Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?

[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...

kccqzy 2 hours ago|||
Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.
ziyadb 4 hours ago||
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.

From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.

hypendev 3 hours ago||
>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

How tho?

Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.

Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.

And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.

Mudbugs 3 hours ago|||
Yep, this feels pretty much it. Looking at the "rare books", it was books that nobody would care about or would just rot away anyway. (The example book of Old books of agriculture is probably not that important today)

Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.

The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.

WarmWash 3 hours ago|||
I think one of the major things you learn as you get older is that there is a huge abundance of people who say the right thing, and a much smaller group of people who do the right thing.

The internet made this even worse by celebrating people who only have to say the right thing.

p-e-w 2 hours ago||
This would be more convincing if there was any agreement on what the right thing supposedly is.

The major lesson I myself learned as I got older is that “you should do the right thing” is a rephrasing of “you should do what I want you to do”.

andsoitis 2 hours ago|||
> This would be more convincing if there was any agreement on what the right thing supposedly is.

The internet has amplified voices who say things to signal rather than do things to change.

That's orthogonal to knowing what "the right thing" is.

Lerc 1 hour ago||
The increase in people who believe that perception is reality has the logical consequences of people shouting their opinions loudly enough so that it becomes 'the right thing'
nmz 1 hour ago||||
Knowing and enacting are two distinct things, "you should do what I want you to do" means you're already doing the wrong thing, and yes, there is an agreement on what the right thing is, its what is ethical, that is, what non profits and archivists are forced to do.
bix6 1 hour ago|||
That’s just not true. There is a clearly a right thing to do in many circumstances.
areoform 1 hour ago||||

    > The example book of Old books of agriculture is probably not that important today
If I may be flippant, not to you but to the sentiment, skill issue.

We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.

And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era?

What if... you just want to learn what your ancestor was doing at such and such time?

What if you want to find clever techniques for robot arms to work with food crops in space?

Or, heck just the alpha from a hedge fund point of view of finding old climate patterns and... :)

Your ability to make the most of knowledge is only limited by your imagination.

    > Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
I don't understand what you're trying to say here.

    > The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
https://en.wikipedia.org/wiki/No-kill_shelter

re: saving books, at a personal level, I try to use the excuse of work to find, read, and do stuff with old books,

https://1517.substack.com/p/powder-and-stone-or-why-medieval

And yes, people still care. And people who care do things.

demibabs 36 minutes ago||||
Also, why are people acting like Anthropic digitizing and destroying a rare book is what makes it inaccessible to the public?

If they were simply buying the books and doing nothing with them, the public wouldn’t be able to access their copies anyway.

smallerize 2 hours ago||||
Well if you're appealing to librarians:

"Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on.

Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...."

https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...

Mudbugs 2 hours ago|||
I don't get the point of this one. We not talking about "retiring books", we talking about "rare but still fine books".

I don't think many Liberians like the idea that they have to do it; it is just one of those things that has to be done, sadly.

tingletech 49 minutes ago|||
I don't know what Liberia has to do with it.

Weeding (deselection works) is a fundamental part of collections management. Every trained librarian is going to understand this.

The interlibrary loan system has mechanisms in place to make sure the member libraries keep two copies of each work in each region. Collection managers consult these databases during weeding to make sure they don't deaccession the last copy.

If the monograph was never collected by a library and it gets caught up in a destructive scanning project then I guess it was pretty "rare" in a literal sense. "Rare Book" in library land is sort of a term of art and I'm not sure if the books in these destructive scanning projects meet the criteria.

smallerize 2 hours ago|||
I don't think that changes the message at all. If a librarian doesn't like routine book destruction, they would still feel worse about this different thing that is happening.
1234letshaveatw 53 minutes ago|||
I'm not a bsky person so I didn't click though but library books aren't "retired" to a farm upstate. My SO works at a library, they are mostly shredded. This is a nothingburger
pwdisswordfishq 2 hours ago||||
They should be happy to have a job at all in the Republic of Liberia.
n4r9 2 hours ago|||
A quick search suggests that most "weeded" books are sold on or donated rather than burnt or sent to landfill. Do you have sources that say otherwise?
Mudbugs 2 hours ago|||
Just go to a local library and ask a librarian or anyone who has a lot of books and tries to give them away; sadly, in most cases, they pick the valuable ones, and the rest just get sent for destruction (Burning).

It is pretty standard procedure.

alightsoul 43 minutes ago||
The difference everyone is missing, is that now there is commercial incentive to burn books aka destroy them after scanning. It's now a profitable business to do so, not something that only has to be done to clear up space
1234letshaveatw 50 minutes ago|||
Ours has a volunteer run used bookstore that sells a few for pennies on the dollar. Vast majority are shredded
dumb_notsmart 2 hours ago||||
Oho you think they're buying common books? Books they can find online already?
starkd 3 hours ago||||
Good to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
sergimansilla 2 hours ago|||
No, they are destroying them because that’s the way that they can keep a digital copy legally.
beering 2 hours ago||
This is not true. Google Books does not destroy books. No court has ruled that you must destroy books to legally keep a digital copy.
tingletech 42 minutes ago||
I think under first sale doctrine you have a much stronger case with destructive scanning. Google Books, HathiTrust, and Internet Archive's book scanning project have had a lot of legal expenses.
voidhorse 2 hours ago||||
You don't know that. A major part of the problem here is the lack of transparency.

Keep in mind we wouldn't even know this was happening were it not for investigative journalism.

andsoitis 2 hours ago||
> we wouldn't even know this was happening were it not for investigative journalism.

Pretty feeble investigative journalism if they cannot give us the names of even 5 such books we are supposed to be outraged about.

shagie 2 hours ago||
While not an authoritative source... https://old.reddit.com/r/books/comments/1vugion/the_federal_... (and not actual book titles)

> Correct. I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades. It would have eventually found its way to the discount rack and then the dumpster. At least this way we're getting some money in that we can use to buy actual cool books/collections, pay salaries and bills, etc.

So... something like Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 - https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...

I make no claims that that's a such a book that's been ordered, but that's the type of book that the reddit post references.

alightsoul 40 minutes ago||
That data is useful as a source of scientific knowledge even if it's not current. Although it's probably already online, they probably don't want to download it and risk getting another copyright lawsuit
fortran77 3 hours ago|||
They are destroying them because it’s easier to scan them if you slice the binding.
eloisant 2 hours ago|||
I didn't fully understand why but apparently there is also a legal reason to destroy the books, it makes it them less likely to be considered copyright infringement.
beering 2 hours ago||
Not really. Selling the book onward does seem legally dubious but legally nothing (yet) prevents you from storing the book in a warehouse. obviously it’s cheaper to dispose of them.
cdkmoose 2 hours ago|||
And very hard to re-assemble after you have done that. They are not in the book binding business.
voidhorse 2 hours ago||||
> This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.

But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.

Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.

If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.

joshstrange 2 hours ago||
Are you under the impression they scan the book, train on it, then destroy the digital copy? Because that's not what's happening. They scan it, and hold it forever to train future models on. The scan still exists, not available to the general public but that's no different than if they had bought the books and kept them in a private library closed to the public.
anjel 1 hour ago||
Destroying the physical copy inflates the value of the retained digital copy
chrisjj 2 hours ago||||
> This way, it will live forever instead, scanned

But unreadable by humans, right?

mjdv 2 hours ago||
They're scans. Those are human-readable. They probably won't make them available to the public, which is the exact same state they would be in if they bought the books and just put them on a bookshelf in a warehouse.
alightsoul 42 minutes ago||
That was due to negligence. Now it is profitable to keep the books private, and actively deny others access to them.
RIMR 2 hours ago||||
Those books at the flea market aren't "rotting", they are available for sale.

You seem perfectly fine living in a world where your flea market is devoid of books.

otherme123 2 hours ago|||
A lot of them are rotting. We are not talking "one of the three living copies of the first edition of Joyce's Ulysses". Rather "1956 statistics of the cultive of yuca in 'some small village from Mexico': a boring analysis". Those books have value to train LLMs as they are 100% free of AI text, but has been collecting dust (or rotting) in someone's room for decades, and no human is buying them even for 10 cents.

Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.

hypendev 1 hour ago|||
No, they are not. A lot of them are rotting and get thrown or given at the end of the day.
ro_sharp 3 hours ago|||
You’re equating something being ubiquitous and affordable with it not being valuable.

Yes, there are plenty of books, many were printed, many have lasted a very long time (plenty over 100 years!).

That says more about the success and utility of the technology than it does about whether individual books should be shredded.

Aurornis 3 hours ago|||
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria,

These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?

Your local library throws out books every year and nobody thought twice about it.

mindcandy 2 hours ago|||
They are ridiculous because the reality of the situation is boring. Boring doesn’t drive engagement. So, all the clickbait headlines and ragebait comments imply scandal. If they didn’t, they wouldn’t get attention.

After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.

This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.

alightsoul 37 minutes ago||||
Does china developing ai powered missiles sound ridiculous? It sounds ridiculous to me because the US is surely doing the same, It would be boring if we knew the us was doing the same. It is emotional to think about china destroying the US when it is just the antithesis of American exceptionalism. It reminds me of what Dario said about their stance on open source.
snickerbockers 3 hours ago|||
>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy?

Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?

>Your local library throws out books every year and nobody thought twice about it.

When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.

BeetleB 2 hours ago|||
No. Almost none of the books you donate to the library get to the shelves. If they can sell it, they will. Otherwise they are thrown away.

I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".

Goronmon 3 hours ago||||
As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices.

What happens when the books don't sell?

pfdietz 1 hour ago|||
Our local library-connected biannual book sale puts a large dumpster, about the size of an 18-wheel truck trailer, outside the warehouse where the sale is conducted. At the end of the sale it gets filled to the brim with waste inked cellulose, and sits there open in the weather until it's taken away for disposal.
p-e-w 2 hours ago|||
They are thrown away, which has been happening since forever, and nobody ever gave a fuck until AI got involved.
snickerbockers 17 minutes ago||
I think it hits different when you're buying large quantities of used books with an eye towards ones which aren't readily available online, and with the specific intention of destroying them.

If all these companies were doing was burning star wars tie-in novels and harry potter sequels nobody would care. That's not their goal because they already have those in their training set. The whole point here is to find rare or underappreciated books from the pre-digital era which nobody ever made publicly available in a digital format.

BTW destroying them isn't even necessary for scanning. It's the easiest way because removing the binding and turning it into a flat stack of papers solves many problems but there are actually dedicated book scanners designed to hold open the book while its photographed, and un-curling pages in post-processing was already a solved problem long before people were using AI to correct images.

johannes1234321 2 hours ago||||
> When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.

Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.

Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.

infecto 2 hours ago||||
As someone who has gone to many many used book sales over decades… many of the books at a sale never get sold, guess what they usually get tossed in the dump. This includes your local library book sales. Books are heavy and worthless. Cheaper to throw away the ones that nobody picks up in a sale.

I know it comes at a shock but truly most books are absolutely worthless.

jonhohle 3 hours ago|||
The average lifespan of a library book is 26 loans. For a popular book that could be less than a year.
brookst 3 hours ago|||
No, that’s the hyperbolic reaction clickbait wants from you.

Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.

Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.

Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.

dd8601fn 2 hours ago|||
I don’t think any of these people know.

All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.

Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.

And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.

andsoitis 2 hours ago|||
> Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.

Then they should list the names of these books otherwise I say they're alarmist.

dd8601fn 1 hour ago||
If they have, I haven’t seen it.

It’s very possible I’ve just missed deeper reporting, obviously.

But otherwise I agree. I’m neither losing sleep over it or just trusting that these (historically kinda scummy) businesses are actually behaving.

If there’s a serious problem I’d like to see something more concrete. Same for hand-waving the question.

pessimizer 2 hours ago|||
I heard from the first stories that virtually all of these books being ordered have ISBN numbers. Books that are rare that have ISBN numbers are rare because no one wanted them 99.9% of the time. Somebody wants every book, but you'd spend many, many years finding that somebody.
VanTheBrand 2 hours ago|||
If they have no value why are they being acquired and scanned?
joshstrange 2 hours ago|||
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.

How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).

> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.

pmarreck 2 hours ago|||
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Hyperbole much?

Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!

Leynos 3 hours ago|||
If you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
leonidasrup 3 hours ago|||
A small change in the copyright law would fix this problem. Something like:

If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.

infecto 2 hours ago|||
Puts the burden on government to store what is probably 90% worthless material.

Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.

There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.

pessimizer 2 hours ago||
> Puts the burden on government to store what is probably 90% worthless material.

That's what governments are for.

infecto 1 hour ago||
Since we are talking about a US perspective do you have evidence that backs this up? It just comes across as an empty statement. The government is the will of the people and I personally like the idea of fixing copyright instead of making the government store how to use windows 95 books.

Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.

brookst 3 hours ago||||
Wait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans?

How does this help anything, except create more work to throw in the trash?

ro_sharp 2 hours ago|||
This is already required for new books published in the US, and has been for more than one hundred years.

It’s called “mandatory deposit”

jjkaczor 2 hours ago||||
Actually - most jurisdictions that issue a publisher a unique root ISBN number have a stipulation that anything new published using that number must have a copy sent to them.

Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.

Semkas 2 hours ago|||
https://www.google.com/search?q=how+many+books+are+published... https://www.google.com/search?q=how+many+megabytes+average+b... https://www.google.com/search?q=what+is+3+megabytes+times+2....
ndriscoll 2 hours ago||||
Wrong end of the pipeline; we should instead demand digital copies of media be sent to the Library of Congress in order to obtain copyright, along with a registration fee to pay for indefinite storage and other costs. Registration should be mandatory if you want copyright. For things like books where a machine readable text format existed, it should be mandatory to include (so no requiring OCR). Access to the archive should be available for research use (including ML training) at cost.
brainwad 2 hours ago||||
The Library of Congress already has a copy of every book published in the US. How would this help?
toast0 1 hour ago||
If the Library of Congress has a digital copy, it would be easier for them to distribute the work after the copyright of the work expires. That would be a public benefit.
bell-cot 1 hour ago||||
IANAL, but I recall copyright law being far too complex for any easy Protected/Not Protected test to exist.
starkd 3 hours ago||||
So now the Library of Congress has to manage all these submissions whenever someone scans something? How do you even go about enforcing such a thing?
Upvoter33 3 hours ago|||
I love this idea. At least make them turn it into some form of a public good.
smalltorch 4 hours ago|||
Surely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.
JKCalhoun 2 hours ago|||
I'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)
pessimizer 2 hours ago||
If they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.
JKCalhoun 1 hour ago||
Thanks.

There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again.

I'll look into libgen.

Filligree 4 hours ago|||
Obviously. Copyright infringement is settled law.
mbeavitt 4 hours ago|||
From their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.
flatline 4 hours ago|||
They cannot scan the books then resell them or donate them under current US copyright law. It’s not clear to me that they could warehouse them if they wanted. In the recent Bartz v Anthropic case the judge ruled this destruction as legal, saying

> The print original was destroyed. One replaced the other.

So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.

jdiff 4 hours ago||
Nowhere in this description did it require destruction of the physical book. This is being done because it's easier to scan a shucked book, and this explanation is circulating because it's easier to blame it on the law and that pesky meddling government.
fc417fc802 3 hours ago|||
So if I scan a book, sell it, and keep using the scan, is that legal? (Spoiler: That's not legal. It's a violation of IP law.)
toast0 1 hour ago|||
Probably, if you sell it after the copyright expires.
jdiff 3 hours ago|||
Selling it is not allowed. The inability to sell it does not require its destruction.
inigyou 3 hours ago|||
But copyright law does, because otherwise you have two copies.

This isn't theoretical, AI companies have finished lawsuits about this and this was the ruling.

jdiff 1 hour ago||
The ruling was that what they did was within the law, not required by law in every detail. They cannot resell the copies. I saw nothing in that ruling that required their destruction, because it is not required.

A person can digitize their own books without destroying the original. So can Anthropic. They are choosing to destroy the books for easier scanning and trying to palm off the blame for it.

flatline 2 hours ago|||
Why would a company keep the hard-copy around at the risk of it being inadvertently given away, resold, etc.? It's a huge outstanding liability given that the illegal copying of works -- the other part of that case -- is what they settled out of court for some huge amount of money. Destruction is the only thing that makes sense.

I'm old enough to have been around when DCMA legislation was under discussion. Many people were dead-set against it and raised concerns over matters exactly like this. In Rainbows End (2006), Vernor Vinge wrote about a similar scenario where a robot went through the university library shredding books, and scanned the shredded pieces to recombined them into a digital archive.

Anthropic may be doing shady things and may have even done this on their own recognizance, we just don't know. As it stand, this is 100% a consequence of US copyright law, much of which was written by large corporations to protect their own assets.

jdiff 1 hour ago||
I agree fully that it makes logistical sense. But it is not a legal requirement, and they should not be permitted to use that as an excuse to wash their hands of their own decisions.
flatline 1 hour ago||
I think I agree with you in spirit. I don’t like what these companies are doing, and Anthropic’s actions can for the most part stand on their own. Copyright law is just a special interest of mine, and I do think it’s important to recognize what external incentives exist and what they prioritize. Because other companies will act in similar manners under the same incentive structure, and the problem is going to cascade and magnify if it hasn’t already. There are active court rulings setting precedence for this behavior - take note!
brookst 3 hours ago||||
Citation please? Bartz v Anthropic seems pretty clear, see also Authors Guild v Google and RIAA v Diamond Multimedia.
jonhohle 3 hours ago|||
1 point by jonhohle 0 minutes ago | edit | delete [–]

You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.

JKCalhoun 2 hours ago||||
I love how sci-fi authors like Ray Bradbury toyed around with a similar issue but then got it so wildly wrong.
silverwind 4 hours ago|||
More importantly: Once Anthropic is gone, all knowlege is lost.
psma_egeliaa 4 hours ago||
It will probably be actioned off in the bankruptcy proceedings.
jonhohle 3 hours ago||
That’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.
wildzzz 2 hours ago||
They cut the bindings off and feed loose leaf books through a document scanner. It would be difficult to store and probably unsellable, it's probably going in the trash.

To legally retain these scans, you must own the original book. You can sell or give away the original book (sans binding) but it's legally dubious as to whether the scan can be transferred along with it. So if Anthropic has no interest in storing thousands of loose leaf books, they are likely destroying both the original and scan as soon as possible.

At the end of the day, the only thing of value Anthropic has is the trained model which is definitely transferable.

thesdev 3 hours ago|||
> working towards the benefit of humanity is not an exclusive right / domain of theirs

That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.

bmelton 3 hours ago||
If they were just chopping them up without scanning them first, then sure, but I think that scanning books and burning books are polar opposites
logseman 3 hours ago||
Burning books and destroying them in a way that nobody else can access the content anymore is a distinction without a difference.
SkyBelow 4 hours ago|||
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.

JKCalhoun 2 hours ago||
Copyright law eating itself…
convolvatron 1 hour ago||
no matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.
LaGrange 2 hours ago|||
> I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity

Look, others talked about how this fetishising of paper books is quite silly (though I don't like it when the destructive scanning is just so one could feed it into a chatbot) but I have to say, all I can do after reading the above sentence is laughing bitterly. Anthropic is an American for-profit company, any talk about "working towards the benefit of humanity" is just marketing lies, and it's always incredible to see people treat those seriously.

_Of course_ Anthropic does that.

JKCalhoun 2 hours ago||
It's probably good to remind ourselves though of how disgusting they can be.
raptor99 4 hours ago|||
I hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic.

They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.

A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.

brookst 3 hours ago||
Are you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
JKCalhoun 2 hours ago||
It's likely more along the "fool me twice" category of wisdom. The opposite would instead be a kind of naive thinking.
michaelsbradley 2 hours ago|||
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Are you referring to the burning of the Serapeum in AD 391 or the warehouse fires in 48 BC?

JKCalhoun 2 hours ago||
Is there a difference with regard to the metaphor?
michaelsbradley 1 hour ago||
Depends on the point being made about “historical precedent” and the lessons to be drawn from such.

Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.

JKCalhoun 1 hour ago||
My take was generally: the loss of Alexandria's collection represents a calamitous loss to our collective culture.

(But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").

michaelsbradley 41 minutes ago||
It's a fair point. I honed in on "the burning of" (original comment) versus more generally thinking in terms of "the loss of", because parent context here is "AI companies destroy…".
jan_m_savage 3 hours ago|||
They have already gotten to Archive.org. Books that were available to borrow are no longer 'available'. SMH
JKCalhoun 2 hours ago||
I've scraped all the stuff I am interested in. There's a whole r/datahoarders so I'm not alone. ;-)
alerighi 4 hours ago|||
Yes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services.

Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).

HeWhoLurksLate 4 hours ago|||
we also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up
JKCalhoun 2 hours ago|||
If you are suggesting that air conditioning, plumbing, democracy, and human rights are bad for society then I am missing the analogy.
yehat 4 hours ago|||
Nobody will ask you, they'll be taken from you, in case you missed what happens around.
brookst 3 hours ago|||
The age-old cry of the aging population, faced with tech that didn’t exist when they were young. Turn back time!

I’m reminded of the screeds about the dangers of novels.

himinlomax 3 hours ago|||
They're not destroying rare manuscripts or incunables.

They're destroying one (1) copy of a mass-produced item for each AI company.

Public libraries destroy millions more yearly as a matter of routine.

This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.

fc417fc802 3 hours ago|||
I was with you until the water use. You're misinformed. There were at one point at least several data centers set to use evaporative cooling on well water.

Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.

SXX 3 hours ago|||
Yep. I was personally misinformed the same way at some point. Never expected evaporative cooling to be so popular, but it is.
himinlomax 2 hours ago|||
It's possible for a datacenter to use scarce well water irresponsibly.

They don't have to, and the vast majority don't.

The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.

8note 34 minutes ago||
> a Chinese Communist Party agenda

is it actually? it feels more like US propaganda that we have to let our oligarchs run roughshod over us because of what we imagine the big bad CCP might want.

i dont think the CCP cares whether there's data centers in rural america.

Regulation that requires closed loop cooling seems simple enough, same with lots of the other problems people have with data centers:

* sound and infrasound under x DB

* no air quality change

* must pay to build out electrical infrastructure

etc

its not to the CCPs benefit or loss to make sure the data centers are built well if they get built

TofuLover 3 hours ago||||
> This is just part of a CCP-aligned moral panic, along with the water use nonsense

What was nonsense about water use?

DaSHacka 2 hours ago|||
That AI consumes it at an abnormally high rate, presumably.

The claim never made sense to me either, I can only assume those that regurgitated such claims never worked with HPC or even general datacenters before.

Was recently talking to a (non-technical) friend about this, she was surprised after talking about the "insane water use for AI datacenters" when I responded that open-loop cooling is pretty rare for a datacenter and I've never actually seen it used before, versus closed-loop (or just regular air-based cooling) which has no real noticable water consumption.

brookst 3 hours ago||||
It was never that dramatic, and it’s declining day by day. It’s a panic over a real but small problem.

Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.

SXX 3 hours ago||
Problem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
himinlomax 2 hours ago||
1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere.

2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.

diseasedyak 2 hours ago||||
That AI data centers are drinking up local ground water for cooling. It isn't (or wasn't) nonsense, though. It was/is a real thing, though it seems to be on the out in favor of closed loop cooling after the massive and still on-going public outcry.
alex43578 3 hours ago|||
The outcry over data centers using a fraction of the water used for things like golf courses or growing alfalfa in a desert.
pfdietz 1 hour ago|||
Nuclear killed itself (vast cost overruns); there was no need for hallucinated foreign influences.

But I understand blaming your energy waifu for its own failure is unacceptable for nuclear bros.

8note 32 minutes ago||
nuclear was killed by russian nat gas money.

running nuclear plants were shut down while running just fine

pfdietz 11 minutes ago||
Nuclear in the US was killed by Russian natural gas money?

What other nonsense do you believe?

chrisjj 3 hours ago|||
> Despite the copyright restrictions that are forcing companies to do this

There are none.

shevy-java 3 hours ago||
> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria

Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.

Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.

odyssey7 1 hour ago||
Big AI companies are leaving an easy opportunity on the table for establishing goodwill with the public.

Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.

Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.

Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.

Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.

Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?

RaffaelCH 1 hour ago|
From what I understand, to work with copyrighted books they need to essentially format shift (i.e., scan and destroy the physical book). So a book vault would not solve this issue.

A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's not worth it.

At this point my only hope is that in the long run these scans make it to the public somehow (leaks, copyright changes/expiration, whatever), where they can then be accessed and preserved by everybody. Then we could have our true digital library of Alexandria.

akk0 3 hours ago||
I imagine they are only buying one copy of each book, thus only significantly affecting the supply of books that were already unfathomably rare. That may still be bad, but doesn't really support the "scan every book you can get your hands on before they are gone" narrative.

That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.

dbspin 3 hours ago||
Doubtful. A robust protocol would be to scan several of each edition (to ensure no scanning errors), and scan each edition. Then too, these books are being purchased in lots with accidental duplicates, and all the major labs are doing it. So we're likely talking about tens of each book. For rare books - anything over a couple of hundred years old or small print runs either, that could well be most or even all copies. This wouldn't be immediately noticed either, especially if the books aren't currently considered noteworthy or well known.
OtherShrezzing 1 hour ago|||
>I imagine they are only buying one copy of each book

The previously struggling second hand bookseller in my town has upgraded their car from a 15 year old hatchback Renault to a brand new Range Rover. Some Canadian company has been buying any book he can provide them for the last year. Their quotes aren't by number of books, or even weight, but by volume. As in, they pay him by the shipping container, and he sends several of those a month.

I think it's reasonable to assume the books in this supply chain which aren't destroyed in digitisation are just pulped and sold to Procter & Gamble for toilet paper manufacturing. I can't see any other fate for Anthropic's second and third copies of The twelfth edition of Vera Lynn's 1980's memoir "We'll Meet Again".

JKCalhoun 2 hours ago|||
"I imagine they are only buying one copy of each book…"

I doubt that corporations of this scale do that extra kind of… book-keeping. They more than likely buy books by the pound.

brookst 3 hours ago||
The dissonance when a cause you support is loudly represented by disingenuous types.

At some point they’ll hit on data center water consumption as yet another reason to support the cause.

ironqcold 1 hour ago||
The scale of problem seems a bit overblown. Anna's Archive paint a picture like AI companies are some movie villains burning books so no one can see them. but in reality they just disassemble them into pages because it's cheaper and faster to scan. Most of these books is highly specialized, they been collecting dust on shelves for decades and nobody need them.

But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.

pmoriarty 4 hours ago||
Because of copyright issues, countless books between around the 1930s until about 2000 were never digitized.

After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.

azatom 4 hours ago|
I dunno, that "at least" worries me, digital actually more easier to be lost if it is under copyright, they just got deleted if they can not produce enough money. Physical may have higher chance to survive.
CamelCaseName 4 hours ago||
You ask "Why destroy physical books?"

I ask "Why save physical books?"

If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.

I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.

zero-sharp 3 hours ago||
>If they are truly rare, then they are likely not valuable,

Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.

tele_ski 4 hours ago|||
So these books are not worth anything because they have so few copies but they are still worth including in only their models? Seems a bit contradictory
gruez 3 hours ago||
Not really. A trivial example: smut novels. I'm sure AI companies want them for training so their models work better as AI girlfriends/boyfriends, but I doubt much would be lost if the bottom 50% (by readership) of such books went into a woodchipper.
wasmitnetzen 3 hours ago|||
You're confusing the worth of the book and its content. A book can be valuable (ie a rare bible print), whereas its content is not (we have all the bible variants copied).
karmickoala 1 hour ago|||
They may have been made scarce by many methods, perhaps a low print run; court rulings; burning in a revolution to suppress dissenting views; scanning to avoid competitors to get the content and prevent other LLMs to reference, search, or train on it. Perhaps it's a particular edition that is rare (e.g., the first edition had a different description and was thoroughly changed in the second edition of the book, which may be important to have all the facts from someone's biography).
shiandow 3 hours ago|||
Good point, why waste time deciphering old badly burnt scrolls when anything worthwhile should have been preserved.
noosphr 3 hours ago||
If the Romans had the printing press we'd have a lot more of those scrolls.
mzhaase 1 hour ago|||
What you personally find important is not what everyone else finds important or inspiring. Destroying something takes it away from every single future human being.
ainiriand 4 hours ago|||
That's what state libraries are for, although I understand that sometimes is hard to wrap around the concept of using public money for something different than producing money.
Goronmon 3 hours ago||
Libraries tend to regularly destroy books as well. And they aren't scanning them first either.

Doesn't that make them even worse?

cormorant 3 hours ago||
"State libraries" such as the Library of Congress. (Which is not regularly destroying books, AFAIK.)
tsukurimashou 3 hours ago|||
> If they are truly rare, then they are likely not valuable

"likely" being the keyword here, what about heavily censored books?

madibo3156 1 hour ago|||
Nobody's said it in this thread so I'll drop it here where you ask "Why save physical books?"—

The problem isn't with morals or copyright. What we're up in arms about is case law. Past rulings have implied that destruction of books significantly contributes to the process being "transformative". This encourages companies to destroy the books. I think this is really dumb.

Why save physical books? It's because the reason to destroy them isn't good. If you think there's too many bad books out there, that's a different argument. Maybe your fight is against consumerism, I don't know.

TaLiTr 2 hours ago|||
Massively missing the point. Having paper books isn't the point. Preserving copies of books is the point, so history isn't lost. Often those books only exist as paper copies due to their age. I don't really care what happens to the paper copies, only that their content is preserved in a way that is accessible. Anthropic's private servers aren't it.
gravypod 3 hours ago|||
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!

yehat 3 hours ago|||
I ask "more clean air", you answer "why, did you deserve it, having clean air is not free, there's a real cost" Somebody else asks "We need more accessible energy", you answer "why we bother with your needs, you're not efficient, energy belongs to more efficient purposes, you can live without that much energy". I can continue with more, but hope you got another viewpoint. Btw, I'm disgusted there are "humans" like you in existence. We definitely don't share the same cultural ancestry, and I hope ours will prevail at the end, rather than cold blooded, mechanical "brains" like those of your kind.
PartiallyTyped 3 hours ago|||
> If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.

The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.

Same can be said about many books.

hahn-kev 1 hour ago|||
Right but is an AI company going to destroy a Gutenberg bible? No, as you said the content is not rare, and that's what they want.
gruez 3 hours ago||||
So old books are valuable because old paper is valuable? I understand why the Gutenberg bible might be valuable but do we really need thousands of mass paperback novels?
jll29 3 hours ago|||
Ask: valuable to whom and why?

- Reader: narrative

- Collector: scarcity of the physical artifact

- AI Company: language samples (quantity, variety), facts

vasco 4 hours ago||
I agree with you for the same reason I think McDonald's is the best restaurant in the world!
tescreal 1 hour ago|
A few points for people: 1. some books are out of print. 2. some books CANNOT return to print. 3. all books prior to the 21st century are products of human minds. 4. copyright extends over the vast majority of printed material due to acceleration of literacy and printing access. 5. not all people value all books equally. 6. most books have a degree of historical interest (even cookbooks, which can say a lot about the economic health of a region when it is printes. culture is also clearly encoded in them). 7. many books are already lost, and historians are the ones who most voice the harms. 8. when a book is absorbed into the machine, it may remain vaugely accessible, but only on the good grace of the ones who pilfered it. 9. if no existant copies remain, then the price for access becomes effectively infinite. 10. removal of books denies human agency over access to information. 11. costs will follow a steepening curve much as ram did.

first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.

More comments...