Top
Best
New

Posted by darccio 7 hours ago

AI companies destroy physical books – let's scan rare books before it's too late(annas-archive.pk)
612 points | 373 commentspage 2
yipinwong 1 hour ago|
The post reads like a propaganda. Evoking feeling over metrics or results

(<- that's what propaganda does by definition)

e.g.

> but ethically, it’s an extremely serious crime against humanity.

Who are you that you consider it a serious crime against humanity? What are you a saint?

> Knowledge is permanently monopolized on private servers.

Throughout human history, it's the knowledge/info that provide one with wealth and advantage over others.

With the author's logic, every single private knowledge the author is not sharing is an extremely serious crime against humanity.

The logic of the writing is not correct.

alightsoul 1 hour ago||
Isn't this what Dario does?
yipinwong 15 minutes ago||
If Dario appeals to feeling, yes.

Provide more context if you want a better feedback not whataboutism.

rcarr 2 hours ago||
Does adding an old book risk making the model worse? Off the top of my head:

- Reinforcing outdated, disproved or otherwise incorrect information.

- Reinforcing outdated forms of communication e.g purple prose.

You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory

thth123123 29 minutes ago|
[dead]
joshuakelleyds 1 hour ago||
I keep seeing headlines, videos, etc and the recent copyright court case, Anthropic v. Bartz (1.5 billion dollars) gives the best context around this. I encourage everyone to read the full thing, but here are some excerpts:

> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order

Also the summary:

> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.

https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

theartfuldodger 2 hours ago||
I collect old books. It's not common to buy books at all yet to buy old books. Library sales as big as concerts exist, little book libraries everywhere but the used book stores are constantly closing.

I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.

1970-01-01 1 hour ago||
This is a good example of immature writing. Buried on the bottom of the page are two links, both revealing internal tickets that are hinting as to how I can actually help. There should be big, bold, easy to follow steps for volunteers.
throwaw12 2 hours ago||
America is interesting.

* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz

* download and publish a book as a company -> fine 1% of revenue

* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations

* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model

tptacek 53 minutes ago||
I think this is a duplicate of a story that ran on the front page yesterday:

https://news.ycombinator.com/item?id=49383026

The big thing here is: libraries and the book trade destroy millions of books every year. If you clean your attic out, box up all your old books, and bring them to your local library donation box, they'll quickly sort through it for things that might actually circulate, and the rest go right to the recycling center.

A lot of people in these comment threads seem not to understand that destruction is part of the natural lifecycle of a book. Books generally don't get preserved.

And then there's the problem that the original reporting that kicked all this off, at 404, is specific about what is meant by "rare books". It's not, as they say, first editions of Oliver Twist. Rather, these are books nobody cares about; that's what makes them rare in the first place. Vanity press stuff, or manuals for old equipment that isn't produced anymore. All these books would naturally end up a dumpster.

branon 4 hours ago||
I have a year membership to AA now, been meaning to contribute for some time, archive.org is next on the list

Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded

afpx 4 hours ago|
archive.org is an incredible blessing that I never really respected enough until the last few years. I have bookmarks going back to the 90s, and for some reason, starting in 2015 or so, sites just started disappearing. I estimate at least 20% of my bookmarks are 404 now.
cormorant 3 hours ago||
archive.org itself reminds me of the Library of Alexandria. It is irreplaceable, and this is a disaster waiting to happen.
pino83 1 hour ago||
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?

And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?

Why is that?

If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?

shrubble 4 hours ago|
AI companies could certainly release the scans of any book that is no longer covered by copyright in the USA, for free download and get some positive publicity for a change.

I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...

dqv 4 hours ago||
The other thing they could be doing is using some of that lobbying money to try to reform copyright law to allow them to release the scans that are still covered.

It would likewise earn goodwill from a lot of people.

shagie 3 hours ago|||
Copyright law is an implementation of international treaties.

Berne Convention https://en.wikipedia.org/wiki/Berne_Convention (182 parties)

TRIPS Agreement https://en.wikipedia.org/wiki/TRIPS_Agreement (164 parties - part of WTO)

The United States can't make copyright weaker than what those agreements require without pulling out of the WTO.

The core of copyright law is about who has the right to redistribute a work. If I buy a print of a photograph, scan it and use that as my desktop image... I can do that. I cannot redistribute the scanned image, and if I was to sell the print later I should delete the scanned image.

Note that format shifting is covered under fair use... which is what training is taking place under. However, that doesn't mean that they can release that format shifted content... nor can then re-release the original work if they are retaining the format shifted content.

https://library.georgetown.edu/copyright/fair-use-reformatti...

    Under § 106 of the Copyright Law of the United States, the owner of the copyright in a work has the exclusive right to make copies of that work, unless an exception applies. When considering reformatting media, please note that individuals do not have an automatic right to reformat a work from one format to another. In order to legally convert media, your use must fall into one of the following categories:

    you own the copyright in the work,
    you have permission from the owner of the copyright, or
    you have done a fair use analysis and have determined that fair use applies
phoronixrly 4 hours ago|||
They could be, however that would weaken their position as the sole source of knowledge which seems to me is half the purpose of their book destroying initiative.
dqv 4 hours ago||
That's why the "but it's illegal" propaganda is so unconvincing for me. It's really "but it's illegal (and we want to keep it that way)".
archonis 3 hours ago||
PR only cares that the company is percieved as unstoppable & inevitable.

If their PR teams did have a response to these actions, it would be something to the effect of:

"would you rather china destroy all the books and gatekeep the knowledge?"

More comments...