Top
Best
New

Posted by darccio 7 hours ago

AI companies destroy physical books – let's scan rare books before it's too late(annas-archive.pk)
668 points | 393 commentspage 3
ryandvm 5 hours ago||
It seems like it would be well within the charter of the Library of Congress to archive a complete scan of every book published in the United States. At least then we wouldn't lose the information. They can figure out distribution and copyright later.
JsonDemWitOster 5 hours ago|
The US Library of Congress is already a book depository, i.e., it has a copy of every book published in the United States. Same for the British Library for the UK and Ireland. Similar depositories exist for most other countries who care for their culture.

Which is really why the outrage cycle over Anthropic's actions is largely misplaced.

pino83 2 hours ago||
What was worse: Putting all of our communication since ~2010 into a commercial walled garden? Or some books that were lying around in some bookstores or whatever (i.e. that nobody was interested in owning so far)?

And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?

Why is that?

If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?

thisisauserid 5 hours ago|||
Second Circuit Court of Appeals ruled (in favor of Google) 2015 that similar actions constituted "fair use".
ZoomZoomZoom 5 hours ago|||
The main question is why aren't they leaking it to AA themselves? Trying to keep an edge with their training sets? Isn't it ridiculous, considering the sheer size of them and statistical insignificance of the set differences?
Aurornis 4 hours ago||
> The main question is why aren't they leaking it to AA themselves?

How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.

They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.

notpushkin 3 hours ago||
> You cannot scan a book and share it without violating copyright law.

Hence the “leaking” part.

embedding-shape 5 hours ago|||
> The main question is why aren't they leaking it to AA themselves?

Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.

Filligree 5 hours ago||
Because that’s illegal.
onionisafruit 5 hours ago|||
Exactly. They set up this operation specifically to comply with the letter of copyright law and defend against publisher law suits. “Leaking” to Anna’s Archive is the last thing they’re going to do.
ZoomZoomZoom 4 hours ago|||
The illegal part is (or ideally should be) using the books in their training. The morally right action is to make it public afterwards.
jmspring 2 hours ago|||
I can't recall the company, it's been several years, but they were scanning rare texts in detail and making them available online. It wasn't Project Ocean/Google Books - that said I don't trust Google to be a good steward here.
clarionbell 3 hours ago|||
In my experience, when someone mentions burning library of Alexandria, they are either exaggerating, or have a poor grasp of history. Usually it's both. This post, the discussion, do not change my mind.
twright 3 hours ago|||
I'm a little confused about the value of scanning rare books since this story came out. I have a small collection of "rare" books and they're not really bounties of information, at least not modern information. I know novel training corpus is important but the information in rare non-fiction books is commonplace or out-dated. And the information in old rare fiction-books are originals for which reprints exist or just uninteresting stories that aren't really worth anyone's time except collectors'.
chistev 2 hours ago|||
What is the deal with AI Companies buying old books to scan and then destroy them?

https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...

tptacek 1 hour ago||
I think this is a duplicate of a story that ran on the front page yesterday:

https://news.ycombinator.com/item?id=49383026

The big thing here is: libraries and the book trade destroy millions of books every year. If you clean your attic out, box up all your old books, and bring them to your local library donation box, they'll quickly sort through it for things that might actually circulate, and the rest go right to the recycling center.

A lot of people in these comment threads seem not to understand that destruction is part of the natural lifecycle of a book. Books generally don't get preserved.

And then there's the problem that the original reporting that kicked all this off, at 404, is specific about what is meant by "rare books". It's not, as they say, first editions of Oliver Twist. Rather, these are books nobody cares about; that's what makes them rare in the first place. Vanity press stuff, or manuals for old equipment that isn't produced anymore. All these books would naturally end up a dumpster.

juiceland 4 hours ago||
Why are people acting like books can’t be reprinted?
letmetweakit 3 hours ago|
You have to have a sample to be able to reprint it.
juiceland 3 hours ago||
Good thing the books are being scanned, then. The copyright owners know just who to go to get the sample.
adamors 2 hours ago||
They’re not giving anybody any samples lol
juiceland 2 hours ago||
Has anyone asked? From the discussion here it seems like the books are extremely valuable.
More comments...