Top
Best
New

Posted by pluc 5 days ago

Microsoft exec called AI scraping 'the largest theft of labor in human history'(techcrunch.com)
950 points | 832 commentspage 2
heaney-555 4 days ago|
LLMs are not compression algorithms. From an information theory perspective, that's impossible given their size.

Thus, a distinction needs to be made between viewing material to _learn_ and viewing material to _verbatim repeat_.

It's not illegal to read the New York Times and then start giving paid advice based on what you learned, as long as you don't repeat the text verbatim.

cush 4 days ago||
It’s irrelevant if it’s legal today or not. This is new technology and may be new precedent.
Timon3 3 days ago|||
I can see how LLMs can't be lossless compression algorithms, but why not lossy?
bustadjustme 4 days ago|||
... but you have to pay to read the NYT. You paid for the information. Guess who didn't.
someguynamedq 3 days ago||
Citation needed. Of course they are compression algorithms
thunkshift1 4 days ago||
This will lead to a massive settlement between the big boys and most people who put stuff out in good faith will be left out of it. And that will be the end of it. We will never hear anything about this ever again and the ‘theft’ will continue like normal.
Weryj 5 days ago||
I think it’s more like ‘The absolute maximum possible degree of theft’ there can’t be larger, it’s everything current and past.
totetsu 5 days ago||
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
iamflimflam1 5 days ago||
I don’t mind these companies scraping my content.

But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.

cmiles8 4 days ago||
The evidence here is quite damning for OpenAI and Microsoft is clearly trying to distance themselves from OpenAI’s behavior here.
American87 5 days ago||
I remember techchrunch.com making the argument that IP Infringment != Theft in the music piracy era.. how quickly the tide turns :)
mitxela 5 days ago|
they did say theft of labor, not theft of the things being trained on
jaybeavers 5 days ago||
You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.
Joel_Mckay 4 days ago|
There is nothing funny about $9Tn in FOSS getting misappropriated, and sold as isomorphic plagiarism tokens... or the estimated $4.6Tn in loan debts these 7 companies incinerated when the bubble pops.

Anyone that lived through the dot-com or housing market bubble know what a collapsing Ponzi scheme does to real businesses, and peoples retirement funds.

Popcorn ready =3

nullbio 5 days ago||
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
rich_sasha 5 days ago||
Yeah, distilled, hosted by OpenAI and charged for. And don’t you try reverse engineer what they did!

If this was all open, I’d maybe half agree.

riskable 5 days ago||
That's what this boils down to: What are our rights?

Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.

If we take away that right, how would the Internet even work? It wouldn't.

Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.

If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!

People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.

The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.

My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.

But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.

anon7000 4 days ago||
> People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.

And profiting on it on a scale that’s hard to fathom. Someone who spent effort creating a great resource or doing some research and maybe got some income via donations, ads, whatever. Now that information from their resource is distilled into a big model. The original author is screwed, the model provider makes money through the effort of everyone else. It worked well for everyone before, because there was recognition, prestige, a sense of doing good for people, even a chance for some income. That’s completely eliminated with AI.

sebastiangrill 5 days ago|
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
simonw_simonw_ 5 days ago|
[dead]
More comments...