Top
Best
New

Posted by pluc 5 days ago

Microsoft exec called AI scraping 'the largest theft of labor in human history'(techcrunch.com)
950 points | 832 comments
haritha-j 5 days ago|
I just don't understand people saying "but a human learning from a book isn't illegal".

How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."

And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.

coffeefirst 4 days ago||
Because it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book.

Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.

ryandrake 4 days ago|||
It's the same bad faith argument as:

"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."

disqard 4 days ago|||
I haven't seen this articulation before, and I'm going to remember this one!

In return, I'll share my own reasoning:

Human Life is finite, and there is a real opportunity cost to learning (say) to copy Picasso's style -- you could've been doing something else in that same Time.

However, if you're training a model on the entire career output of dozens of artists, it only needs more electricity + GPUs to do this.

And, transposing this back to the Human realm, there's no way one person can put in that kind of wide effort to learn how to copy dozens of artists' styles in one lifetime.

I really like the way you put it too -- it's a bit more succinct.

lovich 4 days ago||||
The quantity of an action can have a qualitative difference. I wish I didn’t haven’t to explain that to so many people everytime a new technology is created.
hackerdood 4 days ago||
Quantity has a quality all its own
0x445442 4 days ago|||
I don't know if it's even OK for a cop to just randomly watch a random street corner without a specific reason to do so. If a cop was following a given person around all day every day without probable cause would this not be considered harassment? If we view the Flock Camera network as a single system is this not harassment?
elmomle 4 days ago||
Just to add context and not to comment on rightness, that was one of the early core purposes of policing--to ensure that a people in a particular public space (where either the people or the space are particularly vulnerable) are free from unwanted disturbance. Here is that exact concept shown in a painting from 172 years ago: https://en.wikipedia.org/wiki/The_Gleaners_(Jules_Breton).
4d4m 2 days ago||||
Correct take according to law. A mechanical style royalty would be requisite
j4kp07 4 days ago|||
[dead]
kenjackson 4 days ago|||
"How do people not understand that some laws only make sense at a certain scale? "

Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.

devsda 4 days ago|||
Laws do consider scale and we can see that in everyday law.

4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.

Moving around with a couple of bills is treated differently from carrying huge bundles of cash.

Those laws or exceptions were probably added later as a reaction to abuse of existing laws.

The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.

It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.

ethbr1 4 days ago||
Laws are typically implicitly considerate of scale. Very rarely are they explicitly considerate of scale.

To wit, sharing music digitally, Google scanning books, or Uber/Lyft providing unlicensed taxi services.

It's pretty rare that laws consider what should happen if it were suddenly possible to 10x or 100x preexisting throughput.

And specifically where laws balance multiple, often-opposed, stakeholders' interests, that change can drastically upset the previously negotiated compromise.

Which is why piracy at scale, before it's banned, tends to be a successful foundation for many businesses.

morkalork 4 days ago|||
Murder, terrorism and genocide.

Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.

I'm sure there's other examples too.

mrec 4 days ago||
Theft was historically very scale-sensitive. In England you'd face execution for "grand larceny" if the goods in question were worth more than twelve pence. This was on the books from the 13th century until 1832 (!).
encrypted_bird 4 days ago||
Jesus Christ on a pogo stick, execution for stealing 12 pence worth of stuff? That's crazy! Adjusting for inflation[1][2], that would be the same as nowadays being executed for stealing £5!

[1] According to https://www.bankofengland.co.uk/education/education-resource..., 12 pence (1 shilling, or 1/20 of a pound)) in the old £sd system is equal to 5 pence (or £0.05) in the later decimal system.

[2] According to https://www.bankofengland.co.uk/monetary-policy/inflation/in..., £0.05 (decimal) in 1832 would be equivalent to £5.02 in 2026 Aug.

mrec 4 days ago||
When first introduced it was around the price of a sheep. The homicidal version of bracket creep [1].

[1] https://en.wikipedia.org/wiki/Bracket_creep

encrypted_bird 3 days ago||
Damn, sheep were cheap back then.
leobg 4 days ago|||
Because copyright is a tradeoff. Society wants information to be free. But authors won’t publish if anyone can republish their work for free. Hence the distinction: The ideas you write are not protected - only your expression is.

In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.

vharuck 4 days ago|||
Allowing AI companies to get away with this because we're "better off"¹ is like eating our seeds. Sure, we built something cool with the massive corpus available, but how will it affect future decisions to develop or share creative work?

If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).

1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.

ethbr1 4 days ago||
Agreed. The most damning behavior was Meta.

Specifically, reaching out and discussing licensing material, then pirating because it was too expensive/slow to legally acquire it.

Heaven forbid Meta have to pay for something.

dhx 4 days ago||||
> authors won’t publish if anyone can republish their work for free

There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).

[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.

intended 4 days ago||
Again, incentives.

The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.

People who need renumeration to continue working will not publish.

Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.

make3 4 days ago|||
>Society profits from having AI

Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.

moscoe 4 days ago|||
I really don’t understand why anyone would think they are entitled to any part of a derivative of a published piece. Why publish if not to help advance humanity, just like all who came before you and contributed to your success/ability? This is quite literally the bedrock of the progress of human civilization. As long as they aren’t straight up reproducing a direct copy of the original work.

If you want to keep it to yourself so only you can benefit from it, then keep it private. Otherwise, why don’t you get to work on the next big idea.

rrr_oh_man 4 days ago||
> As long as they aren’t straight up reproducing a direct copy of the original work

I think that part is debatable

gigel82 4 days ago|||
That's eerily similar to the argument used by mass surveillance systems like Flock around expectations of privacy in public.

Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.

DownGoat 5 days ago|||
I agree that there is not an orange to orange comparison, but I have still not seen a law that could scale as well. The example that comes to mind with a proposed law like this is how would you license work that build on another work? What if I decide to publish a blog post after taking some course, that distills whatever I learned in the course for free?
californical 4 days ago||
We already have laws that cover this.

If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place.

If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original.

And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property.

Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans.

I do think copyrights should be shortened to 20 years but that’s another discussion

throwawayk7h 3 days ago|||
Many people's moral systems (and our legal system) are typically deontological. That is to say, the consequence of an action (or scale) aren't relevant to whether it's okay or not, so long as the action doesn't break any particular moral rule.

There's no rule against learning from books in general. (Maybe only for some particular books.) There's no rule against using tools to read books (it's okay to wear glasses)

The only deontological argument I can see against training LLMs on copyrighted data is that some people think it's morally and legally wrong to make derivative works (such as fanfiction) without permission, and the weights of the LLM could be seen as a derivative work.

This may sound stupid but it's the same argument people use in favour of adblockers. When ads were first introduced to the internet, it was understood that different people could view the internet however they liked and they could choose a "user agent" to display content to their preference. So ads were just a nuisance but could be worked around easily -- it was the ad provider who was the fool. Now I often hear people say that ad blockers are unethical; they deprive content creators of their income, or they're deceptive, or criminal. This may indeed be true. But at no point in time did any moral rule suddenly change.

sweetjuly 3 days ago||
> Many people's moral systems (and our legal system) are typically deontological

I think there's also a component of people (HN's audience in particular) trying to approach the law as if it were a program. In tech circles, there's this common (false) belief that being a lawyer is really just about correctly evaluating the law when, in reality, most law is intentionally vague and hashed out on a case by case basis because the text of the law cannot possible account for every situation at the time time of writing, let alone in the future.

tosti 3 days ago||
A reuseable program is similarly intentionally generic. Compare ifupdown from Debian and NetBSD with netplan. Both came out without any knowledge of wireguard. Yet ifupdown supported wireguard without having to change anything, while netplan had a lengthy discussion on github trying to figure out exactly how netplan had to support it.

Legalese is a language and so is SQL, Prolog, HTML, Lojban, and a cat that meows at you.

hn_throwaway_99 4 days ago|||
> How do people not understand that some laws only make sense at a certain scale?

100% agree, and the problem is that technology moves a lot faster than law can keep up. Just look at the Flock brouhaha. Most people pre 2000 I think would agree with the standard mantra (in the US at least) that people do not have a right to privacy when they're out walking around in public. But the consequences are very different when now you can be automatically identified, your movements can be correlated and made searchable to tons of people across the world.

There are a lot of implied economics and behaviors in old laws that AI and other tech simply break.

micromacrofoot 4 days ago|||
This is a similar problem with data brokers. They take public data laws to an extreme and resell easy access to the aggregated data. This easy access has created a tremendous number of problems unforeseen by the original intent of the access.

Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.

Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.

swid 4 days ago|||
My problem with your solution to require author permission to train is that it doesn’t solve anything long term.

What happens when all the licensed information still leads to the creation of demand hoarding AI? Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.

TeMPOraL 4 days ago||
> Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.

Well then they cannot possibly stealing this, and short of creating laws that directly discriminate between algorithmic processing and human consumption - regulating the process, not the subject - this argument is quite literally nonsense.

arendtio 3 days ago|||
I think the core question is what derived work is. I think derivative work should not be forbidden, and an LLM can be that, but it can also be just a copy. So, in essence, it is just a medium, like a sheet of paper.
t0mpr1c3 1 day ago||
Exactly. There is a distinction between learning from a book and memorizing the entire content.

The boundary is not entirely distinct, but if your facsimile is poaching 93% of the revenue of the original then it deserves some scrutiny.

capr 4 days ago|||
Because people want to discover if a law is principled, and "certain scale" is not a principle, is just an arbitrary utilitarian distinction.
tim333 4 days ago|||
If China keeps training AI off the web while US companies have to negotiate with 7,383,654 different rights holders, it's going to handicap the US companies a bit.
solid_fuel 4 days ago|||
You can also use this exact argument to call for a return to chattel slavery, an end to all pollution controls, legalized corporate death squads, and basically any other heinous act you want. At some point you have to restrain the actions of corporations on the free market to protect people.
underlipton 4 days ago|||
Yep. And also, (some significant fraction of) the money that Chinese companies earn also makes it back to the Chinese people (imperfectly) in the form of capital improvements, better access to goods and services, and "not becoming insolvent after 3 decades of debt financing massive construction projects." Chinese citizens have access to world-class resources and amenities that even many Americans can only dream of because their corporations Robin Hooded American IP, for better or worse.

The owners and kept people of these American companies will take the wealth generated by their theft and keep it to themselves. They talk about a "permanent underclass" with disguised glee. Break their operation until they learn noblesse oblige.

tim333 4 days ago|||
I think it's a little different. Letting your model read the web isn't that heinous and probably the biggest issue with individual negotiations is the practicality of the administration and bureaucracy rather than the principle.

(by the way if any AIs want to train on this comment, I give permission in return for $10 sent by paypal per LLM)

intended 4 days ago|||
If the US companies had to negotiate with the 7.3 million rights holders, we wouldn’t be here, and we would have probably got a system where the models were never released in a manner that could be easily distilled.

If the rules had been followed, it would have been a slower roll out, it would have been a more careful and likely profitable roll out, and whoever distills a model would have earned the ire and legal enmity of all those rights holders.

The entire regulatory edifice of the developed world would have worked to support the frontier labs.

Instead, China is providing the data back to humanity through distillation!

maplethorpe 4 days ago|||
I don't think the people pushing that narrative genuinely believe what they're saying. They're just influencing public opinion.
deaton 4 days ago|||
I think that last thing is key. The authors of these things published them with the express intent that other people would read them. Not that they would be used as training data.
GPerson 4 days ago|||
It’s because most people contrive arguments to support what they feel. They don’t examine arguments on their merits and then decide what’s right.
Sohcahtoa82 4 days ago|||
> How do people not understand that some laws only make sense at a certain scale?

Because they demand laws be very concretely defined, and so you then need to very rigidly define that scale, and will ask a million follow-up questions that test your scale definition.

But of course, it's all bullshit. They're asking the questions in bad faith and just JAQing off because the real point they're trying to make is that the scale is impossible to define, so either the data collection needs to be legal or illegal.

SR2Z 4 days ago||
I would hate to live in a world where questions of "legal or illegal" are "impossible to define."
Nasrudith 4 days ago|||
Being anti-permission culture is not the same thing as bad faith. It is just rejecting your arguement as stating scale suddely changes the rules. Plus the whole argument based on supposed impact is missing the point like saying that free speech should be treated differently for being more impactful than anticipated with mass publication. Something throughly rejected.
keeda 3 days ago|||
> How do people not understand that some laws only make sense at a certain scale?

Consider that any other laws you may want in place could be even worse, and what we have with AI is the logical culmination of technology and the laws we as a society have established over centuries of dealing with hairy issues based on sound principles:

https://news.ycombinator.com/item?id=49761887

Tl;dr: AI has harvested that which we as a society have very explicitly decided should belong to the commons.

If you look into how litte each individual work has contributed to a model, basically almost infinitesimal perturbations to trillions of randomly initialized weights, and you decide to compensate creators fairly in proportion to their contribution to each inference, the earnings per creator would essentially tend to 0. Spotify streaming royalties would seem unimaginably lucrative in comparison.

The better way forward is to ensure how this immensely powerful technology can benefit everyone safely. New forms of compensation will need to be evolved, for sure. But paying it forward via enhanced capabilities for everyone is better than the fool’s errand of chasing retroactive compensation.

heresie-dabord 3 days ago|||
When colluding oligarchs agree on a course of action, the laws will bend or be bent.
penguinova 4 days ago|||
[flagged]
93po 4 days ago||
I would argue artists imitating, for example, the artists who revolutionize a genre of music are also in fact massively reducing the demand for the originals by 99%. Imagine if no one ever made a song after the Beetles that sounded any newer. The demand for Beetles music would have remained somehow even more enormous than it was for the last 50 years. Or if no one did abstract paintings after Kandinsky, or impressionism after Monet or cubism after Picasso.
47282847 5 days ago||
“Information wants to be free“.

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

proc0 5 days ago||
"I just wished the collected data was public. "

That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.

Gareth321 4 days ago|||
I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.
ndiddy 4 days ago|||
IP does not protect the little guy. This is nothing new. Draw a picture and then people start putting it on t-shirts and posters without paying you? Great, you can't do anything about it unless you have enough time and money to hire a lawyer to go after them. Self publish a book and then people start uploading PDFs of it? Better hope your real passion is filing takedown requests instead of writing.
ryandrake 4 days ago||||
AI training is the clearest example yet that companies are allowed to get away with what is treated as a serious crime only when an individual does it. There are many more examples of this, of course, but this one seems to be the most stark and obvious.
tstrimple 4 days ago||||
Legislators have consistently picked a lane. Protect the rich and powerful.
underlipton 4 days ago|||
Huey Freeman: "Kim Dotcom was pissed."
blfr 4 days ago||||
Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.
cute_boi 4 days ago||
I don't think this trend of open sourcing LLM will continue for a simple reason: Money.
jrflo 4 days ago||||
I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.
nosyke 4 days ago|||
I think the difference is that people are investing significant amounts of time, effort, and money into transforming their work into something useful, so naturally they would like some return on that investment. Giving away that work for free is not a particularly good business model if those people expect to be compensated for the value they create.
intended 4 days ago|||
There’s too many implicit assumptions in that sentence that run afoul of the conversation.

Just spending money doesn’t mean it’s legal, for example. Criminals expect RoI too.

CJefferson 4 days ago||||
Anthropic getting angry other AIs are trained on their AIs output is, to me, one of the stupidest things I’ve read in a while.
DownGoat 5 days ago|||
On the flip side, output from an LLM is not copyrighted.
tremon 4 days ago||
That has not been decided. The only thing that's been decided is that the LLM itself does not have copyright on its output.
polytely 5 days ago|||
I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
juiceland 4 days ago||
> I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.

Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.

intended 4 days ago|||
How will you redistribute to people in Japan, for example?

The profits and income is earned in America. The idea that America would pay manga artists whose work was copied is … beyond idealistic.

nosyke 4 days ago||||
the problem at this point is none of those AI companies are profitable or are even flirting with the possibility of being profitable
tzs 4 days ago|||
Nope. You are making one or both of these mistakes. (1) Overlooking that the same logic can lead to different outcomes depending on the premises to which the logic is applied. (2) Overlooking that real life is analog, not digital, and so thinking the premises are the same when they are not.
juiceland 4 days ago||
The targeted outcome is the redistribution of wealth away from the capital class to the working class. Capitalist exploitation of labor was analog to begin with, and the same premise does apply: the capital class absorbs the fruits of labor, training, and education that is performed by the masses in order to enrich themselves. The capital class owes a tremendous debt to society and if they don't plan on paying we should plan to seize it.
Nasrudith 4 days ago||
And how has exproproation been working out for you? What is that? Massive poverty and nobody wants to trade with you?
juiceland 4 days ago|||
The exploitation of labor has has a tremendous negative impact on the environment and the health of humans. Recall that it took dragging the factory bosses from their and beating them to death to get an 8 hour work day, a weekend, and restrictions on child labor.

We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.

underlipton 4 days ago|||
Pretty sure we had to sanction/assassinate/goad into self-destructive military campaigns the Red Terror to beat it. And even then, the major survivor still beat us to cyberpunk dystopia (the cool one with hologram skyscrapers, not the uncool one with decaying suburbs).
fwlr 5 days ago|||
Well the future we seem to be getting is “information wants to be free for the first ten thousand tokens, then $1 per million tokens after”.
TeMPOraL 4 days ago||
Your point notwithstanding, that's still a bargain.
alentred 5 days ago|||
> Sharing information is an act of love

Most AI companies are not sharing it, though. They appropriated it and resell it.

pier25 5 days ago|||
There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.
cush 4 days ago|||
Yikes. That’s some deep entitlement.

Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.

If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet

These AI companies actively chose not to get consent from creators and earn billions from their content with no compensation.

r3trohack3r 4 days ago|||
IIUC, the question at hand is: does training require a special, separate, license or can you legally acquire a work and then use it for training?

I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.

cush 3 days ago||
> does training require a special, separate, license or can you legally acquire a work and then use it for training?

No. But it's not about current precedence or legality because the legal framework for accurately (according to general moral and societal acceptance) is decades behind where it needs to be. The courts will decide over the next few years.

armchairhacker 4 days ago|||
In the past, but today fewer people are getting paid less this way.

Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?

avidruntime 4 days ago||
> In the past, but today fewer people are getting paid less this way.

There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.

> Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?

Uploading content online and getting a cut of ad revenue fits this criteria, no?

The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.

TeMPOraL 4 days ago|||
> There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.

And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.

armchairhacker 4 days ago|||
Ad revenue is declining. Most creators make money off sponsorships and Patreon.
JoeJonathan 4 days ago|||
Just because “information wants to be free” is a thing people say doesn’t mean it’s true.
nisegami 5 days ago|||
>"Information wants to be free".

I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.

ex-aws-dude 4 days ago|||
> anyone producing content, everyone’s creativity, is fed by something that others did before

The thing I produce does not replace demand for the original though?

I can’t recite the original for a million people

eli_gottlieb 4 days ago|||
> It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?

DanielHB 5 days ago|||
The key problem is that IP is either proprietary to the creator or it is a commons type of situation.

Even if you agree with the former exploiting the commons for personal profit is... not good.

One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.

Mali- 4 days ago|||
By your argument, I should be able to directly copy a book and sell it myself. The work was already done! What's more, they were fed by everyone's creativity, so of course I should be able to sell an exact copy.

In reality, the short-sighted greed is allowing widespread theft of intellectual property; do you think the number of writers would increase or decrease if there were no protections against content theft?

If you have such "immense gratitude", pay for the work.

capr 4 days ago||
Most likely increase, despite your intuition. This argument has been debunked so many times both intellectually and empirically. For one, read "Against Intellectual Property" by Stephan Kinsella.
tzs 4 days ago|||
> “Information wants to be free“

What about the rest of that quote?

whiplash451 4 days ago|||
I think you’re confusing fan in and fan out.
equality_1138 4 days ago|||
You wouldn't be standing on the shoulders of giants without IP laws, bub. Your "immense gratitude" is a farce.
bcjdjsndon 5 days ago|||
[flagged]
ambicapter 5 days ago|||
Probably because it claims there’s no problem, and then makes a tiny little mention of the BIG problem at the very end.
newsclues 5 days ago|||
The unpopular part is the hypocrisy where their work has already or must be rewarded while other people’s work is not.
aners_xyz 5 days ago|||
[flagged]
jayd16 4 days ago||
So by definition creative labor cannot be stolen? Seems like flawed logic to me.
beering 4 days ago||
Taking the original of a painting from your house is stealing. Copying is only potentially violating the government-granted limited-time exclusivity that allows you to decide who can copy your work.

BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.

jayd16 4 days ago||
It's telling that you need to fall back to non-creative information in your argument.

Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.

Extrapolated further, this is a collapse of creative works as a profession.

sajithdilshan 5 days ago||
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
y-curious 5 days ago||
Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
bcjdjsndon 5 days ago||
But it's copying...how is it theft? Your labour WASNT stolen was it?
Avicebron 5 days ago|||
In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?
cluckindan 5 days ago|||
Even better question: would you be willing to train the AI-driven robot which will end your profession altogether?
px43 5 days ago||
I work in infosec and I would give up everything I own and die happy if infosec became a solved problem and the profession died forever.

It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.

tremon 4 days ago|||
I'm not sure how to square

> It's hard for me to imagine a profession that should exist in a utopian society

with

> People should have instant access to food when they're hungry and housing when the weather gets bad

You don't think that farming, baking, and building are professions?

NekkoDroid 4 days ago|||
I assume they mean that it is all already automated in the utopian society. And I would kinda agree that that would be utopian. But sadly we live in the real world where greed rules pretty much everything, meaning that it will remain an unobtainable dream for (most likely) ever.
ryandrake 4 days ago||
There's unfortunately a huge contingent of people who would strongly oppose a utopia where we all could live as equals in comfortable, luxurious lifestyles, where robotics, replicators, AI, and automation took care of everyone's needs and wants.

They oppose the "live as equals" part.

bcjdjsndon 1 day ago||
> we all could live as equals in comfortable, luxurious lifestyles, where robotics, replicators, AI, and automation took care of everyone's needs and wants.

Until one person decides we should change something about this utopia, arguments break out, populations schism, and we're back to fighting over finite resources once again.

Utopias and conflict-free societies are pipedreams

eli_gottlieb 4 days ago|||
Farming and baking are already very thoroughly automated and their labor force completely proletarianized.
tsoukase 3 days ago|||
There were the fundaments of communism two centuries ago.
SkyBelow 5 days ago|||
Do we normally consider recorded music to be theft from musicians who would have been paid to play music live if we never allowed (or invented) recorded music?

Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.

But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?

Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.

nitwit005 4 days ago|||
> Normally this isn't the case for any technology except for the time it first comes around.

Well, sure, there's no one left to fight once everyone already lost their job and moved to another career. But, that's sort of a "might makes right" resolution. The workers don't have the political sway needed to get the government to intervene.

Businesses succeed in getting that sort of market intervention all the time. Most of the modern changes to copyright law are driven by business lobbying.

bcjdjsndon 20 hours ago||
> Well, sure, there's no one left to fight once everyone already lost their job and moved to another career

You'd rather people still did old, obsolete jobs?

nitwit005 17 hours ago||
It's telling that you're concerned about that case, but not concerned about companies successfully getting legal changes.
max__dev 4 days ago|||
Musicians are paid to record music. If they were not, as is the case for ai data, we would indeed consider it theft.
Windchaser 4 days ago|||
> Musicians are paid to record music. If they were not, as is the case for ai data, we would indeed consider it theft.

I remember reading literature from the 1930s, and there were quite a few folks who thought that the musicians-doing-recordings were stealing from the old-timers who played for live audiences.

History does not exactly repeat itself, but it rhymes

SkyBelow 4 days ago|||
The few who get paid to record are replacing far more who would have otherwise been paid to play if that was the only way to listen to any music. Is it okay if one musician replaces another but not if a non-musician does so? The person getting replaced didn't get paid either way.

Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?

max__dev 4 days ago||
Whether it is okay or not has no bearing on it being theft. And no need to create a hypothetical, this is exactly what happened to musicians.

There is broad evidence that labs have used substantial amounts of pirated data, no need to reach for a new definition of theft.

alex_smart 5 days ago||||
For intellectual property, copying without permission is theft.
aeon_ai 5 days ago||
Depending on the context, it’s fair use.

As it is, in this case.

malfist 5 days ago||
It's as fair use as this scenario:

I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.

aeon_ai 4 days ago||
No, it’s been decided by the courts to be fair use.

I’m not sure why people think they understand IP law better than the courts just because they don’t like the answer

nmeagent 4 days ago|||
> I’m not sure why people think they understand IP law better than the courts just because they don’t like the answer

Because judges are human beings and can be catastrophically wrong; e.g., see https://en.wikipedia.org/wiki/Dred_Scott_v._Sandford

malfist 4 days ago||||
Settled out of court does not mean the court has decided what is fair use in the case.
shagie 4 days ago||
https://www.goodwinlaw.com/en/insights/publications/2025/06/...

https://admin.bakerlaw.com/wp-content/uploads/2025/07/ECF-23...

> For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.

> The third fair use factor favors fair use for the purchased library copies converted from print to digital.

...

> This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.

nosyke 4 days ago|||
Ah yes because “the courts” have an unblemished historical record of never getting anything wrong
m4rtink 5 days ago|||
Well, given these companies are trying to sell it back to you - seems like even worse than stealing. ;-)
mitxela 5 days ago|||
Never ended, just changed in form.
not_a_bot_4sho 4 days ago|||
You're right that it never ended.

But it didn't change form much. Still around 50 million people enslaved nowadays.

midtake 4 days ago|||
Are you comparing modern workplace aches and gripes to literal 1800s slavery?
someguynamedq 3 days ago||
No need to balk. Different forms of coerced labor can all be bad even if some are worse than others.
cindyllm 3 days ago||
[dead]
bcjdjsndon 5 days ago|||
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
TacticalCoder 5 days ago|||
[flagged]
KingMob 5 days ago||
Nobody normal talks like this.

There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.

jasonmp85 4 days ago||
[dead]
xxs 5 days ago||
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft. Then you get the serfdom in the middle ages. Up to recent times humans have been brutally exploited.

The copy part was a recognized right, then taken away.

juvvel 5 days ago||
I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
Draiken 4 days ago|
It's the age old privatize the profits and socialize the losses.

People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...

I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?

sleight42 4 days ago|||
For the life of me, I can't understand why your comment is being downvoted or flagged or whatever makes it go gray on HN.

I'm guessing it's people knee-jerking that you're being political?

I can't understand the people who don't see it.

The data centers strain the power grids then electricity costs go up for everyone else. This is de facto a regressive tax because everyone needs electricity and the poor pay proportionately more of their income for the increased cost.

Live in San Francisco? Probably not now unless you're rich because of the skyrocketing cost of living due to Tech and AI money. Another de facto regressive tax, driving away other people.

Environmental damage? The poor are the most impacted and the least able to absorb the costs. Do they have the property or renter's insurance to protect them from these disasters? Another de facto regressive tax.

Need a new phone? Or a computer? Same problem.

Want to dabble in AI? You're not going to get too far on $20/month. It's mostly a wealthy person's game.

Or there are the statistics that the vast majority of successful founders from up upper middle class families or wealthier.

Wealth centralization is what our economic system does. The purpose of a system is what it does. If it wasn't, the system would have been changed.

juvvel 22 hours ago||
Exactly. That is also why I believe that megacorps should pay an "AI/infrastructure tax" of some sort. Want to use up electricity and water, buy up all the hardware, make workes obsolete? Pay up. Give back to the very society you exploit.
lofaszvanitt 4 days ago|||
People are like sheep. Plus those who work are preoccupied and are too tired to react to these changes.
TutleCpt 5 days ago||
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
vintagedave 5 days ago||
In my experience many execs know what they're talking about.

Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.

Knowledge alone is less often a factor.

Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.

Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.

ekunazanu 4 days ago|||
I think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed.

This sounds rather obvious, but I feel people forget it far too often.

TeMPOraL 4 days ago||
I call this the Hanlon's Handgun:

"Never attribute to stupidity that which can be adequately explained by systemic incentives promoting malice."

Previously: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

soraminazuki 4 days ago|||
You can see the same dynamic playing out in this very thread too.
Sohcahtoa82 4 days ago|||
Oh, I think MS execs (and all other execs, top-level politicians, and pundits for that matter) know what they're talking about, but they'll tell whatever lies they need to enrich and empower themselves.
someguynamedq 3 days ago||
Do you think msft market cap is just an accident, or...?
tom2026hn 5 days ago||
The problem isn’t just “stealing the fruits of human labor”, it’s also driving down the value of human skills and even taking away human jobs.
keeda 4 days ago||
I think you just described "technology" -- would you outlaw all technology then? ;-)
tstrimple 4 days ago|||
Right? I've been automating things and eliminating jobs using technology for over 20 years. No you do not need a team of five to manage your monthly reporting. We can build automated reports that are delivered automatically to everyone who needs them! No you shouldn't be managing all of this data in an excel file on a shared computer. I don't care if 90% of what Dave does during the week is maintain that file. Let's streamline that.

Suddenly it's very different when it's our jobs being automated away.

keeda 4 days ago||
Yes, I don't like it much either -- though I continue to be absolutely fascinated by all this -- but I realize this is unstoppable, and I must trust that on the balance technology has been extremely good for humanity and the source of all our progress, and so I must adapt to a new future.

But I also realize the impact of technology, good or bad, entirely depends on how society uses it, and that is where our focus must lie.

tom2026hn 4 days ago|||
I should add that we must also consider the extent and speed of replacement. If the replacement rate is too high, employment levels can’t recover, both scenarios need to be avoided.
someguynamedq 3 days ago||
Technological innovation can be sped up, but it cannot be slowed down without mechanisms that create worse problems than technology. Adaptation is a political problem.
someguynamedq 3 days ago|||
Every tool that is developed drives down the value of human skills. All sufficiently good tools diminish or take away human jobs.
ambicapter 5 days ago|||
One leads to the other so its simpler to point the root issue.
Trasmatta 5 days ago||
And the cruel irony is that they stole our work to train the AI that devalues our work going forward, and will cause many of us to lose our jobs.

I regret every line of open source code I ever wrote.

bluefirebrand 4 days ago|||
> I regret every line of open source code I ever wrote

And every stack overflow post, every reddit post, everything.

I regret participating in the open Internet.

Here I am anyways, I guess. It's just in my genes or something.

Trasmatta 4 days ago||
Same. I've scrubbed my presence as much as possible from the majority of the internet (except HN for whatever reason) to prevent future models from being trained on my output, but the damage is already done. Part of me exists in pretty much all AI models now, without my consent. And those models are actively stealing my career and passions.
tom2026hn 5 days ago|||
This process can even feel a bit like parasitism, it empties out the host, like in <Alien>.
Kuyawa 5 days ago||
Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours

So no, your cries for regulating others because you are losing the race won't work this time.

cute_boi 4 days ago||
It works if you can bribe politicians.
WarmWash 4 days ago||
How much money have people paid to use these services over the years?

None?

Ok, now you understand the business model.

Kuyawa 4 days ago||
I understand the business model, we are data providers, they are aggregators. That's fine.

What's not right is that they want to limit the use of such data when it's not theirs in first place. They just store it but that doesn't give them a license to prohibit the use by a third party since we all are owners of that data

WarmWash 4 days ago||
No, they are data sellers. You pay them in data. Just like money, it becomes theirs.
Waterluvian 5 days ago||
I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.
ohrus 5 days ago||
Yet we still tell students to buy textbooks. The individual must always pay. The corporation can do whatever the hell it wants.

The hypocrisy of this new world is already catching up to us.

leonidasrup 5 days ago|
In case of programming.

How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?

This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.

Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?

What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.

https://consortiuminfo.org/metalibrary/estimating-the-total-...

menaerus 5 days ago|
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.

IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.

leonidasrup 5 days ago||
LLMs can output near-exact segments of copyrighted code used for training.

https://arxiv.org/html/2408.02487v3

I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.

https://en.wikipedia.org/wiki/Shared_Source_Initiative

menaerus 5 days ago||
I guess you're not writing code much or haven't done much so as your professional career?
leonidasrup 4 days ago||
My argument is if AI companies are ignoring copyright law and looking at all training data as commons, then we should look at LLM output as something that is not protected by copyright law.

Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).

https://www.theguardian.com/technology/2026/may/23/trump-ai-...

https://www.bbc.com/news/articles/c98r8r7dz5no

https://en.wikipedia.org/wiki/Commons

menaerus 4 days ago||
I wanted to understand your background first because what you initially said is a very oversimplified view of LLM mechanics, and generally not quite the way how software is in practice written. Since you didn't answer that question, I will assume that you're not a SWE by a call. To give you an example of what I am trying to convey is: imagine a data-intensive workload hitting your storage/database/kernel implementation, and it's painfully slow, your customers are not happy. Then you as engineer sit down, spend days profiling and understanding the code, researching about existing algorithmic solutions to the same or similar issues found in the wild, you read some open-source implementations of viable approaches, you ditch some, some you take, you also read books, articles, other peoples experiences etc. And finally you end up, let's put it bluntly, with some sharded data structure by which you solve the bottleneck. It's not novel, the technique is so common and is already implemented across many many different products in slightly different flavors so I am wondering why do you think this is not a copyright breach but the LLM, which does more or less the same thing, is?
More comments...