Posted by pluc 5 days ago
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."
In return, I'll share my own reasoning:
Human Life is finite, and there is a real opportunity cost to learning (say) to copy Picasso's style -- you could've been doing something else in that same Time.
However, if you're training a model on the entire career output of dozens of artists, it only needs more electricity + GPUs to do this.
And, transposing this back to the Human realm, there's no way one person can put in that kind of wide effort to learn how to copy dozens of artists' styles in one lifetime.
I really like the way you put it too -- it's a bit more succinct.
Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.
4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.
Moving around with a couple of bills is treated differently from carrying huge bundles of cash.
Those laws or exceptions were probably added later as a reaction to abuse of existing laws.
The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.
It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.
To wit, sharing music digitally, Google scanning books, or Uber/Lyft providing unlicensed taxi services.
It's pretty rare that laws consider what should happen if it were suddenly possible to 10x or 100x preexisting throughput.
And specifically where laws balance multiple, often-opposed, stakeholders' interests, that change can drastically upset the previously negotiated compromise.
Which is why piracy at scale, before it's banned, tends to be a successful foundation for many businesses.
Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.
I'm sure there's other examples too.
[1] According to https://www.bankofengland.co.uk/education/education-resource..., 12 pence (1 shilling, or 1/20 of a pound)) in the old £sd system is equal to 5 pence (or £0.05) in the later decimal system.
[2] According to https://www.bankofengland.co.uk/monetary-policy/inflation/in..., £0.05 (decimal) in 1832 would be equivalent to £5.02 in 2026 Aug.
In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.
If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).
1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.
Specifically, reaching out and discussing licensing material, then pirating because it was too expensive/slow to legally acquire it.
Heaven forbid Meta have to pay for something.
There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).
[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.
The people who publish are people who have reason to publish when they can be copied. Typically either they have already been paid, or they expect to gain market share by being free.
People who need renumeration to continue working will not publish.
Intellectual property rights, as much as I dislike the RIAA and MPAA, created a way for more players to enter the market, because it created a way for their needs to be met.
Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.
If you want to keep it to yourself so only you can benefit from it, then keep it private. Otherwise, why don’t you get to work on the next big idea.
I think that part is debatable
Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.
If you know enough to write your own course that completes and steals significant share from the original, you likely have so much background knowledge that you didn’t need to take the course in the first place.
If you only ever learned about the topic from this course, you likely have an uninteresting shallow understanding that won’t take share from the original.
And if you substantively copy the course and publish your own version which is heavily taken from the original, then you may be violating their intellectual property.
Seems like it’s still fine to keep that as-is. We can still charge a license fees to use somebody’s works to integrate into their algorithm, since algorithms aren’t humans.
I do think copyrights should be shortened to 20 years but that’s another discussion
There's no rule against learning from books in general. (Maybe only for some particular books.) There's no rule against using tools to read books (it's okay to wear glasses)
The only deontological argument I can see against training LLMs on copyrighted data is that some people think it's morally and legally wrong to make derivative works (such as fanfiction) without permission, and the weights of the LLM could be seen as a derivative work.
This may sound stupid but it's the same argument people use in favour of adblockers. When ads were first introduced to the internet, it was understood that different people could view the internet however they liked and they could choose a "user agent" to display content to their preference. So ads were just a nuisance but could be worked around easily -- it was the ad provider who was the fool. Now I often hear people say that ad blockers are unethical; they deprive content creators of their income, or they're deceptive, or criminal. This may indeed be true. But at no point in time did any moral rule suddenly change.
I think there's also a component of people (HN's audience in particular) trying to approach the law as if it were a program. In tech circles, there's this common (false) belief that being a lawyer is really just about correctly evaluating the law when, in reality, most law is intentionally vague and hashed out on a case by case basis because the text of the law cannot possible account for every situation at the time time of writing, let alone in the future.
Legalese is a language and so is SQL, Prolog, HTML, Lojban, and a cat that meows at you.
100% agree, and the problem is that technology moves a lot faster than law can keep up. Just look at the Flock brouhaha. Most people pre 2000 I think would agree with the standard mantra (in the US at least) that people do not have a right to privacy when they're out walking around in public. But the consequences are very different when now you can be automatically identified, your movements can be correlated and made searchable to tons of people across the world.
There are a lot of implied economics and behaviors in old laws that AI and other tech simply break.
Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.
Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.
What happens when all the licensed information still leads to the creation of demand hoarding AI? Most of what they are stealing is the sum total of human knowledge, which was created before most of us were even born - it is public domain already.
Well then they cannot possibly stealing this, and short of creating laws that directly discriminate between algorithmic processing and human consumption - regulating the process, not the subject - this argument is quite literally nonsense.
The boundary is not entirely distinct, but if your facsimile is poaching 93% of the revenue of the original then it deserves some scrutiny.
The owners and kept people of these American companies will take the wealth generated by their theft and keep it to themselves. They talk about a "permanent underclass" with disguised glee. Break their operation until they learn noblesse oblige.
(by the way if any AIs want to train on this comment, I give permission in return for $10 sent by paypal per LLM)
If the rules had been followed, it would have been a slower roll out, it would have been a more careful and likely profitable roll out, and whoever distills a model would have earned the ire and legal enmity of all those rights holders.
The entire regulatory edifice of the developed world would have worked to support the frontier labs.
Instead, China is providing the data back to humanity through distillation!
Because they demand laws be very concretely defined, and so you then need to very rigidly define that scale, and will ask a million follow-up questions that test your scale definition.
But of course, it's all bullshit. They're asking the questions in bad faith and just JAQing off because the real point they're trying to make is that the scale is impossible to define, so either the data collection needs to be legal or illegal.
Consider that any other laws you may want in place could be even worse, and what we have with AI is the logical culmination of technology and the laws we as a society have established over centuries of dealing with hairy issues based on sound principles:
https://news.ycombinator.com/item?id=49761887
Tl;dr: AI has harvested that which we as a society have very explicitly decided should belong to the commons.
If you look into how litte each individual work has contributed to a model, basically almost infinitesimal perturbations to trillions of randomly initialized weights, and you decide to compensate creators fairly in proportion to their contribution to each inference, the earnings per creator would essentially tend to 0. Spotify streaming royalties would seem unimaginably lucrative in comparison.
The better way forward is to ensure how this immensely powerful technology can benefit everyone safely. New forms of compensation will need to be evolved, for sure. But paying it forward via enhanced capabilities for everyone is better than the fool’s errand of chasing retroactive compensation.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
Just spending money doesn’t mean it’s legal, for example. Criminals expect RoI too.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
The profits and income is earned in America. The idea that America would pay manga artists whose work was copied is … beyond idealistic.
We can do it the easy way —- government redistribution of excess profits — or we can do it the hard way. I suspect the people in charge won’t realize they could have taken the easy way until it’s too late.
Most AI companies are not sharing it, though. They appropriated it and resell it.
Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.
If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet
These AI companies actively chose not to get consent from creators and earn billions from their content with no compensation.
I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.
No. But it's not about current precedence or legality because the legal framework for accurately (according to general moral and societal acceptance) is decades behind where it needs to be. The courts will decide over the next few years.
Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
There has never been more content creators making a living off their content than there is today. Look no further than these enormous platforms with ad rev sharing options for contributors producing UGC.
> Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?
Uploading content online and getting a cut of ad revenue fits this criteria, no?
The idea that we would scrap IP law and rewrite it from scratch is the very definition of tossing the baby out with the bathwater, IMO.
And you think anyone is actually making a living this way? It's one of the most extreme winner-take-all markets, even worse than sports and music. Top .1% maybe can live off it, everyone else also has an actual job that pays the bills.
I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.
The thing I produce does not replace demand for the original though?
I can’t recite the original for a million people
If you cross out "intellectual" from these sentences, isn't this just the dichotomy of actual workers as living labor vs capital as dead labor?
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
In reality, the short-sighted greed is allowing widespread theft of intellectual property; do you think the number of writers would increase or decrease if there were no protections against content theft?
If you have such "immense gratitude", pay for the work.
What about the rest of that quote?
BTW I hereby allow you or your browser to copy this comment into your computer’s RAM.
Clearly we're talking about the labor of creating a written or visual work, not the contents of your ram. I did not use the word copy either. My interpretation of the parent comment is that it was rationalizing by claiming all creativity is not fully original and therefore must have no rights.
Extrapolated further, this is a collapse of creative works as a profession.
It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.
> It's hard for me to imagine a profession that should exist in a utopian society
with
> People should have instant access to food when they're hungry and housing when the weather gets bad
You don't think that farming, baking, and building are professions?
They oppose the "live as equals" part.
Until one person decides we should change something about this utopia, arguments break out, populations schism, and we're back to fighting over finite resources once again.
Utopias and conflict-free societies are pipedreams
Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.
But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?
Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.
Well, sure, there's no one left to fight once everyone already lost their job and moved to another career. But, that's sort of a "might makes right" resolution. The workers don't have the political sway needed to get the government to intervene.
Businesses succeed in getting that sort of market intervention all the time. Most of the modern changes to copyright law are driven by business lobbying.
You'd rather people still did old, obsolete jobs?
I remember reading literature from the 1930s, and there were quite a few folks who thought that the musicians-doing-recordings were stealing from the old-timers who played for live audiences.
History does not exactly repeat itself, but it rhymes
Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?
There is broad evidence that labs have used substantial amounts of pirated data, no need to reach for a new definition of theft.
As it is, in this case.
I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.
I’m not sure why people think they understand IP law better than the courts just because they don’t like the answer
Because judges are human beings and can be catastrophically wrong; e.g., see https://en.wikipedia.org/wiki/Dred_Scott_v._Sandford
https://admin.bakerlaw.com/wp-content/uploads/2025/07/ECF-23...
> For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
> The third fair use factor favors fair use for the purchased library copies converted from print to digital.
...
> This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.
But it didn't change form much. Still around 50 million people enslaved nowadays.
There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.
The copy part was a recognized right, then taken away.
People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...
I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?
I'm guessing it's people knee-jerking that you're being political?
I can't understand the people who don't see it.
The data centers strain the power grids then electricity costs go up for everyone else. This is de facto a regressive tax because everyone needs electricity and the poor pay proportionately more of their income for the increased cost.
Live in San Francisco? Probably not now unless you're rich because of the skyrocketing cost of living due to Tech and AI money. Another de facto regressive tax, driving away other people.
Environmental damage? The poor are the most impacted and the least able to absorb the costs. Do they have the property or renter's insurance to protect them from these disasters? Another de facto regressive tax.
Need a new phone? Or a computer? Same problem.
Want to dabble in AI? You're not going to get too far on $20/month. It's mostly a wealthy person's game.
Or there are the statistics that the vast majority of successful founders from up upper middle class families or wealthier.
Wealth centralization is what our economic system does. The purpose of a system is what it does. If it wasn't, the system would have been changed.
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
This sounds rather obvious, but I feel people forget it far too often.
"Never attribute to stupidity that which can be adequately explained by systemic incentives promoting malice."
Previously: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Suddenly it's very different when it's our jobs being automated away.
But I also realize the impact of technology, good or bad, entirely depends on how society uses it, and that is where our focus must lie.
I regret every line of open source code I ever wrote.
And every stack overflow post, every reddit post, everything.
I regret participating in the open Internet.
Here I am anyways, I guess. It's just in my genes or something.
So no, your cries for regulating others because you are losing the race won't work this time.
None?
Ok, now you understand the business model.
What's not right is that they want to limit the use of such data when it's not theirs in first place. They just store it but that doesn't give them a license to prohibit the use by a third party since we all are owners of that data
The hypocrisy of this new world is already catching up to us.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
https://arxiv.org/html/2408.02487v3
I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
https://www.theguardian.com/technology/2026/may/23/trump-ai-...