18 ms·
It already applies to real people, doesn't it? I.e. if you read a book, you're not allowed to start printing and selling copies of that book without permission
by TheSoftwareGuy 2y ago
It already applies to real people, doesn't it? I.e. if you read a book, you're not allowed to start printing and selling copies of that book without permission of the copyright owner, but if you learn something from that book you can use that knowledge, just like a model could.
- m1el 2y agowhen it comes to real people, they get sued into oblivion for downloading copyrighted content, even for the purpose of learning. but when facebook & openai do it, at a much larger scale, suddenly the laws must be changed.
- ryoshu 2y agoCase in point - https://en.wikipedia.org/wiki/Aaron_Swartz https://en.wikipedia.org/wiki/Aaron_Swartz
- JumpCrisscross 2y agoSwartz wasn’t “downloading copyrighted content…for the purpose of learning,” he was downloading with the intent to distribute. That doesn’t justify how he was treated. But it’s not analogous to the limited argument for LLMs that don’t regurgitate the copyrighted content.
- Terretta 2y ago> when it comes to real people, they get sued into oblivion for downloading copyrighted content, even for the purpose of learning. Really? Or do they get sued for sharing as in republishing without transformation? Arguably a URL providing copyrighted content, is you offering a xerox machine. It seems most "sued into oblivion" are the reshare problem, not the get one for myself problem.
- conjectures 2y agoIt does apply to people? When you read a copy of a book, you can't be sued for making a copy of the book in the synapses of your brain. Now, if you have eidetic memory and write out large chunks of the book from memory and publish them, that's what you could be sued for.
- triceratops 2y ago> When you read a copy of a book They're not talking about reading a book FFS. You absolutely can be sued for illegally obtaining a copy of the book.
- deleted 2y ago[deleted]
- tsimionescu 2y agoThis is not about memory or training. The LLM training process is not being run on books streamed directly off the internet or from real-time footage of a book. What these companies are doing is: 1. Obtain a free copy of a work in some way. 2. Store this copy in a format that's amenable to training. 3. Train their models on the stored copy, months or years after step 1 happened. The illegal part happens in steps 1 and/or 2. Step 3 is perhaps debatable - maybe it's fair to argue that the model is learning in the same sense as a human reading a book, so the model is perhaps not illegally created. But the training set that the company is storing is full of illegally obtained or at least illegally copied works. What they're doing before the training step is exactly like building a library by going with a portable copier into bookshops and creating copies of every book in that bookshop.
- visarga 2y agoBut making copies for yourself, without distributing them, is different than making copies for others. Google is downloading copyrighted content from everywhere online, but they don't redistribute their scraped content. Even web browsing implies making copies of copyrighted pages, we can't tell the copyright status of a page without loading it, at which point a copy has been made in memory.
- tsimionescu 2y agoMaking copies of an original you don't own/didn't obtain legally is not fair use. Also, this type of personal copying doesn't apply to corporations making copies to be distributed among their employees (it might apply to a company making a copy for archival, though).
- triceratops 2y agoCan I download a book without paying for it, and print copies of it? Stash copies in my bathroom, the gym, my office, my bedroom etc. to basically have a copy on hand to study from whenever I have some free time? What about movies and music?
- ajross 2y ago> Can I download a book without paying for it, and print copies of it? No, but you can read a book, learn its contents, and then write and publish your own book to teach the information to others. The operation of an AI is rather closer to that than it is to copyright violation. "Should" there be protections against AI training? Maybe! But copyright law as it stands is woefully inadequate to the task, and IMHO a lot of people aren't really treating with this. We need a functioning government to write well-considered laws for the benefit of all here. We'll see what we get.
- zombiwoof 2y agoIf you buy it
- ajross 2y agoNo, even if I steal it. I can teach you anything I know. Congress shall make no law abridging the freedom of speech, as it were.
- tsimionescu 2y agoYes, but this is not the right model. What OpenAI wants is to borrow a book, make a copy of it, and keep using that copy, in training their models. This is fully and simply illegal, under any basic copyright law.
- triceratops 2y agoBut I can't legally obtain the book to read and learn from without me (or a library) paying for it. Let's start there first.
- echelon 2y agoIf models can learn for free, then the models (training code, inference code, training data, weights) should also be free. No copyright for anybody. And if you sell the outputs of your model that you trained on free content, you shouldn't be able to hide behind trade secret.
- crorella 2y ago> just like a model could It is not remotely the same, the companies training the models are stealing the content from the internet and then profiting from it when they charge for the use of those models.
- Terretta 2y ago> the companies training the models are stealing the content from the internet Are you stealing a billboard when you see and remember it? The notion that consuming the web is "stealing" needs to stop.
- crorella 2y agoWe are not taking about billboards here, we are talking about copyrighted works, like books. If you want to do mental gymnastics and call "consuming" the web the act of downloading books without paying for them, then go ahead, but don't pretend the rest will buy your delusion.
- Terretta 2y agoOn the contrary, even telling people which billboards are posted about what, and how to get to them to look at them, is "how it works". But the courts will get to clarify (in today's news): https://www.reuters.com/legal/news-corp-sued-by-brave-software-google-search-engine-rival-2025-03-13/ https://www.reuters.com/legal/news-corp-sued-by-brave-softwa...
- llamaimperative 2y agoThe question is whether it destroys the incentive to produce the work. That is the entire point of copyright and patent law. LLMs do indeed significantly reduce the incentive to produce original work.
- codedokode 2y agoAre you stealing when using a pirated software to run a billion-dollar business?
- 2y ago
- simion314 2y ago>you can use that knowledge, Did OpenAI bought one copy of each book, or did they legaly borowed athe books and documents ? if you copy paste rom books and claim is your content you are plagiarizing. LLMs were provent to copy paste trained content so now what? Should only big Tech be excluded from plagiarizing ?
- pier25 2y ago> just like a model could Not really. You can't multiply yourself a million times to produce content at an industrial scale.
- alabastervlog 2y agoThis is why I think my array of hard drives full of movies isn't piracy. My server just learned about those movies and can tell me about them, is all. Just like a person!
- tsimionescu 2y agoIt doesn't, a real person can't legally obtain a copy of a copyrighted work without paying the copyright holder for it. This is what OpenAI is asking for: they don't want to pay for a single copy of a single book, and still they want to train their models on every single book in history (and song, and movie, and painting, and code base, and anything else they can get their hands on).
- bee_rider 2y agoThese AI models are just obviously new things. They aren’t people, so any analogy about learning from the training material and selling your new skills is off base. On the other hand, they aren’t just a copy of the training content, and whether the process that creates the weights is sufficiently transformative as to create a new work is… what’s up for debate, right? Anyway I wish people would stop making these analogies. There isn’t a law covering AI models yet. It is a big industry at this point, and the lack of clarity seems like something we’d expect everybody (legislators and industry) to want to rectify.
- amelius 2y agoTotally agree. Except the current administration probably will interpret things the way they see fit ...
- codedokode 2y agoModel cannot "learn" because it is not a human. What happens is a human obtains "a free copy" of a copyrighted work, processes it using a machine and sells the result.
- bee_rider 2y ago> Model cannot "learn" because it is not a human. Sure, that’s why don’t like the analogy. > What happens is a human obtains "a free copy" of a copyrighted work, processes it using a machine and sells the result. Right, so for example it is pretty common to snip up small bits of songs and to use in other songs (sampling). Maybe that could be an example of somewhere to start? But, these ML models seem quite different, I guess because the “samples” are much smaller and usually not individually identifiable. And really the model encodes information about trends in the sources… I dunno. I still think we need a new law.
- aiono 2y agoCan I pirate books to train myself?
- amelius 2y agoDo you know Numerical Recipes in C? This discussion reminds me of it.
- sidewndr46 2y agoAnd when I "learn" a verbatim copy of pages of that book, then write those pages out in Microsoft Word & sell those pages its legal?