7 ms·
You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples
by Orygin 2mo ago
You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples used in their tracks.
- echoangle 2mo agoThere’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.
- inigyou 2mo agoSo now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.
- hamandcheese 2mo agoIs Claude's paraphrasing of Lord of the Rings equivalent to the original?
- echoangle 2mo agoIf you just use the abstract idea, you could have done the same thing yourself all the time already.
- pj_mukh 2mo ago"I get a GPL version of Microsoft Office." Is this not..Libre?
- inigyou 2mo agoLibreOffice is a different product from Microsoft Office.
- dec0dedab0de 2mo agobut it's the same idea
- inigyou 2mo agoLibreOffice is not Microsoft Office run through a copyright laundering machine.
- bnj 2mo agoThat’s what a brain is
- Orygin 2mo agoMaybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet. If yours do that, maybe we can redefine what copyrighting and patenting means for humans
- inigyou 2mo agoWhy do companies bother with the Chinese wall technique, then?
- satvikpendem 2mo agoYes exactly. That's basically what an emulator for a games console is for example, a reimplementation of the original.
- inigyou 2mo agoSo a console game loses its copyright if you emulate it?
- echoangle 2mo agoThe game doesn’t but the emulator doesn’t necessarily infringe copyright in itself just because it is based on the original console.
- inigyou 2mo agoSo what does that have to do with a copyright laundering machine?
- satvikpendem 2mo agoA copyright laundering machine has no copyright itself and anything it ingests does not need to concern itself with its own copyright because the output has none anyway.
- satvikpendem 2mo agoA game is a specific work to be copied so no, but the system it runs on can still be without copyright.
- tripzilch 2mo agoThere's also a difference between an MP3 and a FLAC. Again, ask DJs how well they're getting away on that distinction.
- echoangle 2mo agoThat’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.
- tripzilch 2mo agothe "codec" is not really the point. playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference. similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference. the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for). similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books. the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts. to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more. having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.