6 ms·
> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make
by skilled 8mo ago
> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models.
Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?
- tobwen 8mo agoBooks are databases, chars their elements. We have copyright for databases in EU :)
- Bombthecat 8mo agoWho cares? Only Disney had the money to fight them. Everything else will be slurped up for and with AI and be reused.
- RGamma 8mo agoThe chicken is trying to become the egg.
- general1465 8mo agoDid you pirated this movie? No I did not, it is fair use because this movie is nothing more than a statistical correlation to my dopamine production.
- earthnail 8mo agoThe movie played on my screen but I may or may not have seen the results of the pixels flashing. As such, we can only state with certainty that the movie triggered the TV's LEDs relative to its statistical light properties.
- JKCalhoun 8mo agoI saw the movie, but I don't remember it now.
- machomaster 8mo agoI saw the movie, but I did not watch it.
- Ferret7446 8mo agoIndeed, the "copy" of the movie in your brain is not illegal. It would be rather troublesome and dystopian if it were.
- visarga 8mo agoThe problem is when you use your "copy" as inspiration and actually create and publish something. It is very hard to be certain you are safe, besides literal expression close paraphrasing is also infringing, using world building elements, or using any original abstraction (AFC test). You can only know after a lawsuit. It is impossible to tell how much AI any creator used secretly, so now all works are under suspicion. If copyright maximalists successfully copyright style (vibes), then creativity will be threatened. If they don't succeed, then copyright protection will be meaningless. A catch 22.
- HWR_14 8mo ago> close paraphrasing is also infringing, using world building elements, or using any original abstraction (AFC test) World building elements? Do you have more details on that, because that feels wrong to me. Unless you mean the specific names of things in the world like "Hobbits".
- SoftTalker 8mo agoNot yet, anyway.
- thaumasiotes 8mo agoNote that what copyright law prohibits is the action of producing a copy for someone else, not the action of obtaining a copy for yourself.
- bulbar 8mo agoTraining of an LLM however is a lossy compressing algorithm to provide a copy of a variant of the data to the user later on.
- codedokode 8mo agoIf I am not mistaken, the law prohibits producing any unauthorized copies. So if you download a pirated book on a computer, you produce an illegal copy: [1]. If I am not missing anything, ML companies are galaxy-scale infringers. > 106. Exclusive rights in copyrighted works > Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following: > (1) to reproduce the copyrighted work in copies or phonorecords; > 501. Infringement of copyright > (a) Anyone who violates any of the exclusive rights of the copyright owner as provided by sections 106 through 122 or of the author as provided in section 106A(a), or who imports copies or phonorecords into the United States in violation of section 602, is an infringer of the copyright or right of the author, as the case may be. [1] https://www.copyright.gov/title17/92chap5.html https://www.copyright.gov/title17/92chap5.html
- gruez 8mo ago>Did you pirated this movie? No I did not, [...] You're probably being sarcastic but that's actually how the law works. You'll note that when people get sued for "pirating" movies, it's almost always because they were caught seeding a torrent, not for the act of watching an illegal copy. Movie studios don't go after visitors of illegal streaming sites, for instance.
- aucisson_masque 8mo ago> Movie studios don't go after visitors of illegal streaming sites, for instance. They absolutely do, in France we have Hadopi that tracks torrent leecher. Hadopi had been heavily pushed by the movie and music industry.
- gruez 8mo ago>They absolutely do, in France we have Hadopi that tracks torrent leecher You're still uploading even if you don't let it finish and go to "seeding".
- bmitc 8mo agoIt's how the law works for those at the top of the oligarchy.
- ErroneousBosh 8mo agoDid you pirate this movie? No, I acquired a block of high-entropy random numbers as a standard reference sample.
- Elfener 8mo agoIt seems so, stealing copyrighted content is only illegal if you do it to read it or allow others to read it. Stealing it to create slop is legal. (The difference, is that the first use allows ordinary poeple to get smarter, while the second use allows rich people to get (seemingly) richer, a much more important thing)
- threethirtytwo 8mo agoIt does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can definitely see these laws evolving into things that are human centric. It’s permissible for a human to do something but not for an AI. What is consistent is that obtaining the books was probably illegal, but say if nvidia bought one kindle copy of each book from Amazon and scraped everything for training then that falls into the grey zone.
- ckastner 8mo ago> To call training illegal is similar to calling reading a book and remembering it illegal. Perhaps, but reproducing the book from this memory could very well be illegal. And these models are all about production.
- roblabla 8mo agoTo be fair, that seems to be where some of the IA lawsuits are going. The argument goes that the models themselves aren't derivative works, but the output they produce can absolutely be - in much the same way that reproducing a book from memory could be copyright violation, trademark infringement, or generally go afoul of the various IP laws.
- threethirtytwo 8mo agoModels don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.
- empath75 8mo agoThey do memorize some books. You can test this trivially by asking ChatGPT to produce the first chapter of something in the public domain -- for example a Tale of Two Cities. It may not be word for word exact, but it'll be very close. These academics were able to get multiple LLMs to produce large amounts of text from Harry Potter: https://arxiv.org/abs/2601.02671 https://arxiv.org/abs/2601.02671
- nancyminusone 8mo agoWhen you're responsible for 4% of the global GDP, they let you do it.
- qingcharles 8mo agoThey let you just grab any book you want.
- NitpickLawyer 8mo ago> Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor? It makes some sense, yeah. There's also precedent, in google scanning massive amounts of books, but not reproducing them. Most of our current copyright laws deal with reproductions. That's a no-no. It gets murky on the rest. Nvda's argument here is that they're not reproducing the works, they're not providing the works for other people, they're "scanning the books and computing some statistics over the entire set". Kinda similar to Google. Kinda not. I don't see how they get around "procuring them" from 3rd party dubious sources, but oh well. The only certain thing is that our current laws didn't cover this, and probably now it's too late.
- deleted 8mo ago[deleted]
- masfuerte 8mo agoScanning books is literally reproducing them. Copying books from Anna's Archive is also literally reproducing them. The idea that it is only copyright infringement if you engage in further reproduction is just wrong. As a consumer you are unlikely to be targeted for such "end-user" infringement, but that doesn't mean it's not infringement.
- Ferret7446 8mo agoPrivate reproductions are allowed (e.g. backups). Distributing them non-privately is not.
- masfuerte 8mo agoBackups are permitted (and not for all media) when you legally acquired the source. Scanning a physical book is not a permitted backup, and neither is downloading a book from Anna's archive.
- fc417fc802 8mo ago> Scanning a physical book is not a permitted backup On what basis do you claim that? You're also missing critical legal context. When a would be consumer downloads pirated media in lieu of purchasing it he damages the would be seller. When my automated web scraper inadvertently archives some pirated content on my local disk no one is financially harmed. The question is where the boundary between those things lies.
- ThrowawayR2 8mo agoYes, it's been discussed many times before. All the corporations training LLMs have to have done a legal analysis and concluded that it's defensible. Even one of the white papers commissioned by the FSF ( "Copyright Implications of the Use of Code Repositories to Train a Machine Learning Model" at https://www.fsf.org/licensing/copilot/copyright-implications-of-the-use-of-code-repositories-to-train-a-machine-learning-model https://www.fsf.org/licensing/copilot/copyright-implications... ), concluded that using copyrighted data to train AI was plausibly legally defensible and outlined the potential argument. You will notice that the FSF has not rushed out to file copyright infringement suits even though they probably have more reason to oppose LLMs trained on FOSS code than anyone else in the world.
- jkaplowitz 8mo ago> Even one of the white papers commissioned by the FSF Quoting the text which the FSF put at the top of that page: "This paper is published as part of our call for community whitepapers on Copilot. The papers contain opinions with which the FSF may or may not agree, and any views expressed by the authors do not necessarily represent the Free Software Foundation. They were selected because we thought they advanced the discussion of important questions, and did so clearly." So, they asked the community to share thoughts on this topic, and they're publishing interesting viewpoints that clearly advance the discussion, whether or not they end up agreeing with them. I do acknowledge that they paid $500 for each paper they published, which gives some validity to your use of the verb "commissioned", but that's a separate question from whether the FSF agrees with the conclusions. They certainly didn't choose a specific author or set of authors to write a paper on a specific topic before the paper was written, which a commission usually involves, and even then the commissioning organization doesn't always agree with the paper's conclusion unless the commission isn't considered done until the paper is updated to match the desired conclusion. > You will notice that the FSF has not rushed out to file copyright infringement suits even though they probably have more reason to oppose LLMs trained on FOSS code than anyone else in the world. This would be consistent with them agreeing with this paper's conclusion, sure. But that's not the only possibility it's consistent with. It could alternatively be because they discovered or reasonably should have discovered the copyright infringement less than three years ago, therefore still have time remaining in their statute of limitations, and are taking their time to make sure they file the best possible legal complaint in the most favorable available venue. Or it could simply be because they don't think they can afford the legal and PR fight that would likely result.
- postexitus 8mo agoA quite good explanation of what copyright laws cover and should (and should not) cover is here by Cory Doctorow: https://www.theguardian.com/us-news/ng-interactive/2026/jan/18/tech-ai-bubble-burst-reverse-centaur https://www.theguardian.com/us-news/ng-interactive/2026/jan/...
- HillRat 8mo agoIt's not settled law as it pertains to LLMs, but, yes, creating a "statistical summary" of a book (consider, e.g., a concordance of Joyce's "Ulysses") is generally protected as fair use. However, illegally accessing pirated books to create that concordance is still illegal.
- HWR_14 8mo agoCopyright laws are so undefined and NVIDIAs lawyers so plentiful that the statement works in their favor. You're allowed to copy part of a work in many cases, the easiest example is you can quote a line from a book in a review. The line is fuzzy.
- bulbar 8mo agoOf course it does not make sense, it's just the framing of a multi billion dollar industry and people tend to buy those.
- lencastre 8mo agoI would love to see these nvidia designs as mere statistical correlations of graphic card design.