8 ms·
Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or
by 3PS 1y ago
Broadly summarizing.
This is OK and fair use: Training LLMs on copyrighted work, since it's transformative.
This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training.
Overall seems like a pretty reasonable ruling?
- simmerup 1y agoDepends whether you actually agree its transformative
- lesuorac 1y agoFor textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.
- johnnyanmac 1y agoThe nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.
- ticulatedspline 1y agoFor closed models the storage problem isn't really a problem, they can be judged by what they produce not how they store it as you don't have access to the actual data. That said, open weight LLMs are probably screwed, if enough of the work remains in the weights such that they can be extracted (even if it's without even talking to the LLM) then the weight file itself represents a copy of the work that's being distributed. So enjoy these competent run-at-home models while you can, they're on track for extinction.
- ninetyninenine 1y agoWhy doesn’t this apply to humans? If I memorize something such that it can be extracted did I violate the law? It’s only if I choose to allow such extraction to occur then I’m in violation of the law right? So if I or an LLM simply doesn’t allow said extraction to occur, memorization and copying is not against the law.
- ranger_danger 1y agoI think an important distinction here is distribution... did you tell someone else what you memorized? Is downloading a model akin to distributing that same information?
- ninetyninenine 1y agoWhat if I don't download the model and I just communicate with it. Sort of like chatting with another human. That's not a copyright issue right? I mean that's how most LLMs are deployed today.
- ranger_danger 1y agoMy understanding is that it depends on a judge/jury's subjective opinion on how similar the output is to something copyrightable. Perhaps intent may play a role as well.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- ranger_danger 1y agoI wonder if https://en.wikipedia.org/wiki/Illegal_number https://en.wikipedia.org/wiki/Illegal_number comes into play here.
- sidewndr46 1y agoWasn't that just over an arrangement of someone else's photographs?
- lesuorac 1y agohttps://en.wikipedia.org/wiki/Andy_Warhol_Foundation_for_the_Visual_Arts,_Inc._v._Goldsmith https://en.wikipedia.org/wiki/Andy_Warhol_Foundation_for_the... I wouldn't call it that. Goldsmith took a photograph of Prince which Warhol used as a reference to generate an illustration. Vanity Fair then chose to buy a license Warhol's print instead of Goldsmith's photograph. So, despite the artwork being visual transformative (silkscreen vs photograph) the actual use was not transformed.
- thedevilslawyer 1y agoWhat's the steelman case that is transformative? Because prima-facie, it seems to only output original output - "intelligent" output.
- derbOac 1y agoBut those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compared to what it is now, but to me the logic of this ruling is illogical other than "it's ok for a corporation to use lots of works without permission but not for an individual to use a single work without permission." Maybe if they suddenly loosened copyright enforcement for everyone I might feel differently. "Kill one man, and you are a murderer. Kill millions of men, and you are a conqueror." (An admittedly hyperbolic comparison, but similar idea.)
- rcxdude 1y ago>If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I think that's the conclusion of the judge. If Anthropic were to buy the books and train on them, without extra permission from the authors, it would be fair use, much like if you were to be inspired by it (though in that case, it may not even count as a derivative work at all, if the relationship is sufficiently loose). But that doesn't mean they are free to pirate it either, so they are likely to be liable for that (exactly how that interpretation works with copyright law I'm not entirely sure: I know in some places that downloading stuff is less of a problem than distributing it to others because the latter is the main thing that copyright is concerned with. And AFAIK most companies doing large model training are maintaining that fair use also extends to them gathering the data in the first place). (Fair use isn't just for discussion. It covers a broad range of potential use cases, and they're not enumerated precisely in copyright law AFAIK, there's a complicated range of case law that forms the guidelines for it)
- altruios 1y agowhich AFAIN IANAL, copyright and exhaustive rights are completely different. Under copyright, once a book is purchased: that's it. Reselling the same, or transformed (re: highlighted) worked 'used' is 100% legal, as is consuming it at your discretion (in your mind {a billion times}, a fire, or (yes even) what amounts to a fancy calculator). (that's all to say copyright is dated and needs an overhaul) But that's taking a viewpoint of 'training a personal AI in your home', which isn't something that actually happens... The issue has never been the training data itself. Training an AI and 'looking at data and optimizing a (human understanding/AI understanding) function over it' are categorically the same, even if mechanically/biologically they are very different.
- ticulatedspline 1y agoDefinitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly would be legal to do so but only if they had legit copies in the training room. Pirating the movies would be illegal. Though the downside is it does create a training moat. If you want to create the super-brain AI that's conversant on the corpus of copyrighted human literature you're going to need a training library worth millions
- johnnyanmac 1y agoThat's a part of the issue. I'm not sure if this has happened in visual arts, but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. I see elements of that here. Buying copyrighted works not to be exposed and be inspired, nor to utilize the aithor's talents, but to fuel a commercialization of sound-a-likes.
- lesuorac 1y ago> You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet" Keep in mind, the Authors in the lawsuit are not claiming the _output_ is copyright infringement so Alsup isn't deciding that.
- Dracophoenix 1y ago> but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. You're referencing Midler v Ford Motor Co in the 9th circuit. This case largely applies to California, not the whole nation. Even then, it would take one Supreme Court case to overturn it.
- doctorpangloss 1y agoIt’s similar to the Google Books ruling, which Google lost. Anthropic also lost. TechCrunch and others are very aspirational here.
- philipkglass 1y agoDo you mean Authors Guild, Inc. v. Google, Inc.? Google won that case: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... Maybe there's another big Google Books lawsuit that Google ultimately lost, but I don't know which one you mean in that case.
- doctorpangloss 1y agosee, but if you ask a copyright attorney: Google lost. This is what I mean by aspirational. They won something, in very similar circumstances to Anthropic, "fair use," but everything else that made what they were doing a practical reality instead of purely theoretical required negotiation with Authors Guild, and indeed, they are not doing what they wanted to do, right? Anthropic has to go to trial still, they had to pirate the books to train, and they will not win on their right to commercialize the results of training, because neither did Google, so what good is the Fair Use ruling, besides allowing OpenAI v. NYTimes to proceed a little longer?
- dragonwriter 1y ago> Anthropic has to go to trial still, they had to pirate the books to train They did not have to, they had an alternate means available (and used it for many of the books), buying physical copies and destructively scanning them. > and they will not win on their right to commercialize the results of training That seems an unwarranted conclusion, at best. > so what good is the Fair Use ruling If nothing else, assuming the logic of the ruling is followed by the inevitable appeals court decision and becomes binding precedent, it provides a clear road to legally training LLMs on books without copyright issues (combination of "training is fair use" and "destructive scanning for storage and searchability is fair use"), even if the pirating of a subset of the source material in this case were to make Anthropic's existing products prohibited (which I think you are wrong to think is the likely outcome.)
- SoKamil 1y agoWhat if I overfit my LLM so it spits out copyrighted work with special prompting? Where to draw the line in training?
- ninetyninenine 1y agoI mean the human brain can memorize things as well and it’s not illegal. It’s only illegal if said memorized thing is distributed.
- mrguyorama 1y agoBecause humans have rights AI models do not.
- NoOn3 1y agoExactly. If someone wants to compare AI models with humans, maybe then they give AI Models the right to vote and other rights.
- ninetyninenine 1y agoThey use to say the same thing about black people.
- martin-t 1y agoHumans don't scale. LLMs do. Even if LLMs were actual human-level AI (they are not - by far), a small bunch of rich people could use them to make enormous amounts of money without putting in the enormous amounts of work humans would have to. All the while "training" (= precomputing transformations which among other things make plagiarism detection difficult) on work which took enormous amounts of human labor without compensating those workers.
- tartoran 1y agoHumans can only memorize such few texts in comparison so they'd not be scallable in the same sense LLMs are.
- 1y ago
- almatabata 1y agoIf a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?
- bananapub 1y agowhat contract? with who? Meta at least just downloaded ENGLISH_LANGUAGUE_BOOKS_ALL_MEGATORRENT.torrent and trained on that.
- almatabata 1y agoI know, but the article mentions that a separate ruling will be made about that pirating. quote: “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” This tells me Anthropic acquired these books legally afterwards. I was asking if during that purchase, the seller could add a no training close to the sales contract.
- shagie 1y agoWhat contracts? And would it run afoul of first sale doctrine? https://en.wikipedia.org/wiki/First-sale_doctrine https://en.wikipedia.org/wiki/First-sale_doctrine > The doctrine was first recognized by the Supreme Court of the United States in 1908 (see Bobbs-Merrill Co. v. Straus) and subsequently codified in the Copyright Act of 1909. In the Bobbs-Merrill case, the publisher, Bobbs-Merrill, had inserted a notice in its books that any retail sale at a price under $1.00 would constitute an infringement of its copyright. The defendants, who owned Macy's department store, disregarded the notice and sold the books at a lower price without Bobbs-Merrill's consent. The Supreme Court held that the exclusive statutory right to "vend" applied only to the first sale of the copyrighted work. > Today, this rule of law is codified in 17 U.S.C. § 109(a), which provides: > Notwithstanding the provisions of section 106 (3), the owner of a particular copy or phonorecord lawfully made under this title, or any person authorized by such owner, is entitled, without the authority of the copyright owner, to sell or otherwise dispose of the possession of that copy or phonorecord. --- If I buy a copy of a book, you can't limit what I can do with the book beyond what copyright restricts me.
- veggieroll 1y agoBRB, I'm going to download all the TV shows and movies to train my vision model. Just to be sure it's working properly, I have to watch some for debugging purposes.
- ncruces 1y agoYou need to buy one copy of each for the fair use to apply.
- toomuchtodo 1y agoLet everyone donate their DVDs and other physical media. You don’t need to buy it, you just need to possess the media.
- veggieroll 1y agoIndeed, I forsee a "training dataset consortium" arising out of this, whereby a bunch of companies team up to buy one copy of everything and then share it for training amongst themselves (ex. by reselling the entire library to each other for $1).
- toomuchtodo 1y agoLike an Archive? Connected to the Internet?
- veggieroll 1y agoGenius!
- ninetyninenine 1y agoAgreed. If I memorize a book and I am deployed into the world to talk about what I memorized that is not a violation of copyright. Which is reasonable logically because essentially this is what an LLM is doing.
- layer8 1y agoIt might be different if you are a commercial product which couldn’t have been created without incorporating the contents of all those books. Humans, animals, hardware and software are treated differently by law because they have different constraints and capabilities.
- ninetyninenine 1y agoBut a commercial product is reaching parity with human capability. Let's be real, Humans have special treatment (more special than animals as we can eat and slaughter animals but not other humans) because WE created the law to serve humans. So in terms of being fair across the board LLMs are no different. But there's no harm in giving ourselves special treatment.
- layer8 1y agoGenerative AIs are very different from humans because they can be copied losslessly and scaled tremendously, and also have no individual liability, nor awareness of how similar their output is to something in their training material. They are very different in constraints and capabilities from humans in all sorts of ways. For one, a human will likely never reproduce a book they read without being aware that that’s what they are doing.
- jplusequalt 1y ago>So in terms of being fair across the board LLMs are no different Why should "fair" factor into it? The LLMs are not humans, thus they have no rights, and treating them fairly shouldn't come into it. Stop anthropomorphizing linear algebra ffs.
- 1y ago