8 ms·
I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting
by runnig 3mo ago
I'll just leave it here: "Anthropic's downloading of over seven million books from pirate sites like LibGen constituted infringement, the judge ruled, rejecting Anthropic's "research purpose" defense: "You can't just bless yourself by saying I have a research purpose and, therefore, go and take any textbook you want."
https://www.joneswalker.com/en/insights/blogs/ai-law-blog/why-anthropics-copyright-settlement-changes-the-rules-for-ai-training.html?id=102l0z0#:~:text=Anthropic's%20downloading%20of%20over%20seven,take%20any%20textbook%20you%20want.%22 https://www.joneswalker.com/en/insights/blogs/ai-law-blog/wh...
- nicce 3mo agoYet they did not need to destroy the models which were trained with them?
- zaptrem 3mo agoShould we require the destruction of the brains of those that watch pirated movies?
- TightFibre 3mo agoWell I enjoyed this response.
- hmry 3mo agoDifferent situations call for different responses. When someone steals a watch, we force them to give it back. Yet when someone steals a cake and eats it, we don't force them to puke it back up. If you pirate a movie, the court might very well force you to delete all the copies you made of the movie you downloaded, destroy DVDs you burned, etc.
- raverbashing 3mo agoThanks for proving current copyright law makes no sense Here's a better idea, a fixed fee for any work. You can buy the license to read a book for $X (for whatever purpose) in RAND terms - of course publisher/material costs go on top, so if you're buying an actual book you're getting the material costs as well - or streaming fees or whatever
- shakna 3mo agoYou can already buy books today. Doing so for training is currently considered fair use. Anthropic simply considered that cost prohibitive and chose piracy instead.
- nicce 3mo agoHave we already agreed that AI is already equal to human life and not machine?
- ascorbic 3mo agoUsing them was allowed as fair use – it was the downloading of the pirated copies that was infringement. That's why Anthropic switched to scanning paper books.
- realusername 3mo agoIf using the books is fair use, then distilling the model, which is just a derived product of those books is also fair use. These companies are trying to have their cake and eat it too.
- ascorbic 3mo agoProbably, yes. It's likely just a breach in their terms of service. You'll note that they're not suing them – they're trying to get the government to do their work for them.
- drdaeman 3mo agoHmm, training on a book’s text smears the content all over the weights, merging it with all other texts. The original text isn’t intentionally supposed to be reproducible in any larger part (although IIRC models were able to emit fairly large chunks verbatim). Quite unlikely, training on behavior purportedly approximately replicates the behavior. It gets replicated intentionally as a whole. IANAL, but I see significant differences with intent to copy a significant part as a whole into a competing product, surely shouldn’t fit under legal concept of fair use, no matter whether scanning books for LLM training fits or not. Whether such things (behaviors) are copyrightable - and should they be so - is another interesting question. Those aren’t algorithms or databases (stuff clearly and explicitly covered in many copyright laws), those are human expectation models, something like how we train animals or teach our own.
- didroe 3mo agoIt's the exact same training process for both of your examples. I don't really see how you can claim books are not replicated, but that output from other LLMs is.
- gmerc 3mo agoHow many “capabilities” did they “extract” from those books?
- thepasch 3mo agoThe capabilities of the books' writers to produce the text contained within them, which is exactly what Alibaba "extracted" from Claude. The point here is that Anthropic's framing as some sort of sophisticated technological attack is the ridiculous part. It's writing prompts and saving responses. We're all running "distillation attacks" on Claude, every day! Most of us just don't feed that stuff into a training corpus.
- RobotToaster 3mo ago"You're trying to kidnap what I've rightfully stolen!"
- basisword 3mo agoExactly. Couldn't happen to better people. I'm pretty against piracy personally but if we find reliable ways to pirate Anthropic/OpenAI products in the future I'm all for it.
- scientism 3mo agoDon't you find it funny that when you ask for song lyrics these models suddenly remember copyrighted material?
- f6v 3mo agoSome do, others decline to answer.
- rienbdj 3mo agoIn the early days of music streaming, many of the entrants were seeding their service with vast libraries of pirated content. The winners cut deals with the copyright holders and then went after the rest.
- smurda 3mo agoOr the early days of video uploads, YouTube's most watched videos were "pirated" clips from popular shows (e.g. SpongeBob, The Daily Show) and part of the reason I went to YouTube instead of other video hosting sites (e.g. DailyMotion). Viacom sued YouTube, while CBS and Universal ended up licensing their content. https://www.eff.org/deeplinks/2007/03/viacom-v-google-investing-litigation-rather-innovation https://www.eff.org/deeplinks/2007/03/viacom-v-google-invest...
- radicalbyte 3mo agoThey still are. My kids haven't watched a single Simpsons or Family Guy episode but are quoting both regularly. Facebook et al also quite literally stole email contact lists and installed spyware at kernel level on mobile phones which they used to spy on all Android users. Via the phone manufacturers.
- deleted 3mo ago[deleted]