9 ms·
AI Data Laundering
- dkural 4y agoThis reminds me of the Jedi Mind trick of Uber of waving a smartphone to argue that labor & other laws all of a sudden don't apply to them, to the detriment of the public that'll now shoulder the costs.
- daniel-cussen 4y ago
- Havoc 4y agoThe whole thing is a mess but frankly i doubt this genie can be put back in the bottle
- 9wzYQbTYsAIc 4y agoI think it is an a priori fact that the cat is out of the bag. The existing publicly available datasettes, algorithms, and weighted models certainly should be expected to be permanently in the hands of some non-law-abiding parties, at this point. I think that it will be important to ensure that we have symmetric information, going forward, otherwise trying to put the genie back in the bottle may just end up further disadvantaging those that try to follow the rules.
- ROTMetro 4y ago...said the music industry about samplers in the 1990s.
- bo1024 4y agoThe Flickr example is wild. How was nobody sued for that!?
- moyix 4y agoThe Authors Guild v Google decision about Google Books seems relevant: > In late 2013, after the class action status was challenged, the District Court granted summary judgement in favor of Google, dismissing the lawsuit and affirming the Google Books project met all legal requirements for fair use. The Second Circuit Court of Appeal upheld the District Court's summary judgement in October 2015, ruling Google's "project provides a public service without violating intellectual property law." The U.S. Supreme Court subsequently denied a petition to hear the case. [...] > The court's summary of its opinion is: [...] > Google’s unauthorized digitizing of copyright-protected works, creation of a search functionality, and display of snippets from those works are non-infringing fair uses. The purpose of the copying is highly transformative, the public display of text is limited, and the revelations do not provide a significant market substitute for the protected aspects of the originals. Google’s commercial nature and profit motivation do not justify denial of fair use. https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... This doesn't touch on the ethics of course – at minimum I think allowing people to exclude themselves or their work from a dataset is necessary.
- authpor 4y ago> I think allowing people to exclude themselves or their work from a dataset is necessary. or they could open it all up for everybody and stop protecting the rights of death people (authors dead less then 70 years ago) then again, that will make the publishers starve... but why pretend publishing corporations need food?
- ad404b8a372f2b9 4y agoThis is larger than publishers, this is every artist, film-maker, photographer, every writer, every engineer, anybody who has ever written or created something and shared it publicly is liable to have their work assimilated and an infinite amount of derivatives produced with no control over how they're used and by whom. Comment generated with gpt-neox prompt: Comment about AI and data collection and generation and its pitfalls, expressing concern, emphasis on professions, emphasis on automation, written by Stephen King, creative writing, award winning, trending on reddit, trending on hacker news, written by Greg Rutkowski, written by Zola, written by Voltaire, written by authpor, written by moyix. (Just kidding, it wasn't AI generated but you see my point.)
- killjoywashere 4y ago> It’s currently unclear if training deep learning models on copyrighted material is a form of infringement What? It's clearly a derived work.
- alar44 4y agoAlright, then every piece of music you've ever heard is also a derived work. Unless the composer grew up in a void.
- kevingadd 4y agoI think you're trying to be ridiculous, but lawsuits of that sort do come up periodically - accusations of composition theft that in practice are probably not theft and just the fact that there aren't THAT many unique note progressions you can put into a song, and it's not that odd for a composer to accidentally imitate something they heard a long time ago when picking one.
- lolinder 4y agoHow does AI change the playing field here? Accidental infringement is a thing either way, and the creator should be careful to avoid accidental infringement regardless of whether they're using an AI to do the creation.
- wodenokoto 4y agoI'm pretty sure I can count the number of words in Harry Potter without breaking copyright law. It is absolutely not clear when statistical models stops counting ngrams and starts making a derived works.
- tiborsaas 4y agoYou can also read the HP series and write summaries and reviews about each book as wodenokoto. You can probably create HP looking artwork and write stories that could fit into the HP universe. You can't call any of your work HP. This sounds obvious, but if it's done by a machine then some people think it's a different question. I can write code to get a list of characters in the book, get their page numbers analysed and draw graphs to help me create my own version. Am I breaking copyright laws? Most likely not. It's a truly grey area which lawmakers never saw coming. I believe if events unfold well we'll see and treat AI tools to be like sharp knives eventually. It will be up to the user what they do with it.
- RosanaAnaDana 4y agoThe horse seems well out of the gate.
- VanTheBrand 4y agothe horse is out of the gate on photocopiers and before them printing presses but that doesn't make using them without the rights to what you are copying legal
- pixl97 4y agoI would go as far to say a vast amount of our creative economic output would cease tomorrow if we had strict 'right to read' copyright enforcement. You can also thank Disney for extending copyright beyond all reason.
- VanTheBrand 4y agobut we aren't talking about right to read. we are talking about right to read and then regurgitate and distribute/sell a derivative work of what you read with substantial similarities such that it harms the market for the original work.
- deleted 4y ago[deleted]
- learndeeply 4y ago> But then Meta is using those academic non-commercial datasets to train a model, presumably for future commercial use in their products. Weird, right? This is a very strong and likely inaccurate presumption.
- brrrrrm 4y agoyep. This class of fallacy has its own wiki article: https://en.wikipedia.org/wiki/Appeal_to_probability https://en.wikipedia.org/wiki/Appeal_to_probability
- nerdponx 4y agoIs it? Maybe they have their own internal version they are using, but who's to say they aren't fine tuning the model and applying it somewhere?
- gfd 4y agoWas this term coined on HN? I remember first seeing it (used in an AI context) from this 2019 comment under "Cool stuff that's still completely unregulated": https://news.ycombinator.com/item?id=21167689 https://news.ycombinator.com/item?id=21167689 Most of the predictions in that first comment came true.
- lachlan_gray 4y agoWilliam Gibson mentions data laundering as an illicit activity in the Neuromancer books! It’s plausible that the phrase itself was coined there
- noduerme 4y agoI've got a couple examples of Stable Diffusion replicating watermarks along with similar swatches of imagery into scenes from the same prompt [1]. A single case of this should be enough to file a massive lawsuit if the art were recognizable to the creator. [1] https://news.ycombinator.com/item?id=33061707 https://news.ycombinator.com/item?id=33061707
- danielbln 4y agoThe model learns all attributes of the images it's trained on, including that some have a watermark. The fact that it generates a watermark in some images doesn't mean that that is a 1:1 image from the training set, it just means to the model some images seem to have a watermark, so it will add it sometimes. Often you can just add "no watermark" (or add it as a negative prompt with some weights) and re-use the same seed to get the same image without the watermark.
- ROTMetro 4y agoNo, it means that it is reproducing the original work and is not producing a new original work. It is basically a really fancy Instagram lense, but it is still 100% derived from the underlying works and therefor derivative instead of a newly created non-derivative work.
- danielbln 4y agoI'm not sure how you can make this argument just based on the model synthezing a watermark that is has learned about in the original dataset. Don't forget, the model is only 4GB in size, and while it's not out of the question that it could regurgitate an image from its data set, considering the size of the training which is a few magnitudes larger it is highly unlikely.
- noduerme 4y agoIt may or may not be a 1:1 image, but I think it's significant that in both cases, with different seeds, what is directly behind / right of the watermark is a pretty similar building with different distortions applied to it. I'm not sure what the difference is between "learning" from a particular image and encoding that image with a lot of compression, when in either case the usage more or less reliably reconstructs the image algorithmically. If I have a photographic memory and I memorize the Coca Cola logo and then draw it into a commercial work by decoding the firing of my neurons into muscle movements, the storage and retrieval method I used has no bearing on whether I infringed on their copyright.
- nojvek 4y agoBig Tech has really big datasets esp Google. With YouTube, Photos, Music, Gmail, Docs, Maps, Books, Waymo, Search … they have giant multimodal datasets that capture essence of all human knowledge. They have 10+ products with more a billion users creating data for them. If Google Brain/DeepMind were to crack AGI, it would make Google/Alphabet crazy rich at the detriment of millions of YouTubers, Book authors, musicians, drivers. AI will concentrate power and wealth to fewer individuals.
- mirker 4y agoAds companies getting rich off of AGI seems a bit sensational when they’re already getting rich off of the boring type of AI. They’ve already gotten rich indexing the web and all the data we have years ago.
- patcon 4y agoNot sure laundering it the right term. Laundering private things through the commons feels not as shady as laundering in private networks. The commons benefits too. It's more like open source that money laundering
- krab 4y agoAre we heading towards voiding most of current copyrights or is there a way out of this mess with another patch to the laws?
- theGnuMe 4y agoIt’s definitely fair use. One question I have though is Mickey Mouse protected by copyright or trademark or ? I assume someone other than Disney can’t sell mickeys likeness or is that wrong in art? And if the AI makes a movie?