8 ms·
I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on
by mattcantstop 2y ago
I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to.
If it is trained on private information, then I would have issue with it.
- __loam 2y agoYou're probably among friends on this site, but outside tech coded spaces, most people understand that publicly available is not the same thing as an unlimited license to do whatever you want.
- ben_w 2y agoWhile true, likely more pertinent that most people don't have a clue what's possible legally or technically until it gets in the news. Can't give informed consent if you don't know what the EULA means or what the machines can do.
- Ekaros 2y agoThis site is weird when you compare it to open source software projects and these same companies selling those as a service on their platforms it is again huge massive problem and exploitation... When the license explicitly allows that, without single legal question. I wonder if things would be different if software could be copied and then recreated by these models by the mega corps. Would there still be such push in favour of it?
- Dylan16807 2y agoThe people arguing in favor of AI use are not in fact arguing for "an unlimited license to do whatever you want" so that solves that apparent hypocrisy nice and quick. Also your theoretical software cloner would also make clones of proprietary software, right? I think that would be welcomed just fine.
- __loam 2y agoThey're apparently arguing for the legal right to use all content on the internet to create a product that is commercial and competes with the original content.
- Dylan16807 2y agoAs long as it's only borrowing very small amounts from any particular source work, I think it's fine for a new work to be commercial and compete with the originals.
- __loam 2y agoThey ingested the whole work.
- Dylan16807 2y agoIs that a problem? Ingesting entire works is the norm, no matter how much gets used.
- qwery 2y agoWhy? A statement that extraordinary would be interesting if it had some reasoning alongside it. Also, Facebook posts aren't really "in the public sphere" / publicly accessible, but that's a nitpick.
- deleted 2y ago[deleted]
- uoaei 2y ago"If they can, they will."
- ipaddr 2y agoThe question become are post with a limited reach (friends) the public sphere.
- trimbo 2y agoI don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by keeping it very similar to the original but just different enough to avoid being sued. Despite this, I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Why would we embrace this now that a computer can do it and there's a level of deniability? I don't understand how this benefits anyone.
- autoexec 2y ago> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimately "AI did it" should never be allowed to be used as an excuse. If a company pays for a marketing guy who rips off someone's work and they can be sued for it, then a company that pays for an AI that rips off someone's work should still be able to be sued for it.
- segasaturn 2y agoAFAIK, AI models have no way of differentiating high quality input from garbage. If it's fed peer-reviewed, academic papers as well as a paranoid, violent person's Facebook manifesto it treats them with equal weight as long as the sentences are coherent.
- potato3983 2y agoOn some level it needs to be fed some amount of garbage because it takes in all sorts of garbage inputs like we do. AI that needs painstakingly curated training data isn't interesting in the same way that early lightbulbs that used precious metals and cost too much to be commercially viable aren't interesting.
- swatcoder 2y agoThat sounds compelling when you borrow the marketing term "AI" and position the work as part of a sweeping revolution into some beautiful sci-fi future. It's less compelling when you see the technology as noisy content generators that will flood the network with spam and devour the livelihood and opportunity to learn for low-market artists and programmers. In the former perspective, you may look at this is "well, what's the best way we can make this happen?" while the latter sees it more like "So you insist on making this happen. Are you sure there's a suitably responsible way for you to do that?"
- deleted 2y ago[deleted]
- casenmgreen 2y agoWhen I speak to my friends, it's a conversation not wholly private - after all, I've shared whatever I'm saying with them - but it certainly isn't wholly public. In all our conversations, we have and we understand there are degrees of privacy; that which we share with family, that with friends, that with strangers. When I post on-line, I both expect and expected that my conversations would be between me and the group of people I conversed with. I knew who was reading, and I was fine to write whatever I was writing to that group. I may be wrong, but I think this is generally how people feel, how they act, what they expect, how they are, as humans. We think about who we are writing to. It does not come naturally to imagine that third parties are listening in, or will listen in, in the decades to come. This brings us to now, with a third party, reaching back over ten or fifteen years, for absolutely everyone, everywhere, taking copies of everything it can get access to, for its own use, whatever that may be. I profoundly reject Microsoft, and Google, and all entities and companies which act in such ways, these smiling evils, with their friendly icons and bright colours, happy faces and hundred page T&Cs to utterly obscure and obliterate the truth of their actions.
- mupuff1234 2y agoI think the question is what is included in the "public sphere". If I'm a Facebook user, I definitely don't see posts that I meant to share among friends as something that should be considered part of the public sphere.
- limit499karma 2y agoWe need to distinguish between modalities of machine intelligence and proceed to set policy tailored to each specific type. (Arguably) we can further include the variable of private, public control; and the orthogonal matter of private or public service. Machine intelligence is anthropomorphic in utility, that is it serves as either surrogate or substitute for a human cognitive capability. This permits enumeration of AI utility categories. Broadly we can distinguish between creative, knowledgeable, analytical, judicial, predictive, and directing. As an example use of this approach, consider the case of the AI trained on all public domain material and optionally having had training access to private matter (think Vatican archives). Such an instance should generally not be afforded creative rights, but we would be remiss to restrict its utility as a knowledge base. The other parameters noted in terms of dual of public|private can of course have bearing on setting type specific constraints.
- 1vuio0pswjnm7 2y agoIs there anything in the "public sphere" that is not (a) published to the web and (b) under a license that allows Meta to use it for training "AI". It seems that "AI" is biased toward (1) only bits, and (2) only bits that are published to the internet.
- _heimdall 2y agoAre you concerned that this approach will lead to the abandonment of an open web? If AI companies, and whatever comes next, are expected to take advantage of everything shared online, regardless of copyrights, it seems reasonable that people will stop sharing most things of value.
- dyauspitr 2y agoIf you never display your work online, you’ll probably never gain any traction as an artist.
- _heimdall 2y agoI'd be concerned with getting traction if my art is online and anyone can feasibly copy my style. Maybe that's unimportant and no different than being able to make physical copies, though good forgeries haven't always been so easily done and the forgery is meant to be an identical copy of the original work. It could just be me, but the idea of an attempted identical copy of a well known work feels different than a new creation being passed off as the work of a well known artist. For example, you can claim to have a really god copy of the Mona Lisa but that wouldn't be as valuable as claiming you have a previously unknown, unique work from the artist.
- sli 2y agoThis is just copyright infringement reworded to pretend it's not. I own the things I write, and publishing it on the internet doesn't negate that. OpenAI doesn't have the right to claim it, no matter what they think, and neither does anyone else.
- bruce511 2y agoFirstly publishing something on Facebook explicitly gives them the right to "copy" it. It certainly gives them the right to exploit it (it's literally their business model.) Secondly, Facebook is behind a login, so it's not "public" in the way HN comments are public. You'd have gained more kudos had you argued that point. Thirdly this article I about MetaAI not OpenAI. So, no, OpenAI isn't claiming anything about your Facebook post. I'll assume however that you digressed from the main topic, and were complaining about OpenAI scraping the web. Here's the thing. When you publish something publically (on the internet or on paper) you can't control who reads it. You can't control what they learn from it, or how they'll use that knowledge in their own life or work. You can of course control republishing of the original work, but that's a very narrow use case. In school we read setwork books. We wrote essays, summaries, objections, theme analysis and so on. Some of my class went on to be writers, influenced by those works and that study. In the same way OpenAI is reading voraciously. It is using that to assign mathematical probabilities to certain word pairings. It is studying published material in the same way I did at school, albeit with more enthusiasm, diligence and success. In truth you don't "own the things you write" not in the conceptual sense. You cannot own a concept, argument or position. Ultimately there is nothing new under the sun (see what I did there?) and your blog post is already a rehash of that which came before. Yes, you "own" the text, to the degree to each any text can be "owned" (which is not much.)
- alok-g 2y agoBeautifully written. Thanks.
- jprete 2y agoOpenAI is not reading voraciously, it is not a human being. It makes copies of the data for training. If there were an actual AI system which was trained by continuously processing direct fetches from the Web, without storing them but directly using it when for internal state transitions, then that might make the reading analogy work. But then AI engineers couldn't do all the analysis and annotation steps that are vital to the training process.
- yapyap 2y ago[dead]
- 7bit 2y agoYet, you signed an EULA for every publicly available service that legally prevents you from doing anything the company doesn't want you to do. So why do you want them to be legally using your data without, while they explicitly deny you things like scraping data from their platform.