8 ms·
Openai need to train their models based on these books, not stackoverflow or reddit.
by revskill 1y ago
Openai need to train their models based on these books, not stackoverflow or reddit.
- burkaman 1y agoThey do: https://xcancel.com/vxunderground/status/1888019174133276846 https://xcancel.com/vxunderground/status/1888019174133276846, https://www.theverge.com/2023/7/9/23788741/sarah-silverman-openai-meta-chatgpt-llama-copyright-infringement-chatbots-artificial-intelligence-ai https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o... The tweet only names Meta, but it would be very surprising if OpenAI didn't do the same thing.
- CamperBob2 1y agoAnyone who doesn't train on all material available, legal or otherwise, will be outcompeted by teams that do, including those based in countries that don't respect Western copyright law. It's that simple. Either this is practice is judged (or legislated) to be fair use, or copyright is done. It's also that simple.
- alfalfasprout 1y agoSo, what? Authors and rights holders are supposed to just take it? Copyright law exists for a reason. Trying to improve an LLM doesn't give you the right to flout our legal system. Yes, other countries might have an advantage in LLM training as a result but so be it.
- crazygringo 1y ago> Authors and rights holders are supposed to just take it? If it's judged as fair use, then yes. And then it's not flouting anything. Remember the whole point of fair use is to benefit society by allowing reuse of material in ways that don't directly copy large portions of the material verbatim. For example, nonfiction authors already "just take it" when reviews describe the main points of their book without paying them a cent. The justification is that it's for the greater good, and rights are limited.
- bfrankline 1y ago> Remember the whole point of fair use is to benefit society by allowing reuse of material in ways that don't directly copy large portions of the material verbatim. How do you think masked language models work?
- atrettel 1y agoJudges have recently ruled [1] that training on legally obtained materials constitutes fair use, but we will have to see in the long term if that ruling holds up. [1] https://www.404media.co/judge-rules-training-ai-on-authors-books-is-legal-but-pirating-them-is-not/ https://www.404media.co/judge-rules-training-ai-on-authors-b...
- Night_Thastus 1y ago>the whole point of fair use is to benefit society I'll stop you right there - I really don't think that applies at all. Does 'society' really benefit when the whole thing is a funnel for enormous amounts of wealth to go to already-gigantic companies like Microsoft?
- CamperBob2 1y agoYes, if it helps me get my own job done more effectively, efficiently, and economically. That's how our society works. You and I benefit from this, too, not just Microsoft. If you don't like it, there's a process for changing how it works, but don't expect an easy path to success. Various people will object, and will have to be won over to your way of thinking.
- alfalfasprout 1y ago> If you don't like it, there's a process for changing how it works Except the converse is true. Copyright law today governs how fair use works and even so, how material can be obtained, licensed, etc. To change it to explicitly allow what you're suggesting would require changing copyright law.
- 1y ago
- hyperman1 1y agoCopyright law indeed exists for a reason. And that reason was that church and crown felt threatened by the power of printing presses to distribute ideas they couldn't control. 'To promote the usefull arts' has always been a way to sell the idea to the masses.
- CamperBob2 1y ago"...but so be it." That phrase is carrying a lot of water, isn't it? Trillions of dollars worth by some estimates.
- atrettel 1y agoI'm not convinced that LLMs and other AI models need to train on all material available. A representative sample is better. I'll ignore the legality aspects in my response. I think coming up with a representative sample of all relevant information would be better in the long term (teams will not be outcompeted on long time horizons). Why don't the companies do this? Because it is easier to just "carpet bomb the parameter space" and worry about the potential confounding [1] and sampling bias [2] later. Coming up with a representative sample requires domain expertise and that is expensive in terms of time and money. But it reduces the total amount of training data and should reduce the amount of time and resources it takes to build the models. That may matter now that models are quite large. This is definitely a design decision with tradeoffs on both sides. I can entertain the notion that we don't have time to sample things, but I think we are all too often dismissing the long-term benefits of proper sampling. (In terms of the legality aspects, judges are trying to "split the baby" [3] in my opinion by saying that training on stuff you got legally is OK but training on pirated material isn't. So nobody is going to recommend training on pirated material in the first place.) [1] https://en.wikipedia.org/wiki/Confounding https://en.wikipedia.org/wiki/Confounding [2] https://en.wikipedia.org/wiki/Sampling_bias https://en.wikipedia.org/wiki/Sampling_bias [3] https://www.404media.co/judge-rules-training-ai-on-authors-books-is-legal-but-pirating-them-is-not/ https://www.404media.co/judge-rules-training-ai-on-authors-b...
- CamperBob2 1y agoPerhaps, but it seems safe to assume that the most valuable training material will be the 'illegal' material that is copyright-encumbered.
- spaceport 1y agoQuality. The tranformable value in all data is not equal.
- 9dev 1y agoOr none of both happens and the corporations will just continue to evade laws and taxes to their benefit.
- sigseg1v 1y agoOutcompeted in the competition of what, exactly? How quickly they can produce inaccurate garbage?
- bugufu8f83 1y agoThey do, don't they? I think OpenAI uses libgen. Meta managed to get into a private ebook torrent tracker called Bibliotik a few years ago to use for training Llama and the resulting publicity essentially killed the tracker.