5 ms·
Pretty amazing to see training data being discussed more openly
by esha_manideep 2y ago
Pretty amazing to see training data being discussed more openly
- WiSaGaN 2y agoIndeed. I think part of the reason when they are not discussed openly may be that much of the data used is copyrighted, which introduces some legal ambiguities.
- YetAnotherNick 2y agoIANAL but hiding something doesn't make someone legally immune. Any company could sue LLM companies and they can't hide it during the case. e.g. there is already a similar case on OpenAI.
- fl0id 2y agoYes, but it at the very least delays any findings while you rake in the cash and try to create a favorable environment. OpenAI even stated that think using copyrighted texts is necessary and should be covered by fair use.