5 ms·
They never said model training is a violation of copyright. The ruling says model training on copyrighted material for analysis and research is NOT copyright in
by gitremote 1y ago
They never said model training is a violation of copyright. The ruling says model training on copyrighted material for analysis and research is NOT copyright infringement, but the
commercial use of the resulting model is:
"When a model is deployed for purposes such as analysis or research… the outputs are unlikely to substitute for expressive works used in training. But making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially where this is accomplished through illegal access, goes beyond established fair use boundaries."
- Workaccount2 1y agoThe vast trove of copyright work has to refer to training. ChatGPT is likely on the order of 5-10TB in size. (Yes, Terabyte). There are college kids with bigger "copyright collections" than that...
- gitremote 1y agoNo. The paragraph as a whole refers to the "outputs" of vast troves of copyrighted work. Disk size is irrelevant. If you lossy-compress a copyrighted bitmap image to small JPEG image and then sell the JPEG image, it's still copyright infringement.
- nickpsecurity 1y agoI won't say it's irrelevant. How much you use is part of fair use considerations. Their huge collections of copyrighted works make them look worse in legal analyses.