5 ms·
My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist. I think it's
by simonw 8d ago
My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist.
I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone.
OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable.
My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.
- smcg 8d agoyou massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
- ryan_n 8d agoLAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation. OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
- needfish 8d agoPersonal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.
- ryan_n 8d agoWhat I meant more is that the dataset they scrape is openly available for download by anyone, unlike any of the frontier labs. Not that they don’t scrape copyrighted content. Still sketch, but at least they don’t call themselves “OpenLAION”. Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.
- ccgreg 8d agoCommon Crawl is text-only.
- simonw 8d agoAnthropic too, and evidently others. Amazon were recently confirmed to be doing the same thing: https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/ https://www.404media.co/we-tracked-a-shipment-of-rare-books-...
- dgellow 8d agoAnd that’s bad, right?
- simonw 8d agoIt's legal. I wouldn't do that myself, but I guess that's why I don't train models for a frontier AI lab.
- mahogany 8d agoThe thread is not really about what's legal; the topic is integrity. It sounds like, based on the fact that you wouldn't do it yourself, you agree that it's not a good thing to do.
- ryan_n 8d agoYou're right, it was an over simplification. I think public exchange of data is great for innovation and research (Common Crawl/LAION). But I still think scraping proprietary data without consent or attribution is generally bad (also Common Crawl/LAION). Then you have OpenAI etc.. who build these multi-billion (trillion??) dollar machines and sell them back to people, using everyone's proprietary data, and (among other things) tell everyone it's going to take their jobs. That combination of things doesn't scream integrity to me. Still, it's undeniable that these machines could be beneficial for humanity (cancer research and such). So, I'm sure many people would say the good out-ways the bad. I don't know. Seems that would set a risky precedent for future companies, but maybe not.