5 ms·
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two an
by ryan_n 8d ago
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation.
OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
- needfish 8d agoPersonal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.
- ryan_n 7d agoWhat I meant more is that the dataset they scrape is openly available for download by anyone, unlike any of the frontier labs. Not that they don’t scrape copyrighted content. Still sketch, but at least they don’t call themselves “OpenLAION”. Also my understanding was they’re not storing the actual music, but the metadata and a link to the YouTube video.
- ccgreg 8d agoCommon Crawl is text-only.