6 ms·
Well open AI raised eye brows by crawling the internet and using everyone's data to make a commercial product One day some new startup will train on all of li
by JimmyRuska 3y ago
Well open AI raised eye brows by crawling the internet and using everyone's data to make a commercial product
One day some new startup will train on all of libgen and torrent networks, but it will be very hard to prove. You'll keep getting these gaps up in questionable morality and legality, and even openai will complain about playing fair
- fragmede 3y agoGoogle Classroom, teenager's essays, written by humans, for learning what it means to be human, and graded by humans, is a richer dataset than anything else I can think of that anyone else couldn't get their hands on.
- londons_explore 3y agoAn awful lot of teachers can grade a 10 page essay in about 90 seconds... Skim read it, mark out some grammar errors, assign it a grade based on the quality of the opening and closing paragraphs.
- fragmede 3y agoYup, and they're doing it the whole country over, and putting that data in to Google Classrooms for Bard to know "this is C-grade work" and "this is A-grade work". Knowing what's deemed good and bad writing is where I'm thinking this dataset shines for training LLMs.
- why_only_15 3y agoMany people train on libgen/torrent in the form of books3 (e.g. LLaMa does this).
- pas 3y agoThePile already contains some content from a torrent, and there's as lawsuit alleging that Meta has committed copyright infringement by using it. https://www.theverge.com/2023/7/9/23788741/sarah-silverman-openai-meta-chatgpt-llama-copyright-infringement-chatbots-artificial-intelligence-ai https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o...