5 ms·
It is, it's libgen + commoncrawl + wikidump + a bunch of other datasets. OpenAI claim that commoncrawl is roughly 60% of its total training corpus and they also
by dirheist 4y ago
It is, it's libgen + commoncrawl + wikidump + a bunch of other datasets. OpenAI claim that commoncrawl is roughly 60% of its total training corpus and they also claim they use the other datasets listed. They probably also have some sort of proprietary Q&A/search query corpus via Microsoft.
- humanistbot 4y ago> It is, it's libgen + commoncrawl + wikidump + a bunch of other datasets. I'm having trouble finding a source for the libgen claim. Is that confirmed or just rumor?
- mandmandam 4y agoThe ChatGPT Prompt book by LifeArchitect.ai is where I saw it: https://docs.google.com/presentation/d/17b_ocq-GL5lhV_bYSShzUgxL02mtWDoiw9xEroJ5m3Q/edit#slide=id.g1b8e0b333f6_0_124 https://docs.google.com/presentation/d/17b_ocq-GL5lhV_bYSShz...
- dblitt 4y ago> Informed 'best guess' only. > Sources: https://lifearchitect.ai/papers/ https://lifearchitect.ai/papers/ Doesn't seem too convincing to me