6 ms·
The open training data is a huge differentiator. Is this the first truly open dataset of this scale? Prior efforts like The Pile were valuable, but had limitati
by WeirderScience 1y ago
The open training data is a huge differentiator. Is this the first truly open dataset of this scale? Prior efforts like The Pile were valuable, but had limitations. Curious to see how reproducible the training is.
- layer8 1y ago> The model will be fully open: source code and weights will be publicly available, and the training data will be transparent and reproducible This leads me to believe that the training data won’t be made publicly available in full, but merely be “reproducible”. This might mean that they’ll provide references like a list of URLs of the pages they trained on, but not their contents.
- WeirderScience 1y agoYeah, I suspect you're right. Still, even a list of URLs for a frontier model (assuming it does turn out to be of that level) would be welcome over the current situation.
- glhaynes 1y agoThat wouldn't seem reproducible if the content at those URLs changes. (Er, unless it was all web.archive.org URLs or something.)
- dietr1ch 1y agoThis is a problem with the Web. It should be easier to download content like it was updating a git Repo.
- TobTobXX 1y agoWell, when the actual content is 100s of terabytes big, providing URLs may be more practical for them and for others.
- layer8 1y agoThe difference between content they are allowed to train on vs. being allowed to distribute copies of is likely at least as relevant.
- sschueller 1y agoNo problem, we have 25 Gbit/s home internet here. [1] [1] https://www.init7.net/en/internet/fiber7/ https://www.init7.net/en/internet/fiber7/
- evolvedlight 1y agoYup, it’s not a dataset packaged like you hope for here, as it still contains traditionally copyrighted material