7 ms·
I was under the impression that all of them try to train their models on all data they can possible ibtain. I'd think what ever data this Newsmax has - it woul
by konart 1mo ago
I was under the impression that all of them try to train their models on all data they can possible ibtain.
I'd think what ever data this Newsmax has - it would be in all datasets by now or soon enough. Was I wrong?
- cyanydeez 1mo agothat was the first stab at it; but it's pretty clear that you try to get good data, especially when you see how much the open source chinese models are doing with far less compute and parameter sizes. These are the open source datasets: https://huggingface.co/datasets/HuggingFaceFW/fineweb https://huggingface.co/datasets/HuggingFaceFW/fineweb I don't think they do that anymore, but at the beginning, yes, they threw all kind of bullshit in there to get a proper gradient.