5 ms·
Precision: « pre-training data is exhausted » everyone has been saying that for a while now. The graph plotting body mass against brain mass… what does it say
by _l7dh 2y ago
Precision: « pre-training data is exhausted » everyone has been saying that for a while now. The graph plotting body mass against brain mass… what does it say exactly? (where is the link to the prior point on data?). I think we would all benefit from being more critical here and stop idealizing these figures. I believe they have no more clue that any other average ML researcher on all these questions.
- XenophileJKO 2y agoThe other thing that bugged me is the built in assumption that today's model have learned everything there is to learn from the Internet corpus. This is quite easy to disprove. Both in factual retention, but also meta cognition on the context of the content.
- jebarker 2y agoYeah, exactly. A human can learning vastly more about, say, math from a much smaller quantity of text. I doubt we're anywhere close to exhausting the knowledge extraction potential from web data.
- bbor 2y agoWhich is exactly the point he's making, I believe; that simply collecting more data isn't the next step. That we've reached a local plateau in scaling ability based on corpus size. Which was assumed by pretty much everyone outside the DL elite the whole time, AFAIU
- esperent 2y agoRight, but there's nothing new in that statement. I've been hearing that we're running out of data for training AIs for two years at least.
- cma 2y agoAlso left out the Baidu scaling laws paper from 2017, and his circle has a history of a kind of citation ring type thing leaving earlier stuff out https://research.baidu.com/Blog/index-view?id=89 https://research.baidu.com/Blog/index-view?id=89