6 ms·
Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem
by probe 2mo ago
Do you think 3 is better than 1 & 2 as context gets larger? I think for smaller data sets its mostly fine no. It's an interesting bet by TM. End state does seem some form of continual learning (model weights update like dreaming)