7 ms·
> Combined with our latest 30T-token multimodal pre-training corpus [...] Is the optimal formula still 20x the amount of model params in tokens for training? C
by garo-pro 21d ago
> Combined with our latest 30T-token multimodal pre-training corpus [...]
Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?