6 ms·
The interesting delta here is that this proves that we can distribute the training and get a functioning model. The scaling factor is way bigger than datacenter
by fabmilo 1y ago
The interesting delta here is that this proves that we can distribute the training and get a functioning model. The scaling factor is way bigger than datacenters
- refulgentis 1y agoThe RL, not the training. No?
- itchyjunk 1y agoRL is still training. Just like pretraining is still training. SFT is also training. This is how I look at it. Models weights are being updated in all cases.
- refulgentis 1y agoSimplifying it down to "adjusting any weights is training, ipso facto this is meaningful" obscures more light than it sheds (as they noted, RL doesn't get you very far, at all)
- comex 1y agoBut does that mean much when the training that produced the original model was not distributed?
- deleted 1y ago[deleted]