6 ms·
I've been training different architecture 1.33b and 3.33b models using nvidia megatron for about a month. You can train a 1.33b on a 46gb mac and 3.33b is about
by hadlock 6d ago
I've been training different architecture 1.33b and 3.33b models using nvidia megatron for about a month. You can train a 1.33b on a 46gb mac and 3.33b is about as big as you can go with a single 96gb blackwell with regular checkpoints etc