5 ms·
You're not wrong that the dataset and compute are important, and if you browse the author's previous work, you'll see there are datasets available. The reprodu
by w1nk 4y ago
You're not wrong that the dataset and compute are important, and if you browse the author's previous work, you'll see there are datasets available. The reproduction of DALL-E 2 required a dataset of similar size to the one imagen was trained on (see: https://arxiv.org/abs/2111.02114 https://arxiv.org/abs/2111.02114).
The harder part here will be getting access to the compute required, but again, the folks involved in this project have access to lots of resources (they've already trained models of this size). We'll likely see some trained checkpoints as soon as they're done converging.