Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stefanbaumann
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
stefanbaumann
3y ago
The models presented in the paper are trained on class-conditional ImageNet (where the input is Gaussian noise and one of 1000 classes, e.g., "car") and unconditional FFHQ (where the input is only Gaussian noise).
2.
▲
by
stefanbaumann
3y ago
Not yet, we focused on the architecture for this paper. I totally agree with you though - pixel space is generally less limiting than a latent space for diffusion, so we would expect good performance inpainting behavior and other editing ta
3.
▲
by
stefanbaumann
3y ago
Both Latent Consistency Models and Adversarial Diffusion Distillation (the method behind SDXL Turbo) are methods that do not depend on any specific properties of the backbone. So, as Hourglass Diffusion Transformers are just a new kind of b
4.
▲
by
stefanbaumann
3y ago
The "input image" is just the noisy sample from the previous timestep, yes. The overall architecture diagram does not explicitly show the conditioning mechanism, which is a small separate network. For this paper, we only trained o
5.
▲
by
stefanbaumann
3y ago
Thanks a lot! Yeah, the main motivation was trying to find a way to enable transformers to do high-resolution image synthesis: transformers are known to scale well to extreme, multi-billion parameter scales and typically offer superior cohe
6.
▲
Direct pixel-space megapixel image generation with diffusion models
(crowsonkb.github.io)
280 points
by
stefanbaumann
3y ago
|
47 comments
7.
▲
by
stefanbaumann
3y ago
It's already a thing [1]. They also have a project website [2] with some nice videos, although the code hasn't yet been released. [1] https://arxiv.org/abs/2308.09713 [2] https://dynamic3dgaussians
8.
▲
Open3D v0.17
(github.com)
2 points
by
stefanbaumann
4y ago
|
0 comments