7 ms·
I know it’s probably using < 1000x compute of the real Sora, but “pretty good” is stretching it
by zdyn5 2y ago
I know it’s probably using < 1000x compute of the real Sora, but “pretty good” is stretching it
- thriftwy 2y agoJust run all key frames through stable diffusion and it should be quite good.
- resource_waste 2y ago>Just Adding the word 'just' doesnt make it any easier. Something I've noticed is that people who have never done something themselves and are telling someone to do an difficult task, will use: "Just" in-front of it. This is particularly relevant in tech.
- PhoenixFlame101 2y agoIt's very likely the comment you replied to said it in a joking sense
- sertraline 2y agoIt is extremely difficult to tell if the person is joking in a field full of people who think AI is some sort of magic.
- nineteen999 2y agoWell I mean calling any of this diffusion/LLM stuff "AI" is a misnomer to begin with.
- tetris11 2y agoDepends if it's already been done before, in which case "just" would then have been just used quite justly.
- uh_uh 2y agoI don't think there's such a thing as key frames here, just frames. And if you run SD through every frame, the output will be janky because SD doesn't know about temporal coherence.
- ehsankia 2y agoHonestly neither does OpenSora it seems, as it is pretty damn janky already.
- mejutoco 2y agoYou can pass a few frames as a single image grid. Then you will get coherence, although it will be very limited by gpu ram.
- thriftwy 2y agoAs soon as it's an MP4 it will have key frames all right. You could add AI upscaling to your encoder. People are making fun of "Just", but I believe I could take apart ffmpeg to add this feature (PoC) in two weeks or less. Provide somebody pays for my labor and for the HW.
- loudmax 2y agoDepends on your frame of reference. Compared to anything else I've seen generated on a consumer grade GPU, I'd say these are are indeed pretty good. Here's their example gallery: https://hpcaitech.github.io/Open-Sora/ https://hpcaitech.github.io/Open-Sora/ Compared to the outputs from other models run on consumer grade GPUs, I'd say those are very good.
- maxerickson 2y agoWhat's the useful frame of reference? Looking better than other things that are also bad is sort interesting in that it represents progress in some direction, but it isn't very interesting to people outside of the topic.
- roenxi 2y agoWe seem to be in an exponential uptick phase of tech driven by hardware improvements; a few years ago this was impossible on consumer grade GPU. So in some sense there isn't a useful frame of reference, state of the art should improve out-of-sight about every 2 years and eventually I'd expect iPhones to be outgenerating Disney at movies.
- clwg 2y agoFor me I use the Will Smith video[0] from just over a year ago. Compared to the examples it's a pretty stark difference. https://arstechnica.com/information-technology/2023/03/yes-virginia-there-is-ai-joy-in-seeing-fake-will-smith-ravenously-eat-spaghetti/ https://arstechnica.com/information-technology/2023/03/yes-v...
- Lockal 2y agoYes, people learned not to generate other people eating. Current SOTA models still have no concept of walking (left leg, right leg, left leg, right leg; it's so complicated?), there is no reason to believe that they have learned the peculiarities of food consumption.
- forkerenok 2y ago