20 ms·
Unsupervised learning of probably symmetric deformable 3D objects from images
- ijpsud 6y agoNot loading properly for me :( I think server may be being overloaded by HN traffic?
- jcims 6y agoStuff like this can obviously be used to make things like deepfakes 'better'. But i think it might be cool for creating virtual meeting rooms, where you can take a webcam shot of a persons face, normalize the skin tones for lighting, map to a 3d surface, then relight it for the virtual room. When you can rig the meshes to drive each other you could wear 'masks' of other peoples faces (or critters).
- st_goliath 6y ago> But i think it might be cool for creating virtual meeting rooms, where you can take a webcam shot of a persons face, normalize the skin tones for lighting, map to a 3d surface, then relight it for the virtual room. You might be interested in project HeadOn at TU München: https://www.niessnerlab.org/projects/thies2018headon.html https://www.niessnerlab.org/projects/thies2018headon.html Justus Thies gave a presentation at our University about a year ago. IIRC they don't use any fancy NN stuff but instead extract the face geometry using a stereo camera and use interpolation to project the movement onto a target mesh. Using stereo goggles for a VR meeting and various other applications were discussed during the presentation, but the main focus was of course on entertaining the audience with fake videos. On a side note: This is probably the closest to a real-life Max Headroom that we have so far. Not sure if that influenced the name.
- jcims 6y agoOh man, this plus the obs thread earlier plus zoom. The mind boggles haha.
- vidarh 6y agoOne of Vernor Vinge's novels describes 3d models transmitted as part of video conferencing, and re-skinning used to fool an adversary (while depending on causing a noisy, reduced bandwidth connection to make it plausible why the model is imperfect, I seem to remember).
- Frost1x 6y agoThese photogrametric, structure from motion, structured light, etc. techniques have been around awhile so I don't think this changes too much though it may make it a bit easier to generate realistic depth maps for other purposes. At some point you can probably reconstruct large portions of scenes in existing movies and change perspectives, especially if you had good techniques (perhaps AI based) for filling occlusions in the data.
- deleted 6y ago[deleted]
- coding123 6y agoI think we slammed it dead.
- hirako2000 6y agoIf it was truly processing on the browser, we wouldn't have jagged it so easily.
- itronitron 6y agobecause we can
- thomasahle 6y agoVideo here: https://youtu.be/5rPJyrU-WE4 https://youtu.be/5rPJyrU-WE4 Can anyone explain why people would use text to speech for something like this, when they have perfectly good voices themselves?
- dointheatl 6y agoBecause reading from a script for five minutes is likely to require multiple takes for someone who isn't a practiced voice actor, while text to speech requires no extra effort on their part?
- thaumasiotes 6y ago> Because reading from a script for five minutes is likely to require multiple takes for someone who isn't a practiced voice actor This depends on how much you can tolerate speech errors. Most listeners will gloss over them, preferring the human voice to the speech synthesizer while not even really noticing the errors.
- deleted 6y ago[deleted]
- Tade0 6y agoNone of the authors appear to be native English speakers, so perhaps they're self-conscious about their accents?
- vidarh 6y agoMy thought as well. The TTS is good enough that it won't take much of an accent before the accent is harder to understand than the TTS as well. I know my own accent is strong enough that I'd have to put in very conscious effort to be easier to understand than this video.
- Frost1x 6y agoCould also have speech problems. Could be lazy. Could want to save time. Could be useful at producing consistent CC information across mediums. Could allow people to choose arbitrary voice synthesis in the future which super futurists may like the idea of. Could have used a translator to produce the text (I haven't listened) and not know English atall. Personally, I'll take the human voice unless you literally cannot speak (e.g. disability) or feel uncomfortable.
- deckar01 6y ago"We store a copy of the uploaded image ..." Is this really happening with a "deep network in the browser"? It looks like it is happening on a server, then the 3D result is viewed in a browser.
- Karuma 6y agoYeah, I'm glad they changed the completely inaccurate title. Right now it doesn't even work due to server overload...
- h91wka 6y agoPlease fix the title of submission. It is " Unsupervised Learning of Probably Symmetric Deformable 3D Objects from Images in the Wild". No anime examples to be found in the paper :(
- michaelt 6y agoThis looks impressive - shame there aren't any examples larger than postage-stamp-sized in the paper or the video.
- nalaka 6y agoLooks like HN has broken their server. I am getting a "Failed to send request to server!" error when I upload an image.
- awinter-py 6y agoif you rotate it 180 deg Z axis, it's like she's watching you
- schoen 6y agoVery striking! This effect is well-studied and also happens with a real physical mask when viewed from the inside: https://en.wikipedia.org/wiki/Hollow-Face_illusion https://en.wikipedia.org/wiki/Hollow-Face_illusion
- EZ-Cheeze 6y agoMore symmetric things: -Motorcycles -Airplanes -Horses Can it work on things other than human and feline faces? From oblique angles? Can it be generalized with enough examples? At any rate, kudos to the authors for drawing attention to the value of symmetry and approximate symmetry.