Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andrew-w
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
andrew-w
8mo ago
We have not released the weights, but it is fully available to use in your websites or applications. I can see how our wording there could be misconstrued -- sorry about that.
2.
▲
by
andrew-w
8mo ago
So glad you enjoyed it! We've been able to significantly reduce those text hallucinations with a few tricks, but it seems they haven't been fully squashed. The /imagine command only works with the image at the moment, but we&
3.
▲
by
andrew-w
8mo ago
Not something we had thought to do tbh, but would definitely enhance the experience. And, should be reasonable to do. Thanks!
4.
▲
by
andrew-w
8mo ago
We have not released the weights, but it is fully available to use in your websites or applications. I can see how our wording there could be misconstrued -- sorry about that. You can absolutely create a vTuber persona. The link in the post
5.
▲
by
andrew-w
8mo ago
Haha, I kind of get that reaction. Convincing the world "this was hard to do" is generally not easy. Re: user uploads, we're operating in good faith at the moment (no built-in IP moderation). This hasn't been an issue so
6.
▲
by
andrew-w
8mo ago
Thank you! Impressive demo with OVA. Still feels very snappy, even fully local. It will be interesting to see how video plays out in that regard. I think we're still at least a year away from the models being good enough and small enou
7.
▲
by
andrew-w
8mo ago
I wonder how it would come across with the right voice. We're focused on building out the video layer tech, but at the end of the day, the voice is also pretty important for a positive experience.
8.
▲
by
andrew-w
8mo ago
Thanks for the feedback. The current avatars use a STT-LLM-TTS pipeline (rather than true speech-to-speech), which limits nuanced understanding of pronunciations. Speech-to-speech models should solve this problem. (The ones we've tried
9.
▲
by
andrew-w
8mo ago
This isn't natively supported -- we are continuously streaming frames throughout the conversation session that are generated in real-time. If you were building your own conversational AI pipeline (e.g. using our LiveKit integration), I
10.
▲
by
andrew-w
8mo ago
Thanks! And sorry! I can see how our wording there could be misconstrued. With a real-time model, the streaming infrastructure matters almost as much as the weights themselves. It will be interesting to see how easily they can be commoditiz
11.
▲
by
andrew-w
8mo ago
Yep, the model is running on Hopper architecture. Anything less was not sufficient in our experiments.
12.
▲
by
andrew-w
1y ago
Thanks for trying it out! character.ai has put their model behind a waitlist, so it's hard to compare. As far as I can tell, they don't appear to make any specific claims about speed or interactivity in their press release.
13.
▲
by
andrew-w
1y ago
thanks for trying us out!
14.
▲
by
andrew-w
1y ago
Just added a signup at the bottom of the technical report: https://lemonslice.com/live/technical-report
15.
▲
by
andrew-w
1y ago
glad to bring a little joy into the world :)
16.
▲
by
andrew-w
1y ago
It works with any style of character! Check out the embedded videos in our tech report. Peachy and the toilet are my favorite. https://lemonslice.com/live/technical-report
17.
▲
by
andrew-w
1y ago
Thanks! We think we can cut down the latency to <2s which should make it feel even more natural.
18.
▲
by
andrew-w
1y ago
Thanks! What kind of use case are you thinking about?
19.
▲
by
andrew-w
1y ago
It's something we are considering. What use cases do you have in mind?
20.
▲
by
andrew-w
1y ago
We've been very inspired by interactive character experiences powered by traditional VFX + puppetry (turtle talk with crush is a favorite). I think that sort of interactive entertainment will become more commonplace as tech like ours c
21.
▲
by
andrew-w
1y ago
Thanks for the feedback. This is definitely a demo where every piece matters for maximizing the enjoyment factor. We spent the most effort on optimizing video quality and latency, but not a lot on tweaking the character prompts that go into
22.
▲
by
andrew-w
1y ago
Not relying on facial keypoints means we can animate a wide range of non-humanoid characters. My favorite is talking to the Doge meme.
23.
▲
by
andrew-w
1y ago
One way this differs is in the model architecture. Our approach relies on a single pass of a diffusion transformer (DiT), whereas Live Portrait relies on intermediate representations and multiple distinct modules. Getting a DiT to be real-t
24.
▲
by
andrew-w
1y ago
I spent about 2 hours recording videos with different characters. Of course, the one I made as a joke for myself and never intended to share was the most enjoyable to watch :)
25.
▲
by
andrew-w
1y ago
Just added as a public character :)
26.
▲
by
andrew-w
1y ago
Thanks for the feedback. Optimizing for speed meant we had fewer LLMs to choose from. OpenAI had surprisingly high variance in latency, which made it unusable for this demo. I think we could probably do a better job with prompting for some
27.
▲
by
andrew-w
1y ago
We're back online! One of our cache systems ran out of memory. Oops. Agree on improved messaging.
28.
▲
by
andrew-w
2y ago
We've talked about doing something like that. Feels like it should work in theory.
29.
▲
by
andrew-w
2y ago
Cool use case! Thanks for sharing your thoughts.
30.
▲
by
andrew-w
2y ago
Makes sense, thank you!
More ›