Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
JonathanFly
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
JonathanFly
3y ago
According to https://news.ycombinator.com/item?id=38539167#38539479 Twitch is the most popular streaming service in Korea, just slightly ahead of Afreeca.
32.
▲
by
JonathanFly
3y ago
Are bandwidth fees just as high for any domestic Korean company creates a Twitch competitor? Or do the 3 major telecoms operate their own streaming services, so only companies under their umbrella can afford a site like Twitch?
33.
▲
by
JonathanFly
3y ago
>Mentality: SK Terran is a highly aggressive style that is intended to pressure the Zerg with a large number of MnM and Science Vessels. The ideal scenario is to split apart a big MnM ball and engage in guerrilla tactics around the map w
34.
▲
by
JonathanFly
3y ago
>Entertainment and the infrastructure that delivers it is an important pillar of soft power. Might be a miscalculation if the goal is soft power as in cultural export (the so called Korean Wave). At least in the gaming/streaming spa
35.
▲
by
JonathanFly
3y ago
>Are infill and outpainting equivalents possible? Do you mean outpainting as in you still what words to do, or the model just extends the audio unconditionally the way some image models just expand past an image borders without a specifi
36.
▲
by
JonathanFly
3y ago
It autodetects language by default but you can set to a specific one. Though you'd still have that problem on a multi-lingual input video.
37.
▲
by
JonathanFly
3y ago
Genuinely impressive used as intended. Spectacular when not - English to English translation. https://twitter.com/jonathanfly/status/1711607166371561805
38.
▲
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
(stability.ai)
374 points
by
JonathanFly
3y ago
|
203 comments
39.
▲
by
JonathanFly
3y ago
>So how long will it be until we will be able to download something of this quality onto a future-gen Raspberry Pi which can do some AI processing, where we make an HTTP call and it starts speaking through the audio out in a perfect voic
40.
▲
by
JonathanFly
3y ago
Yeah I often try to think about what might be in a YouTube caption when finding prompts that work in Bark. But pipe character isn't one I remember seeing on YouTube. Maybe it's part of some other audio dataset though. Or maybe it&
41.
▲
by
JonathanFly
3y ago
Interesting that SoundStorm was trained to produce dialog between two people using transcripts annotated with '|' marking changes in voice. But the exact same '|' characters seem to mostly work in the Bark model out of t
42.
▲
by
JonathanFly
3y ago
I just realized we already chatted briefly about banned tokens (I'm the Grover tongue twister guy) but I somehow completely missed this gist at that time. Total facepalm moment, would have been helpful reference.
43.
▲
by
JonathanFly
3y ago
AHhhhhhhhhhhhhhhhhhhhhhhhhhhhh. That's me screaming. I constantly wonder why this stuff does not exist in LLMs. But my technical depth and competence is quite low. Way lower than the people implementing the models and samplers. So I ju
44.
▲
by
JonathanFly
3y ago
What's the difference between "Search" and "Edit song"? It seems like Search is just about length, and Edit also considers sections you marked? Or does Edit do the same thing but with many edits, rather than a few?
45.
▲
by
JonathanFly
3y ago
Replying to myself in an old thread as a little easter egg. Bark is just so fun I can't resist a teaser. Suno is seriously underselling the power of the fully operational Bark model. I'm already cranking out "French Obamas&
46.
▲
by
JonathanFly
3y ago
>The way it tends toward context-sensitive shifts in delivery is kind of amazing A blessing and a curse! Super cool though.
47.
▲
by
JonathanFly
3y ago
Leverage sure, but you still basically need a model that does the opposite. You can't like take OpenAI Whisper, which turns speech to text, and just run the model backwards and generate audio. For example. I mean you probably could wit
48.
▲
by
JonathanFly
3y ago
Must be some low hanging fruit to optimize in Bark. It would be somewhat close to realtime if it was close to 100% and scaled linearly.
49.
▲
by
JonathanFly
3y ago
Yeah, it's kind of hand crafted. There's more to the story and more results. I would normally just Tweet but I think it's actually so interesting that it deserves more than a tweet, at least a thoughtful writeup or a youtube
50.
▲
by
JonathanFly
3y ago
Actually I was just checking, and Bark isn't that close to maxing out GPU utilization. Running two instances on a 3090 seems like a throughput increase and the models fit. Update: And getting weird CUDA issues. Hmn...
51.
▲
by
JonathanFly
3y ago
I barely touched that, it's just from the Serp cloning github, but people kept asking so I put it in. Their clone just isn't really doing much though, it's just loading up the coarse model with the encoded wav file as a fake
52.
▲
by
JonathanFly
3y ago
Thanks. I'll consider it but I haven't deployed a model or service like that, so it'd be more of a second project in itself than a funding mechanism, probably. I was just realizing this morning how far behind I am on paying w
53.
▲
by
JonathanFly
3y ago
Not used at all, at least in my case.
54.
▲
by
JonathanFly
3y ago
>Yeah I was trying to figure out how good it was in Korean. The cadence and flow was pretty good but there was kind of artifacts in the audio. Then I check the samples of the default audio prompts for Korean any my god, they were godawfu
55.
▲
by
JonathanFly
3y ago
I'll link my Bark fork with long audio generation and other features on the root thread, I suppose: https://github.com/JonathanFly/bark There's going to be a big update this week with some new stuff I haven&#
56.
▲
by
JonathanFly
3y ago
It's just a characteristic of some of the default voices. I made some perfect French voices by having somebody on Discord check for native accents, because otherwise I can't tell. Once somebody does this for every language it shou
57.
▲
by
JonathanFly
3y ago
It's like 50% realtime on a 3090, not quite real time on a 4090. You can also use smaller models and it's a bit faster. 3080 is same speed, you don't need the extra memory.
58.
▲
by
JonathanFly
3y ago
>I'm curious, how did you generate the David Attenborough voice? The repo says: >> Bark tries to match the tone, pitch, emotion and prosody of a given preset, but does not currently support custom voice cloning Check back late
59.
▲
by
JonathanFly
3y ago
>when i tried a similar on elevenlabs, it sounded a lot worse in terms of the metallic s-s at the end. That's interesting. When I'm judging Bark I'm looking at my own random samples, but for eleven I'm seeing stuff pe
60.
▲
by
JonathanFly
3y ago
I actually think Bark actually beats Eleven right now. Bark tends to add a bit more of a metallic twinge towards the end of longer audio segments (I wonder if this is a bug and not a limitation actually...), but Bark is more expressive. Th
More ›