21 ms·
While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It'
by tmjdev 2y ago
While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.
- rob 2y agoProbably because you're expecting it and looking at a demo page. Put these voices behind a real video or advertisement and I would imagine most people wouldn't be able to tell that it's AI generated at all.
- Veen 2y agoIt'd be annoying to me whether it was AI or human. The faux-excitement and pseudo-bonhomie is grating. They should focus on how people actually talk, not on copying the vocal intonation of coked-up public radio presenters just back from a positive affirmation seminar.
- xnx 2y agoAgreed. To be fair, I also get annoyed by fake/exaggerated expression from human podcasters.
- iNic 2y agoIt sounds like every sentence is an ad read.
- JoblessWonder 2y agoYeah... It isn't that it doesn't sound like human speech... it just sounds like how humans speak when they are uncomfortable or reading prepared and they aren't good at it.
- semitones 2y agoI suppose it doesn't matter if it is a human, or a bot delivering the message, if the message is boring
- beoberha 2y agoTotally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.
- amelius 2y ago> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.
- bryanrasmussen 2y agoI'm hoping it will be a lot of Ok Guv'ner and right you ares in the style of Dick Van Dyke.
- kelseyfrog 2y agoRight mate
- mindcrime 2y agoGor blimey lad, that's the problem now innit???
- Dilettante_ 2y agoCheeky bugger, you are
- lebuffon 2y agoee by gum
- KineticLensman 2y ago> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English https://en.wikipedia.org/wiki/Multicultural_London_English
- onion2k 2y agoThat could just be the context though. Listening to a clip that's a demo of what the model can produce is very different to listening to a YouTube video that's using the model to generate speech about something you'd actually want to watch a video of.
- kaibee 2y ago> Example of a multi-speaker dialogue generated by NotebookLM Audio Overview, based on a few potato-related documents. Listening to this on 1.75x speed is excellent. I think the generated speaking speed is slow for audio quality, bc it'd be much harder to slow-down the generated audio while retaining quality than vice versa.
- moralestapia 2y agoIt's due to the histrionic mental epidemic that we are going through. A lot of people are just like that IRL. They cannot just say "the food was fine", it's usually some crap like "What on earth! These are the best cheese sticks I've had IN MY EN TI R E LIFE!".
- shermantanktop 2y ago“I’m OBSESSED with the dipping sauce. So good.”
- deleted 2y ago[deleted]
- hyperific 2y agoIt's like their training set was made up entirely of awkward podcaster banter.
- yapyap 2y agothey all sound like valley-people, complete with the raspy voice and everything
- deleted 2y ago[deleted]
- swatcoder 2y agoWe're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clusters of speech behavior that lets our brain go "yeah, I'm used to things fitting together in this way sometimes" These don't have that. Instead, they seem to mostly smudge together behaviors that are just generally common in aggregate across the training data. The speakers all voice interrupting acknowledgements eagerly, they all use bright and enunciated podcaster tone, they all draw on similar word choice, etc -- they distinguish gender and each have a stable overall vocal tone, but no identity. I don't doubt that this'll improve quickly though, by training specific "AI celebrity" voices narrowed to sound more coherent, natural, identifiable, and consistent. (And then, probably, leasing out those voices for $$$.) As a tech demo for "render some vague sense of life behind this generated dialog" this is pretty good, though.
- lancesells 2y agoAgreed. To me it sounds like bad voice-over actors reading from a script. So the natural parts of a conversation where you might say the wrong thing and step back to correct yourself are all gone. Impressive for sure.
- pvarangot 2y agoIt's because it's probably trained with "professional audio", ads, movies, audiobooks, and not "normal people talking". Like the effect when diffusion was mostly trained with stock photos.
- gwbas1c 2y agoI get the feeling that this is useful for something that someone half-listens to.
- narag 2y agoWhile it is impressive and I like to follow the advancements in this field... Please don't think that I'm trying to suggest... anything . It's just that I'm getting used to read this pattern in the output of LLMs. "While this and that is great...". Maybe we're mimicking them now? I catch myself using these disclaimers even in spoken language.
- tmjdev 2y agoI like to preface negativity with a positive note. Maybe I am influenced in my word choice but my intent was to point out that this is a very, very impressive feat and I don't want to undermine it.
- ljf 2y agoWhen I got to the bit where they referred to the smaller training set of paid voice actors, that hit it for me. It certainly sounds like they are throwing the 'um' and 'ah's in to a script - not naturally. This is good, but certainly not yet great.
- MrSkelter 2y agoI agree. It’s profoundly sad that so much energy is being invested in solving the non-problem of making long documents accessible. To think that people will ignore carefully written work for the “chat show” output of an LLM is horrifying and a harbinger of a societal slide into happy stupidity and willing ignorance.
- nl 2y agoWhilst I don't doubt you feel like that the general response to the notebook LLM podcast feature (which uses this) has been very well received generally. In general people find the back and forth between the "hosts" engaging and also gives people time to digest the contents.
- Cthulhu_ 2y agoI tuned it out instantly because I have that feeling with most Americans / podcasts / etc already. That said, it's a convincing enough analog for that kind of content I think.
- chrismorgan 2y ago> Audio clip of two speakers telling a funny story, with laughter at the punchline. In similar vein, I’m glad they told me it was a funny story, because otherwise I wouldn’t have known.
- pmontra 2y agoIt doesn't feel any different to me than listening to a random radio station where I don't know who is speaking. I didn't feel any uncanny valley but I'm not an English native speaker so I might miss some nuances. However there are relatively few English native speakers around the world so this might not be a problem for us. The problem is that people talking over each other is not a format I long to listen to.
- lokimedes 2y agoFor me it isn’t uncanny from a lack of humanity. Rather, it triggers all my “fake and shallow” personality biases. It certainly sounds human enough, just not the type of humans I like.
- jeksicjjdjisos 2y agoThere’s a certain fakeness to the rhythm of the space between words. Particularly the “uh” and “um” filler sounds. To me it sounds like they always either come in abnormally early or late after speaking those sounds
- vel0city 2y agoI got a similar feeling. I think it was overdoing the ums and uhhs for something trying to sound like an even slightly professional podcast kind of sound.