4 ms·
Yeah, it's an exciting use case. Although all these models, even seemingly GPT-Live-1 doesn't actually hear your pronunciation, it seems they all get passed tra
by embedding-shape 6d ago
Yeah, it's an exciting use case. Although all these models, even seemingly GPT-Live-1 doesn't actually hear your pronunciation, it seems they all get passed transcripts, so for learning to speak another language, they're still not there seemingly.
But, it's close! You can control their pronunciation, make them speak slower/faster, and obviously great at anything text, so many use cases work great for language learning with LLMs. Just wish they solved this last mile thing too!
- peab 6d agoSome of the models claim to be audio to audio, like one of the Gemini models. But I've tested and it does seem that you're right, it's not getting all the nuance at all
- embedding-shape 6d agoYeah, sadly "audio to audio" seems to mean "we transcript it automatically for you internally which gets passed to the model", otherwise we'd be seeing models that are able to hear nuance in the input voice and pronunciation, which AFAIK, no model does yet.
- jimmySixDOF 5d agothe latest Google Translate features based on the all voice real-time 3.5 Live model is as good as I have tried in Live Modes it's not perfect but I can put it down on a table of four or five people conversing and get a reasonable amount of it translated into my earpiece.