8 ms·
This is a super cool device. Note that the decoding is highly limited: they decode into one of five different sentences. This is easier than five words for exam
by tbenst 2y ago
This is a super cool device. Note that the decoding is highly limited: they decode into one of five different sentences. This is easier than five words for example as there is more information to distinguish.
Unfortunately the media is blowing this way of out proportion as the larynx alone does not contain sufficient information to decode silent speech.
If you also sense the lips, tongue articulators, and jaw, then general English decoding becomes possible with high accuracy (eg see our recent work here: https://x.com/tbenst/status/1767952614157848859 https://x.com/tbenst/status/1767952614157848859). It’s not in the preprint but I’ve done experiments with only the larynx recorded and performance is pretty abysmal on even a 10 word vocabulary—-hence why they did a five sentence task.
- irviss 2y ago> If you also sense the lips, tongue articulators, and jaw, then general English decoding becomes possible with high accuracy A bit OT but I see this frequently and I'm curious. Why do you English speakers (or just a US phenomenon?) tend to use the word "English" instead of "language", "linguistic" or one of its related words to refer to a general concept?
- atopal 2y agoThere are about 6000 spoken languages around the world with an extreme variety in how they produce meaning. How could you make sweeping statements about all of them?
- x1798DE 2y agoNot OP, but as a native English speaker and former scientist (though not in this area), I would interpret "x does y on English tasks" to mean "we tested this in English and don't know if the effect generalizes to other languages".
- thaumasiotes 2y agoIn this case we do know if the effect generalizes to other languages. It cannot fail to; the larynx, lips, tongue, and jaw are almost all there is. For example, vowels are conventionally defined by jaw position ("height"), tongue position ("frontness"), and lip configuration ("rounded" or not). You might miss some things like creaky voice or ejectives, you'll probably miss aspiration, but all that does is give you a worst-case scenario analogous to a native speaker trying to understand someone with a foreign accent. Extremely high accuracy will be possible.
- AlecSchueler 2y agoThis is a reasonable hypothesis but if only English has been studied then it would be unscientific to extrapolate at this time.
- thaumasiotes 2y agoSure, in the same sense that it would be "unscientific" to conclude that someone's amputated leg didn't regenerate by chance, because the sample size is only 1. If you know how you're recognizing English, and you know that other languages do not differ from English in relevant ways, then you know you can recognize those other languages. Pretending you don't know something you do know is not scientific.
- metabagel 2y agoOther languages have different sounds which aren’t present in English.
- thaumasiotes 2y agoSo? They don't have sounds that are produced in a manner other than arranging the lips, tongue, and jaw. (Actually, they do. So does English; I already mentioned aspiration. But those are minor elements.)
- wizzwizz4 2y agoThey're minor elements in English – and even then, you can construct sentences where the meaning changes based on aspiration.
- thaumasiotes 2y agoWell, no, they're minor elements everywhere. You don't need to be able to capture every phonemic distinction in a language to get a near-perfect transcription, as witnessed by the fact that people understand foreign accents without difficulty. The much larger problem in understanding foreign speech is the odd word choices and lack of grammaticality, but those problems don't arise when you're transcribing native speech. For some comparisons, think about the fact that Semitic languages are traditionally written without bothering to indicate the vowels, or that while modern English has a phonemic distinction between voiced and unvoiced fricatives, this has a very uneven correspondence to the same distinction as it exists in the writing system. In the case of the interdental fricatives, the writing system does not even contemplate a distinction. And there's nothing particularly problematic about this; if you delete all the voicing information from a stretch of English speech, it stays about as intelligible as it was before. (A voicing difference in stops is not even audible to English speakers. It's audible in fricatives, but no one is going to be confused.)
- khazhoux 2y agoThis is your misperception. In the instances where a person says "English" in this kind of context, it catches your attention and you infer that the person is an English-speaker, and possibly American. But when a person uses the generic word "language", you don't notice it. This leads you to believe that English speakers "tend to use the word English," when that's not the case necessarily. I don't know what this perceptual fallacy is called, but there's probably a word. In English :-)
- johnisgood 2y agoI have not noticed this. I just assume that they are specifically talking about a language, in this case: English.
- roenxi 2y agoI'd speculate English speakers are used to being part of a society where non-English speakers are present and politically important. It is polite not to assume that English = language. Even on the British Isles English isn't a universal thing. Let alone somewhere like America where it isn't even native. "Language" just doesn't mean "English". In Australia if someone is talking about "language" on its own I'd assume they're Aboriginal advocates.
- AlecSchueler 2y ago> Even on the British Isles English isn't a universal thing. Let alone somewhere like America where it isn't even native. English isn't native to all of those isles either only Great Britain.
- ImHereToVote 2y agoI bet if you listened to the feedback you could teach yourself to talk using the larynx and surrounding muscles.
- jvanderbot 2y agoWhy can't the muscles of the larnex and perhaps chest / diaphragm, be monitored and mapped to vocal chord noises, rather than full speech? Just put the noise in the throat and let the rest of the body make it work.