5 ms·
Our devices are always listening, if you have them set up to listen for hail words like "Siri" or "Hey Google," so I think the Overton window on listening alrea
by bensyverson 6d ago
Our devices are always listening, if you have them set up to listen for hail words like "Siri" or "Hey Google," so I think the Overton window on listening already shifted about 10 years ago. The device actually recording/transcribing what it hears is newer, but you're right that it's becoming more and more normalized. With Granola, I think most people assume that any video call is being transcribed, even if there's no indication in the meeting.
Whatever side you come down on, privacy doesn't stand a chance against a small amount of convenience.
- RASBR89 6d agoIt depends what you mean by ‘listening’. If you look at how this actually works - it’s basically only ever looking for a certain waveform ‘Listening’ sort of implies intelligence is paying attention to what is said. It’s more ‘hearing’ than ‘listening’ There are documents on how these activation words work
- prophesi 6d agoTo look for a certain waveform, wouldn't you need to obtain all of the mic data at all times? I think that's what's implied by "listening" here, but semantics are rarely an interesting discussion.
- zamadatix 6d agoYeah, I can see maaaany layers to have to differentiate between when discussing. E.g., for just a few random examples: - The microphone always vibrates from the sound waves - The microphone is powered in a way that the sound waves change electrical readings in some way - The electrical readings are sent to another component reading them - A component receiving the data does some kind of unbuffered processing related to triggered actions (e.g. a clapper or a activation wave) - Some component temporarily uses a buffer of the data but not for permanent storage (e.g. live, unstored transcription for the deaf or a 'nevermind' after a triggered activation) - Some component stores or sends data generated by the sound, but not necessarily the original audio or even any attributions of who (e.g. voice trigger web search sends the search query as text) - Some component generates a stored copy of the transcription with attribution of who - Some component stores the actual audio in a way that can be later replayed I'd say this "sounds" like a mess to deal with, but then I'd be worried about falling into a category ;).
- theshrike79 6d agoA Clapper[0] from the 80s "looked for a certain waveform", but definitely didn't "listen". What the smart speakers and devices do with the keyword is closer to the clapper than an actual transcription. [0] https://en.wikipedia.org/wiki/The_Clapper https://en.wikipedia.org/wiki/The_Clapper