6 ms·
The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I sus
by cyp0633 1y ago
The same happens with whisper-large-v3 on Chinese transcription: silence is transcribed to something like "please upvote, share and favourite this video". I suspect they trained the model on some random YouTube video without carefully picking really useful data.
- st_goliath 1y agoThat's interesting, the few times I tried playing with whisper, I had the impression that YouTube style videos or random cellphone videos was something it did particularly bad with (compared to movies). My guess at the time was that most of the training material might be sub titles and raw screen plays. The videos I tried to transcribe were also Mandarin Chinese, using whisper-large-v3. Besides the usual complaints that it would phonetically "mishear" things and generate nonsense, it was still surprisingly good, compared to other software I played around with. That said, it would often invent names for the speakers and prefix their lines, or randomly switch between simplified and traditional Chinese. For the videos I tested, intermittent silence would often result in repeating the last line several times, or occasionally, it would insert direction cues (in English for some reason). I've never seen credits or anything like that. In one video I transcribed, somebody had a cold and was sniffling. Whisper decided the person was crying (transcribed as "* crying *", a cough was turned into "* door closing *"). It then transcribed the next line as something quite unfriendly. It didn't do that anymore after I cut the sniffling out (but then the output switched back to traditional Chinese again).
- isoprophlex 1y agoIndeed, with another model I would get persistent transcriptions of silent parts into 'Thanks for watching!' or '[MUSIC]'. Pretty dumb that this failure mode wasn't caught in some QA process, and there are now multiple transcription models suffering from the same issue. Having silent parts in your input audio seems like it should be a very common occurrence...
- rollcat 1y agoWhen I was taught mathematics, the zero value was always considered the most important edge case. You prove something for N=0 (or N=1), then for N=M+1. It's even more important in audio DSP: processing near-zeroes can end up being extremely CPU intensive, look up denormal/subnormal floats.
- inglor_cz 1y agoYeah, I studied mathematics (algebra and number theory) and zero is the point, often sporting discontinuities, or weird asymptotic behavior. Quite a lot of algorithms use some form of division and zero is the only number in our typical structures (Z, Q, R, C), that cannot be used to divide with.
- isoprophlex 1y agoWell, now in this brave new age of AI we can enjoy computer programs crashing with an Error: division by please upvote, share and like!
- xyproto 1y agoThis also works; I upvoted your comment.
- o1bf2k25n8g5 1y agoI have discovered a truly marvelous proof of how to smash that like and subscribe button, which this comment box is too small to contain.
- wahnfrieden 1y agowhisper MUST be combined with silence detection / VAD
- pferde 1y agoAh, the good old "you're holding it wrong". What good is a speech recognition tool that literally hears imaginary voices?
- Xmd5a 1y agofaster-whisper has a min_silence_duration_ms option
- wahnfrieden 1y agoThere are much higher quality VAD solutions available
- DANmode 1y agoPlease name a couple to get someone started who's hacking on webapps? I'd really appreciate it.
- DANmode 1y ago(as would future readers, I'm sure)
- DANmode 1y agohttps://github.com/ten-framework/ten-vad https://github.com/ten-framework/ten-vad
- wahnfrieden 1y agoI last used silero but haven’t kept up with stage of the art so didn’t mention it
- ttflee 1y agoIn Chinese, it always added something like "For study/research purpose only. Please delete after 48 hours." This is what those volunteers added in subtitles of (pirated) movies/shows.
- codedokode 1y agoFair, if AI companies are allowed to download pirated content for "learning", why ordinary people cannot.
- snickerdoodle12 1y agoThere is so much damning evidence that AI companies have committed absolutely shocking amounts of piracy, yet nothing is being done. It only highlights how the world really works. If you have money you get to do whatever the fuck you want. If you're just a normal person you get to spend years in jail or worse. Reminds me of https://www.youtube.com/watch?v=8GptobqPsvg https://www.youtube.com/watch?v=8GptobqPsvg
- 4gotunameagain 1y agoIf you owe the bank $1,000 you have a problem. If you owe the bank $100,000,000 the bank has a problem. We live in an era where the president of the United States uses his position to pump crypto scams purely for personal profit.
- kyleee 1y ago10% for the big don
- alphan0n 1y agoNo one (in the US) has been jailed for downloading copyrighted material.
- snickerdoodle12 1y ago
- xigoi 1y agoPray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?
- madcaptenor 1y agoI am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.
- Workaccount2 1y agoHaving zero exposure to any form of computation for your entire life, as the vast majority of people in the early 19th century were.
- JonChesterfield 1y agoWhat's the defence for the current population?
- immibis 1y agoI can. He was asking if Babbage was cheating. You put in 2+2 - the right figures. The machine says 4 - the right answer. If you put in the wrong figures, like 3+3, will the machine still say 4? It's easy to make a machine that always says 4. The people who asked him that question, however, probably got a different scam demonstrated to them every every. Remember the Mechanical Turk? Babbage's reply paints him very honestly. It shows that he couldn't even conceive that someone might try to trick the royal court (or whoever it was) into accepting a fake device.
- deleted 1y ago[deleted]
- mmcwilliams 1y agoSimilar in the English model. Pretty clear they trained on YouTube videos where creators will put that in otherwise silent sections to ensure it shows up for people with CC on.
- probably_wrong 1y agoThe number one hallucination in my transcriptions was "Subtitles by the Amara.org community".
- indrora 1y agoWhen YouTube began building automatic transcriptions for captions, it regularly flagged any noise or music -- typically industrial noise -- with "[foreign]" If it couldn't understand it, it was "foreign" for the longest time.
- stndef 1y agoYeah, I can confirm seeing that a fair bit specifically during non-verbal parts of videos when someone is using a tool.
- TurkTurkleton 1y agoCan confirm as well, although to my recollection it just shows up as if it's a word the transcription model heard, not "[foreign]" in brackets like with "[Music]" or "[Applause]". It's especially weird to me because I recall the auto-transcriptions being reasonably serviceable when they first rolled them out, only to degrade over time to the point where it was hallucinating the word "foreign" and dropping letters from words or using weird abbreviations (like "koby" for "kilobyte", "TBTE" for "terabyte", or, most memorably weirdly, transcribing the phrase "nanosecond-by-nanosecond" as "nond by nanc") if it didn't decide it heard another one entirely. I also noticed a couple of months ago that YouTube seems to have quietly rolled out a new auto-transcription model that can make reasonable guesses at where capitalization, punctuation, and sentence boundaries should go. It seems to have degraded even more rapidly than the old one, falling victim to the same kinds of transcription errors. Although the new one has a different hallucination in silence and noise that it wasn't able to classify (which, incidentally, its ability to recognize things like music and applause seems worse than the old one's): where the old model would have hallucinated the word "foreign", the new one thinks it's hearing the word "heat", often repeated ("Heat. Heat.").
- the_af 1y agoHey, Netflix occasionally still puts in its English subtitles "[foreign music]", it always cracks me up.
- 1y ago
- tonyhart7 1y agolmao
- philipwhiuk 1y ago> I suspect they trained the model on some random YouTube video without carefully picking really useful data. They trained the model on every YouTube video they could, and hoped the aggregate was useful data.
- danirod 1y agoThis is totally happening with other models too, at least with Spanish. Many transcriptions will end with something that roughly translates to "Thanks for watching!" even if it's never present in the original audio.
- PhasmaFelis 1y agoThis reminds me, some years ago as Google was expanding its translation service, someone tried translating text into and out of an obscure African language (don't recall which) and it always came out as weird Biblical-sounding semi-gibberish. My revelation was that machine translation needs a corpus of bilingual documents to learn from, and if the language is sufficiently obscure, there may not be any bilingual documents except for the Bible, which missionaries have translated into just about every language on Earth.
- horseradish7k 1y agooh yeah this happens a lot on reddit on videos in foreign languages