8 ms·
Show HN: Cleanvoice – Automated Podcast Editing
- autoencoders 5y agoHey HN! I like podcasting, but I hate editing them. I tend to stutter and have a lot of filler words in my podcast. That's why I created Cleanvoice, in order to spend less time editing them. Cleanvoice is an ML tool which removes filler words, mouth sounds, stuttering and dead air from your podcast. To use it, just upload your podcast - wait some minutes - download the cleaned audio. It's still not perfect, but it's at a stage where I can blindly use it on every single one of my podcast. I would love to hear your feedback!
- qmmmur 5y agoWithout giving away your secret sauce, what are your approaches to the cleaning process? Is it a combination of different passes of algos or is it something more generic and "sausage machine-like" like a neural network?
- jwuphysics 5y agoBased on the OP's username, surely one of the deep learning algorithms is a denoising autoencoder, right?
- autoencoders 5y agoThe audio is edited in several phases. It uses different algorithms, but most of them are deep learning based. It is surely overengineered, but as a Data Scientist, ML is the most fun part for me.
- qmmmur 5y agoDo you do any audio segmentation to remove the filler words and such?
- nmstoker 5y agoHow is the latency and, if it's sufficiently low, could this realistically be applied to "nearly live" content? That scenario seems really appealing for conferences, even if it just quietens down the verbal ticks, but I'm guessing if the lag is too great it would get like a bad lip sync issue
- Fogest 5y agoNvidia RTX voice does similar. It's pretty similar to other technology though where it focuses more on removing background noise. It actually works very well. It would definitely be interesting to see it also filter speech itself. But I feel like this would be hard to do without introducing extra latency. If someone is saying "umm" or some other filler before a word you kinda need to know what that word will be to determine if it's filler or not. So it almost can't be done without introducing latency as it would need some future speech to determine if filler or not.
- pessimizer 5y agoTo do this, the speaker would have to wear an EEG cap. You're talking about cutting the mic before a verbal tic happens. With an EEG cap, though, I bet a smart person familiar with the methods could bash something together in a day that would work.
- autoencoders 5y agoTrue. You don't even need a full CAP. Just some channels in the visual cortex. (With more advanced AI) So you would just need to hear a headband or one of those EEG which look more elegant.
- autoencoders 5y agoIt would be a huge engineering endeavour, which I wouldn't be capable of doing. That said, things like background noise and some sounds can be removed. See Krisp.ai
- qmmmur 5y agoIzotope plugins already do some of these things but not all. In particular their de-clicking algorithm is pretty good but definitely not automatic or low latency.
- telesilla 5y agoHave you compared this to other commercial options such as Descript? Looks really great at a glance, thanks for sharing!
- autoencoders 5y agoI tried to use Descript for my podcast, but it has some issues. 1) It doesn't work well if you have a strong accent. As an non-native speaker, the transcription were quite bad, making the editing quite bad. 2) Cleanvoice works with multiple languages, descript doesn't. 3) Cleanvoice can remove stutters (not always, but it tries) and mouth sounds like lip smacking, teeth clicking. Descript can't. This is not a big deal for most, but since I stutter alot this was essential. My approach is different from Descript. They use a transcription service, and then they edit the audio based on the text. I work directly on the phonetics level. Allowing me to have more control over audio. Depending on the needs, either one is better. I guess you should try it for yourself and compare.
- telesilla 5y agoI will try it thanks! We work with lots of accents and we've found the same with Descript that it fails for example with a strong French accent. Translstion is really key for us also, looking forward to seeing systems trained on more accents.
- ckdarby 5y agoI use Descript and it is absolutely lovely. There are a bunch in this space that I would not be surprised being merged or acquired. Would love to see Descript & GetWelder merging together. While Cleanvoice has some niche features that Descript doesn't offer I would not be surprised to find them rolling these features out in the next major release they're doing. IMO the founder of Cleanvoice should sell/join Descript.
- wpietri 5y agoNeat! I love products that come out of a personal need. Is it possible for you to do a live, personal demo? No logins or anything. I'm thinking something where you tell people to start up their audio and then give them a quick prompt like "Describe your breakfast yesterday." Record for 30 seconds, and then let them play back the original and cleaned versions. You could limit them to, say, 5 goes, with a different prompt each time. I suggest it because a) a little personal investment makes it more likely they'll give you their email address for signing up, and b) many potential customers underestimate how much they need something like this.
- autoencoders 5y agoI like your idea, makes sense. My biggest fear is that without login, people will start abusing it in ways that I don't expect. Definitely considering it. Thanks you!
- wpietri 5y agoThat's a good fear to have. That's the kind of thing I would set up some monitoring for and then wait to see. You might get a few jerks. But those same jerks might also be the sort of people who would sign up with a bunch of fake emails, so gating on an email address may not be much better than gating on a fresh-issued cookie. Thanks for listening, and good luck with your project!
- undoware 5y agoI literally just bought your product, thank you very much, I needed this and wondered why no one had made it yet.
- autoencoders 5y agoI appreciate it! If you have any issues or need help, feel free to reach out. (You can use the chat in the app.)
- gus_massa 5y agoIs the example in the page really made by the computer? In my opinion the pauses in where the filler words were are slightly too long. Is it possible to configure this? Is it possible to keep some filler words? I make something similar (but not professionally), and sometimes I like too keep a few of them.
- autoencoders 5y ago> Is the example in the page really made by the computer? Yes. >In my opinion the pauses in where the filler words were are slightly too long. Is it possible to configure this? I agree, however, if you use it in an interview. The edits sound better. In an unnatural setting, you get unnatural results. Currently, there is no way to set it for now. But customization is planned for Q2 next year. >Is it possible to keep some filler words? For now no, but keeping some filler sounds to keep it authentic is something which I plan.
- gus_massa 5y agoI agree that the correct length of the pause after the word is removed is very tricky. Perhaps your configuration is the better than my imaginary magical edition. In other comment, eganist posted a link to https://cleanvoice.ai/integrations https://cleanvoice.ai/integrations It looks interesting because I can choose which to keep and even use it to sink with video [with some additional work]. I didn't see it the first time in the page.
- autoencoders 5y agoADL Support is also around Q2, so you could just import it in your audio/video editor without issue. Thank you point out. I'll put Integrations on the homepage as well.
- deleted 5y ago[deleted]
- axhl 5y agoCongratulations on launching. How are you finding using termly.io for the legal side of things?
- autoencoders 5y agoIt's not ideal. See the comment talking about the terms. I have a meeting with a lawyer soon. But I guess is better than no terms.
- fareesh 5y agoWhat's the high level approach required to build something like this yourself? Does it involve relying on speech to text with timestamps and then a series of cuts based on that?
- spicybright 5y agoI'm going to sound like a negative nancey, but I wish podcasters/youtubers would just practice their speaking skills instead of rely on series of really quick jump cuts. Worst offenders are those that can't get through a sentence without splicing it 2+ times... Perhaps you could have a mode to detect how much one stutters, and parts worth redoing without spending as much time combing the whole thing.
- cube00 5y agoEspecially ones who won't set their background LED lights to a stable color. The smooth flowing gradient becomes very distracting when you jump cut the heck out of it.
- intrasight 5y agosynesthesia?
- pfortuny 5y agoClassically professionals learnt their discourses by heart. That stands out when you see it. I remember fondly a student of mine who seemed unable to express himself properly. I told him to memorize his final project dissertation because otherwise it would be a wreck (OK, I did not say this last part, it was more of a suggestion). BOY: did he memorize it. He got an honors and I did think “this guy has really done it, and it sounds like music!” When you do it well, it tells.
- intrasight 5y agoSome podcasts I listen to are over-edited. I'd always assumed that a) it was done manually and b) it was done to keep the length below some threshold. Now I'm curious if they are using software to automate the editing. I find the cadence very unnatural when all the spaces between phonemes are removed.
- ghaff 5y ago>I find the cadence very unnatural when all the spaces between phonemes are removed. Any editing can be overdone and, while I do a modicum of editing out umms, you knows, and other verbal ticks when I'm putting together a podcast interview, I'm not fanatical about it. You do occasionally get someone who just speaks quite slowly and it is sort of annoying to listen to as audio. So I've done some automated gap reduction is a couple cases.
- notafraudster 5y ago"Free 30 Minutes Trial" is not native English. "Free 30 Minute Trial" would be better; but I think the sentence is a little confusing. I presume you mean you can convert 30 minutes of audio for free, not that the trial account is only valid for 30 minutes from creation. I would do "Clean 30 minutes of audio for free. No Credit Card needed." or similar. The sale page which says "Get 30 minutes credit to try the service out." is better, and "30 minutes" does sound correct on that page. In your FAQ, you say: "Currently we remove lip smacks, saliva crackle, mouth clicks and harsh parts of breathing (not the whole breath). If you want to remove a particular mouth sound (ex. Chewing), write us in the chat as a feature request." I don't think most English speakers would understand what "harsh parts of breathing" are. Typically a parenthetical example in English would be written "(e.g. chewing)" not "(ex. Chewing")". Your question "What filetype and sizes do you support?" doesn't answer what filetypes you support, and I suspect the singular "filetype" was a grammar error. You also write "We have an audio file size limit of 1.5G per file or in case you are uploading multi-track and a total file size of 2 GB. ". The part that says "or in case you are uploading multi-track and" doesn't make any sense in English. I think you mean "We support file sizes up to 1.5GB per file for single-track files, or 2GB if you are uploading a multi-track file as separate files." but I'm not sure. In general I don't understand why each selling point has a separate FAQ page but the FAQs are often not related to the selling point. I don't think people think the "Mouth Sound Remover" page is the one that lists file size support, while the "Stutter Remover" page is the one that lists the maximum number of tracks per project. Your integrations page lowercases "cleanvoice" whereas other pages write it as "Cleanvoice". Under integrations, you have a section called "Markers Export". This should probably be "Export Markers" or "Marker Export". Under "How to Export Edits", you probably don't want to capitalize "Results" or "Editor" unless these are supposed to be title cased, in which case you probably want to title case all of them. Under your pricing FAQ you have "Does my credit expire at end of the month? Your credit will reset every billing month. Unused credit will be lost." This is needlessly confusing. You use the verbs "expire", "reset", and "be lost" to describe the same thing, and you don't actually answer the question. Also you don't want "at end of the month", you want "at month's end" or "at the end of the month". I would rewrite as "Does my credit expire at the end of each month? Yes. Credit resets every month and cannot be carried over to future months. Unused credit will be lost." This is a terrible business model, though, and so I suggest you not do this. Either sell as a subscription or sell as a credit model, not both, this is gross. In general I think you want to pay someone who is a professional English copywriter to fix your website. Cheers. Edit: I just noticed your changelog is powered by a service called Headway. I am not sure if you also made Headway, but Headway's website is also in need of English copyediting.
- hs86 5y agoReminds me of https://auphonic.com/ https://auphonic.com/ Their pricing is also similar, but Auphonic allows both subscription and prepaid "credits".
- autoencoders 5y agoYes, the idea is to bring also prepaid credits soon. Auphonic and Cleanvoice go well together. I guess the idea is to have your podcast edited by Cleanvoice and then the audio post-processing with Auphonic.
- ghaff 5y agoAuphonic's volume equalization is almost a must-have for podcasts. I used to spend a lot of time getting volumes right. With Auphonic it's quick and easy. I definitely prefer pre-paid credits to a subscription given my podcast production varies a lot.
- throwaway1777 5y agoOvercast has features to do some of this on the listener side. I prefer having the AI on the listener side so I can go back to the raw version if the AI messes up for some reason.
- monroewalker 5y agoSounds similar to Descript https://www.descript.com/ https://www.descript.com/
- erichdongubler 5y agoI'm a very happy user of the free tier of Descript right now, and will definitely pay once my transcription limit is reached. It seems like this particular product might do a better job of the automated editing specifically, but Descript has a ton of other features (speaker identification, transcription, real-time editing based on text edits, asset management, and uploads), and I definitely wouldn't trade them for marginally better auto-removal of noise and filler. Does anybody who develops Cleanvoice have any commentary here?
- autoencoders 5y agoAdrian from Cleanvoice here. Before building Cleanvoice, I tried to use Descript for my podcasts. 1) It doesn't work well if you have a strong accent. As an non-native speaker, the transcription were quite bad, making the editing quite bad. 2) Cleanvoice works with multiple languages, descript doesn't. 3) Cleanvoice can remove stutters (not always, but it tries) and mouth sounds like lip smacking, teeth clicking. Descript can't. This is not a big deal for most, but since I stutter alot this was essential. However, if none of these apply for you. There is no reason to change from Descript.
- geuis 5y agoYour demos don’t play on iOS safari.
- autoencoders 5y agoUps! Thank you for point it that out. I'll check it.
- mijustin 5y agoHey! Justin (from Transistor.fm) here. This looks really interesting. Two questions: 1. Any plans for an API and bulk pricing? 2. Any plans to add loudness normalization, balancing, etc to the processing?
- autoencoders 5y agoHey Justin! Love your podcast. 1) API Access will come end of Q1. 2) In the next 6 months, No. However, Auphonic would be a good fit for you.
- bozhark 5y agoSo my current podcast stack looks like Zencastr RECORD Transistor.fm HOST but looks like I'll be adding Cleanvoice AI CLEAN UP Auphonic AI EQ any other suggestions for optimum output from multiple inputs all over the world?
- mikepechadotcom 5y agoReally cool project, I wish you great success! Could be useful for my (german) podcast agency! Out of curiosity: Which ai-technology did you use? OpenAI? Google API? Or did you train the models yourself with Python (sth. like Tensorflow)? Cheers, Mike
- autoencoders 5y agoHallo Mike, freut mich dich kennenzulernen! I trained my own models. No OpenAI/Google API. Liebe Grüße, Adrian
- stavros 5y agoThis is excellent, well done! I'd be curious to know how it's done, as I don't know much about deep learning and this looks like magic to me.
- arendtio 5y agoThat logo is very similar to the Cisco logo: https://www.cisco.com https://www.cisco.com
- pwned1 5y agoI suspected something like this was happening with podcasts. I've noticed lately that some podcasters have unnaturally short pauses between speakers (question and answer) or between sentences. It really annoys me. It makes it almost unlistenable.
- carols10cents 5y agoYes, the worst is when so much silence is removed that it sounds like someone is laughing over themselves.
- autoencoders 5y agoI agree, as if they don't breathe! This is not the case with my app. I keep the edits longer than shorter, since I also find that unlistenable.
- sdoering 5y agoNot sure were you are located, but if you are giving access to people protected by the GDPR your cookie notice does not fullfill the requirements set by European Regulations. Additionally, if you are located in a country that (like Germany for example) has regulations on the necessity of an imprint, this might also be missing.
- autoencoders 5y agoIt should be ok, since I use strictly essential cookies, which don't require consent. (But users need to be informed) Or do I misunderstand the law? [1] Strictly necessary cookies — These cookies are essential for you to browse the website and use its features, such as accessing secure areas of the site. Cookies that allow web shops to hold your items in your cart while you are shopping online are an example of strictly necessary cookies. These cookies will generally be first-party session cookies. While it is not required to obtain consent for these cookies, what they do and why they are necessary should be explained to the user. [1] - https://gdpr.eu/cookies/ https://gdpr.eu/cookies/
- nateweiss 5y agoLooks cool! Would this also work for "explainer" type videos, showing how to use a software product or similar? If yes, you might consider a page or callout about that use-case, as it might attract some additional users. Just a thought.
- tyingq 5y agoThat seems like it would be tricky, as the video and audio would get out of sync. You would have to remove, then "fill" to keep the timing. Though this product does mention it works with multiple speakers on different tracks...so they are already somewhat in that space.
- autoencoders 5y agoFor video is quite tricky. One thing with Video is that you don't want to over edit the audio, since its then very hard to keep the video synced. That said for explainer video it should work ok, but for a Video Podcast it would be horrible. I have an idea how to deal with this, but this is not now available.
- eganist 5y agoThis is awesome. Can I suggest the ability to export as project files for popular editors for your roadmap? It'd cut professional workflows down substantially, which would be worth an (even higher) upcharge. (It wasn't immediately obvious to me if you already did this) Edit: https://cleanvoice.ai/integrations https://cleanvoice.ai/integrations seems pretty close. I'd honestly charge more for integrations and provide a base tier for just exporting sound. I imagine most indie users would benefit from finished exports enough to pay, while project files would command a higher fee from editors looking to speed up their workflow to take more clients. That's where I'm coming from on pricing tiers and upcharging for professional features.
- autoencoders 5y agoADL Support will come around Q2, so you can import it in lot of audio and video editors. For now, we have these export files which you mentioned. Regarding Pricing, that's a good point. I will definitely consider it, thank you!
- tim-- 5y agoTo add to this, it might also be a good feature to output EDL files for video editors. https://www.rev.com/blog/how-to-import-an-edit-decision-list-edl-into-davinci-resolve https://www.rev.com/blog/how-to-import-an-edit-decision-list... This could help when you have for example multiple camera angles, to switch between/do morph cuts (https://helpx.adobe.com/premiere-pro/using/morph-cut.html https://helpx.adobe.com/premiere-pro/using/morph-cut.html) for video interviews.
- autoencoders 5y agoThank you for the suggestion. EDL/ADL is definitely on the list.
- daenney 5y agoThe Terms of service seem worrisome. > By posting your Contributions to any part of the Site or making Contributions accessible to the Site by linking your account from the Site to any of your social networking accounts, you automatically grant, and you represent and warrant that you have the right to grant, to us an unrestricted, unlimited, irrevocable, perpetual, non-exclusive, transferable, royalty-free, fully-paid, worldwide right, and license to host, use, copy, reproduce, disclose, sell, resell, publish, broadcast, retitle, archive, store, cache, publicly perform, publicly display, reformat, translate, transmit, excerpt (in whole or in part), and distribute such Contributions (including, without limitation, your image and voice) for any purpose, commercial, advertising, or otherwise, and to prepare derivative works of, or incorporate into other works, such Contributions, and grant and authorize sublicenses of the foregoing. It sounds an awful lot like "we are allowed to do anything and everything we want with the content you upload to us". Maybe I'm misunderstanding something, but I'd be extremely hesitant to upload any content I create to a service with those kinds of terms.
- autoencoders 5y agoI agree. The terms will be changed. I used an auto-generated Terms generator for now (termly.io) I would like to rewrite it. What I do is just keep your files on the server for a week. In case you have an issue, I will look into your file to fix your issue. And if you want, you can give consent for me to further improve the service. (Say you have an accent which the AI is bad and I can use your audio file to understand why it failed.)
- throwthere 5y agoWith this statement you’ve now shown that your site doesn’t take contracts seriously and opened the door to people arguing future contacts are also invalid. I’d delete this response asap.
- giansegato 5y agoWhy? They can change the policy and ask for a confirmation, as every service out there is already doing.
- xipho 5y agoCan anyone recommend similar for removing ums etc. in videos? IIRC there is a workflow in some professional software, but being able to train and throw the algorithim right at the video itself (especially locally) would be useful.
- mijustin 5y agoYes; Descript.com does this.
- autoencoders 5y agoFor now, Descript would be the best option. You can still make it work with the integrations, but it is a lot of effort. That will change in Q2, when I add support for video.
- nickjj 5y ago> Can anyone recommend similar for removing ums etc. in videos? For single camera floating head style videos where you're continuously talking about 1 topic it's going to be very jarring if you start cutting out filler words. You'll end up with a bunch of jump cuts where it looks like video frames are dropped.
- abdik 5y agoThe logo is similar to ours https://www.lovo.ai/ https://www.lovo.ai/
- bryans 5y agoWhile turning it into a heart may be clever branding, you've only slightly modified a ubiquitous icon representing audio, and countless startups used that before you.
- nickjj 5y agoAs someone who has personally edited over a hundred 1-2 hour podcasts with a new guest every time removing umms, ahhs, dead air and filler words is soul crushing. It has gotten to the point where after 2 years of running my podcast[0] I'm seriously considering stopping the show because I'm getting burnt out from editing and without sponsors it's not feasible to hire an editor, but even with the show making no money I would happily pay triple your asking price if I could click a button and have the problem solved in a way that matched a human's ability to edit out filler words. It really is the difference between being able to edit a 1 hour episode in 1 real life hour (editing at 2x speed) vs literally spending 5 hours to edit 1 hour when there's a lot of filler words or ums. That's due to having to stop every few seconds, think about when to cut it and perform the cut. This is using a heavily optimized keyboard shortcut focused workflow too. I hope you don't mind constructive criticism but in my opinion your "after" version doesn't sound natural. This isn't an attack on your service specifically, because the outcome is the same with all of the automated tools I've tried. I haven't tried them all but I did play with a few of them. For example in your case the pause between "Removing" and "filler" doesn't match the pace of the rest of the sentence and the transition from "very" to "time" has a very hard cut. This is also a 10 word clip that's about 6 seconds. If you listened to a 1 hour podcast episode that was edited like this it would be much more noticeable. There's so many intricate and subtle details around when and what to cut to remove these things in a way where it's not noticeable. Are there any paths moving forward in AI / ML that can lead to this being indistinguishable from being humanly edited? I debated deleting this comment before posting it because it's a combination of feedback but also saying the service isn't something I would buy in its current state but I'd like to think it's more beneficial to post this to show there is a real demand for this service if it can be executed flawlessly. [0]: https://runninginproduction.com/ https://runninginproduction.com/
- autoencoders 5y agoThe edit on the page is not the best. I agree!. Mainly, if your recording is unnatural (like that one) the edit is also unnatural. However, the tool works better in an interview podcast. I would strongly recommend to just upload a sample, and you would see a big difference. Regarding if ML would be indistinguishable from humanly edit. Hard to tell. I think it will be like self-driving cars in the future. 98% edits good 2% bad edits.
- tokamak-teapot 5y ago“The algorithm can also work with accents from other countries, such as Australian ones or Irish.” Other than which country, though? Presumably an English speaking one - UK? New Zealand? Canada? US?