Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
punchingwater
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
punchingwater
8y ago
Audiobooks are definitely possible for ASR training. Indeed the largest open ASR training dataset before Common Voice was LibriSpeech ( http://www.openslr.org/12/ ). Also note, the first release of Mozilla's DeepSpe
2.
▲
by
punchingwater
8y ago
the article states: > “But Reitze counters that the complete data from that first run is already available online. According to Shoemaker, this includes the relevant time series data and the programs used, but "it's not a trivi
3.
▲
by
punchingwater
8y ago
Just to add my two cents (I work for Mozilla on Common Voice): without help from linguists, Common Voice would have made some very different and very bad decisions about all sorts of things like: accents, dialect segmentation, corpus curati
4.
▲
by
punchingwater
8y ago
We have no plans to allow users to download the "raw" data from s3 (ie. before we perform the train/dev/test split). But we want to eventually build some tools to automate this. See here for some background: https:/
5.
▲
by
punchingwater
8y ago
> The other thing is that it's very cool to see the "you helped us reach out x% goal" thing but it locks up all the previous / next shortcuts which means I have to switch back to the mouse after 5 entries. We have a w
6.
▲
by
punchingwater
8y ago
Just to note, we will never require your email address to contribute. There will always be an anonymous contribution workflow. But adding new languages to Common Voice is a bit complicated at the moment, and we haven't built a way to
7.
▲
by
punchingwater
8y ago
Thank you for bringing this up. Indeed Common Voice is not for everyone. We try to make it clear in our Privacy Policy [1] what pieces of data we collect and why. We do not publish email address or names with the data, and we even strip spe
8.
▲
by
punchingwater
8y ago
In the early days of this project, before we shipped the website (ie. ~March of 2017), we did some explorations around Mechanical Turk. The problem with the Mech Turk approach is that for recording your voices you need a lot of different
9.
▲
by
punchingwater
8y ago
There has been some discussion around this, but no real movement yet: https://github.com/mozilla/voice-web/issues/336
10.
▲
by
punchingwater
8y ago
Would you mind filing an issue? https://github.com/mozilla/voice-web/issues
11.
▲
by
punchingwater
8y ago
> Also... (too lazy to check right now) - if I create an account, can I see the 'yes/no' ratings of my own submissions? Not yet, but this is something in the works. You can explore our new experience with the evergreen lin
12.
▲
by
punchingwater
8y ago
We do have a issue filed to allow users to tag recordings with certain metadata, like noisy or male/female voice. https://github.com/mozilla/voice-web/issues/814 It is something we are still working on.
13.
▲
by
punchingwater
8y ago
We also keep the README in the repo: https://github.com/mozilla/voice-web/blob/master/docs/corpus...
14.
▲
by
punchingwater
8y ago
We used some of the research around Mechanical Turk to find best practices for limiting trolling (e.g. [1]). Our approach thus far has been the two-thirds rule: if two out of three people say the clip is good/bad, we trust that. Also n
15.
▲
by
punchingwater
9y ago
Don't forget Kaldi! https://github.com/kaldi-asr/kaldi
16.
▲
by
punchingwater
9y ago
Thank you so much! I also want to emphasize the importance of listening (validating) as well as recording. Validation is an big part of the puzzle for building machine learning viable data.
17.
▲
by
punchingwater
9y ago
Yup, this is an excellent point. We have, and will continue to explore ways to allow Common Voice users to speak more organically (for instance by answering a question, or responding free-form to some other sort of prompt). The problem with
18.
▲
by
punchingwater
9y ago
Noted. Again thanks for the feedback :)
19.
▲
by
punchingwater
9y ago
This is a bug with our website [1]. We actually are trying to collect non-native speakers (as well as native). We are looking into clarifying this on the site. 1.) https://github.com/mozilla/voice-web/issues/2
20.
▲
by
punchingwater
9y ago
Good question. Sounds like we should add an "Other" to that drop down, and make it clear that we are looking for all accents?
21.
▲
by
punchingwater
9y ago
Exactly! Part of the goals of Common Voice is to make voice recognition work better for non-north american men (which is where the vast majority of the training data comes from). If you are a non-native speaker, we need your voice!
22.
▲
by
punchingwater
9y ago
Sorry about the 503s! We were adding servers to our cluster to handle the hacker news load, and a few 503s are hard to avoid. If this is consistently happening for you, please file a bug and we'll look at it. https://github.
23.
▲
by
punchingwater
9y ago
Great feedback, we can look into clarifying on our homepage that our entire goal is to create a dataset in the public domain. We want people to donate not just to Mozilla, but to the world :)
24.
▲
by
punchingwater
9y ago
Common Voice is only about collecting a large public database of voices. We do have a separate project around speech-to-text [1]. We haven't done much work around speaker recognition (AFAIK) or voice synthesis, but they are both very i
25.
▲
by
punchingwater
9y ago
thanks for the vote of confidence! yes we will absolutely open this data up, and it's just a matter of collecting enough data to be useful, and then building the UI. we have a goal of achieving this by the end of 2017, so stay tuned!
26.
▲
by
punchingwater
9y ago
I can tell from your comment (and it's responses) that the language on our homepage is a bit confusing, so thank you for the feedback. To answer you question: Common Voice is about building a collection of labelled voice data (ie. sent
27.
▲
Snowden and Fareed Zakaria Debate on Govt's Right to Access to Encrypted Devices
(debatesofthecentury.org)
3 points
by
punchingwater
10y ago
|
0 comments
28.
▲
by
punchingwater
11y ago
> (quote from imurdock's twitter) Maybe my suicide at this, you now, a successful business man, not a NIGGER, will finally bring some attention to this very serious issue. This sounds a lot more like 4chan trolls than it does one of