HN Hall of Fame Weekly email

Project Common Voice

voice.mozilla.org Research & data Datasets AI & data Candidate
Screenshot of voice.mozilla.org captured 2026-07-20
Page preview · captured 2026-07-20

The original page was unavailable during review. The primary button opens a verified 2020 archived copy; the original address is retained as provenance.

Resurfaced independently across 2 calendar years, with breakout response in 2 of them.

submissions
7
submitters
7
observed span
2017–2019
peak thread · 57 comments
205 pts
latest 20+ return · 2019-10-16
174 pts

Submission timeline

2007–2026

One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.

First comments on top threads

HN comment order

The terminology is a bit confusing. They are saying that they want to build voice recognition but it seems like they actually might want to build a speech recognition engine. Speech recognition is about recognizing the speech, the spoken words. Voice recognition is about recognizing the speakers voice, i.e. identifying the speaker. Also, maybe they also want to build a text-to-speech (TTS) system but I'm not sure. No matter what, the collected data might be useful for all of that…

The goal of making a large, publicly available training corpus for ASR is incredibly admirable, but this approach is problematic. People speak entirely differently when reading from a script. Models trained on read speech (like LibriSpeech) generally don't perform well on spontaneous speech test sets (like Switchboard). Transcribing speech that was read from a script isn't a particularly interesting problem. This effort would be more interesting if it could collect speech data in a more specific domain, like web…

gok·174-point thread·

The project is a great idea and an open source way to decentralize the diverse training speech record data which previously was available only with big corporations like Google, FB & Amazon. This open data repo created using the common voice project will help startups & researchers to create generalized speech recognition algorithms and compete with other big players. The project is truly an amazing initiative to promote the open innovation on the web. Hope all will participate and contribute…

The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.

Breakout years
2

100+ points or 50+ comments

Total points
396

reference only — not used in Hall rules or ranking

Total comments
107

reference only — not used in Hall rules or ranking

Every submission

DateTitle as submittedByPointsComments
2017-06-21Donate your voice to Project Common Voice. Moz://a is opening voice recognitioninterweb30
2017-06-26Project Common Voicetype010
2017-06-28Project Common Voicem_eiman20
2017-07-16Project Common VoicehashtagMERKY40
2017-07-17Mozilla: Project Common Voiceibotty71
2017-07-18Project Common VoiceFirst breakout · Best threadmhr_online20557
2019-10-16Common Voice – Mozilla's initiative to help teach machines how real people speakLatest 20+ point returnmjlee17449