The latest versions of Common Voice datasets, Scripted Speech v27.0 and Spontaneous Speech v5.0, are available for download on the Mozilla Data Collective! A giant thank you to all of our contributors for their hard work putting together these datasets.
What’s new in this release:
Version 27.0 of Scripted Speech includes datasets for 295 languages. These datasets contain 32,349,999 voice clips, making approximately 42,593 hours of speech data available for developing and improving speech technology.
Version 5.0 of Spontaneous Speech includes datasets in 80 languages. These datasets contain 89,754 voice clips of free-form answers to questions, of which 302 hours have been transcribed and validated.Since the last release, the Common Voice community has welcomed 1 language to Scripted Speech - Pa’O (blk) - and 6 languages to Spontaneous Speech - Chinese (China, zh-CN), Swahili (sw), Palauan (pau), Sundanese (su), Bengali (bn) and Shan (shn).


We are in the middle of adding our language, Achomi to CV :)
I wish CV had some cool badges so we could show off with them :)))))
Alright I created an issue on the repo: https://github.com/common-voice/common-voice/issues/5525