Skip to main content
zenodoopen

PB2007 French acoustic-articulatory speech database

<p><strong>PB2007 acoustic-articulatory speech dataset</strong></p> <p>Badin, P.,Bailly G., Ben Youssef A., Elisei F., Savariaux C., Hueber T.&nbsp;<br> Univ. Grenoble Alpes, CNRS, Grenoble INP, GIPSA-lab, 38000 Grenoble, France<br> <br> LICENSE:<br> ========<br> This dataset is made available under the Creative Commons Attribution Share-Alike (CC-BY-SA) license</p> <p><br> CREDITS - ATTRIBUTION:<br> ======================<br> If using this dataset, please cite one of the following studies (all of them exploit this dataset)&nbsp;<br> - Ben Youssef, A., Badin, P., Bailly, G. &amp; Heracleous, P. (2009). Acoustic-to-articulatory inversion using speech recognition and trajectory formation based on phoneme hidden Markov models. In Interspeech 2009, vol., pp. 2255-2258. Brighton, UK.<br> - Ben Youssef, A., Badin, P. &amp; Bailly, G. (2010). Can tongue be recovered from face? The answer of data-driven statistical models. In Interspeech 2010 (11th Annual Conference of the International Speech Communication Association) (T. Kobayashi, K. Hirose &amp; S. Nakamura, editors), vol., pp. 2002-2005. Makuhari, Japan.<br> - Hueber T., Bailly G., Badin P., Elisei F., &quot;Speaker Adaptation of an Acoustic-Articulatory Inversion Model<br> using Cascaded Gaussian Mixture Regressions&quot;, Proceedings of Interspeech, Lyon, France, 2013, pp. 2753-2757.&nbsp;</p> <p>&nbsp;</p> <p>DATA FILES DESCRIPTION:<br> =======================<br> /_seq/:&nbsp;<br> &nbsp;&nbsp; &nbsp;Electro-magnetic Articulography data, recorded at 100Hz<br> &nbsp;&nbsp; &nbsp;Sensors :<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR01 : LT_x (lower incisor, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR02 : tip_x (tongue tip, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR03 : mid_x (tongue dorsum, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR04 : bck_x (tongue back, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR05 : LL_vis_x (lower lips, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR06 : UL_vis_x (upper lips, x coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR07 : LT_z (lower incisor, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR08 : tip_z (tongue tip, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR09 : mid_z (tongue dorsum, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR10 : bck_z (tongue back, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR11 : LL_vis_z (lower lips, z coordinate)<br> &nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp;PAR12 : UL_vis_z (upper lips, z coordinate)</p> <p>/_wav16:&nbsp;<br> &nbsp;&nbsp; &nbsp;subject audio signal, synchronized with the EMA data<br> &nbsp;&nbsp; &nbsp;Format: PCA wav, 16kHz, 16bits</p> <p>/_lab: phonetic segmentation using the following set<br> __ (long pause), _ (short pause), a, e^ (as in &quot;lait&quot;), e (as in &quot;bl&eacute;&quot;), i, y (as in &quot;voiture&quot;), u (as in &quot;loup&quot;), o^ (as in &quot;pomme&quot;),x (as in &quot;pneu&quot;), x^ (as in &quot;coeur&quot;), a~ (as in &quot;flan&quot;), e~ (as in &quot;in&quot;), x~ (as in &quot;un&quot;), o~ (as in &quot;mon&quot;), p, t, k, f, s, s^ (as in &quot;CHat&quot;), b, d, g, v, z, z^ (as in &quot;les Gens&quot;), m, n, r, l, w, h, j, o, q (schwa)<br> &nbsp;</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics