Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
24
datasets available to search
ShareScore release 0.9.0
Dataset results
24 results for “phoneme”
Transfer of sensorimotor learning reveals phoneme representations in preliterate children - Dataset
<p>This file provides formants values in each speaker and for each trial of the experiment described in the article : Transfer of sensorimotor learning reveals phoneme representations in preliterate children.</p> <p> </p>
CLDF dataset with phoneme inventories from the "Journal of the IPA", aggregated by Baird et al. (2021)
<p>Cite the source of the dataset as:</p> <blockquote> <p>Baird, L., Evans, N., & Greenhill, S. J. (2021). Blowing in the wind: Using 'North Wind and the Sun' texts to sample phoneme inventories. Journal of the International Phonetic Association, 1–42. doi:10.1017/s002510032000033x</p> </blockquote>
Raw and post-processed data for the microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
<p>The current dataset consists of three main folders:</p> <ul> <li><strong>01-Stimuli/</strong>: Contains the three sets of noises (white noise, bump noise, MPS noise) for the 12 study participants (S01 to S12).</li> <li><strong>02-Raw-data/fastACI/</strong>: Contains the raw data as obtained for each participant, which are also available within the GitHub repository of the fastACI toolbox, using the same directory tree. The results for each (anonymised) participant (under: <strong>publ_osses2022b/data_SXX/1-experimental_results/</strong>) include their audiometric thresholds (folder: <strong>audiometry</strong>), the results for the Intellitest speech test (folder: <strong>intellitest</strong>), and for the phoneme-in-noise test /aba/-/ada/ for the three noises (savegame files in MAT format).</li> <li><strong>02-Raw-data/ACI_sim/</strong>: Contains the raw data as obtained for the artificial listener, i.e., the model osses2022a.m (available within the fastACI toolbox). Twelve sets of simulations (using the waveforms of participants S01 to S12) were run for the three types of test noises. The results of the simulations of the phoneme-in-noise test are stored in the savegame MAT files. The template derived from 100 repetitions of /aba/ and /aba/ at an SNR=-6 dB in white noise is also included (template-osses2022a-speechACI_Logatome-abda-S43M-trial-1-v1-white-2022-7-15-N-0100.mat). The same template was used in all simulations.</li> <li><strong>03-Post-proc-data/ACI_exp/</strong>: Auditory classification images (ACIs) derived from the participants' data (folder: <strong>ACI_exp</strong>) and from the simulations (folder: <strong>ACI_sim</strong>). For each participant (or artificial listener) there are three ACIs (MAT files) for each of the corresponding noises. Cross predictions are also included with performance predictions across 'participants' (Crosspred.mat, 12 cross predictions for each noise) or across 'noises' (Crosspred-noise.mat, 3 cross predictions for each participant). The cross predictions all have the same names but are stored in dedicated directories.</li> </ul> <p><strong>Use these data:</strong></p> <ol> <li>Download all these data, place them in a local directory of your computer. If you have MATLAB and you downloaded a local copy of the fastACI toolbox (open access at: <a href="http://github.com/aosses-tue/fastACI">GitHub</a>) you can recreate the figures of our paper.</li> <li>After initialising the toolbox (type 'startup_fastACI;', without quotation marks in MATLAB) and then type either of the following commands, to recreate the figure you want. To recreate the figures in the main text:</li> </ol> <pre><code class="language-javascript">publ_osses2022b_JASA_figs('fig1','zenodo'); publ_osses2022b_JASA_figs('fig2a','zenodo'); publ_osses2022b_JASA_figs('fig2b','zenodo'); publ_osses2022b_JASA_figs('fig3','zenodo'); publ_osses2022b_JASA_figs('fig4','zenodo'); publ_osses2022b_JASA_figs('fig5','zenodo'); publ_osses2022b_JASA_figs('fig6','zenodo'); publ_osses2022b_JASA_figs('fig7','zenodo'); publ_osses2022b_JASA_figs('fig8','zenodo'); publ_osses2022b_JASA_figs('fig8b','zenodo'); publ_osses2022b_JASA_figs('fig9','zenodo'); publ_osses2022b_JASA_figs('fig9b','zenodo'); publ_osses2022b_JASA_figs('fig10','zenodo');</code></pre> <p>To generate the figures of the supplementary materials (Appendix in the BioRxiv preprint):</p> <pre><code class="language-javascript">publ_osses2022b_JASA_figs('fig1_suppl','zenodo'); publ_osses2022b_JASA_figs('fig2_suppl','zenodo'); publ_osses2022b_JASA_figs('fig3_suppl','zenodo'); publ_osses2022b_JASA_figs('fig3b_suppl','zenodo'); publ_osses2022b_JASA_figs('fig4_suppl','zenodo'); publ_osses2022b_JASA_figs('fig4b_suppl','zenodo'); publ_osses2022b_JASA_figs('fig5_suppl','zenodo'); publ_osses2022b_JASA_figs('fig5b_suppl','zenodo');</code></pre> <p><strong>References:</strong></p> <ul> <li><strong>Preprint</strong>: Alejandro Osses, Léo Varnet. "A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception." BioRxiv.</li> <li><strong>fastACI toolbox</strong>: Alejandro Osses, Léo Varnet. fastACI toolbox: the MATLAB toolbox for investigating auditory perception using reverse correlation (v1.2). Zenodo. doi:<a href="https://doi.org/10.5281/zenodo.7314014">10.5281/zenodo.7314014</a>. Supplement to: <a href="http://github.com/aosses-tue/fastACI/tree/v1.2">https://github.com/aosses-tue/fastACI/tree/v1.2</a></li> </ul>
Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns"
<p>The .pdf file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> </ul> <p> </p> <p>The .zip files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p> </p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots. </li> </ul>
Duhumbi Phonology: Origin of non-native and marginal phonemes
<p>These files present the origins of the Duhumbi non-native and marginal consonant and vowel phonemes a supplementary material to section 2.2.1, 2.2.2 and 4.10.1 and 4.10.2 of the Duhumbi grammar.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for commercial purposes <em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration & payment for access, or sites that rely on advertisement (including YouTube) </em>is <strong>not</strong> permitted without <strong>specific written consent</strong> from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on <a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim (at) gmail (dot) com</p>
Changes in neuronal representations of phonemes in the ascending auditory system and their role speech recognition
<p>This dataset comprises neural responses to a set of speech sounds from several brain regions. Auditory nerve data was simulated using a computational model of the auditory nerve. Also included are multi-unit extracellular recordings or responses to the same stimuli in the inferior colliculus and auditory cortex of anaethetised guinea pigs.</p>
Eesthetic: Estonian Paradigms in Phonemic Notation
<p>Eesthetic is a collection of Estonian verbal and nominal paradigms, in phonemic and orthographic notation. They are suited for both computational and manual analysis. The dataset conforms to the Paralex standard</p>
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)
<p>Datasets from Main study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)
<p>Datasets from Pilot study of <strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong> (project linked here: https://doi.org/10.17605/OSF.IO/S6VDN).</p>
Data from: Neural tracking of syllabic and phonemic time scale
Open the record for dataset details and reuse information.
Re-evaluating phoneme frequencies: Supplementary materials S6–S7
<p>Supplementary materials (Sections S6–S7) for the paper <em>Re-evaluating phoneme frequencies</em><em> </em>(<a href="https://www.frontiersin.org/articles/10.3389/fpsyg.2020.570895/full">Macklin-Cordes & Round, 2020</a>)</p> <p>These materials include:</p> <ul> <li>S6: Data viewer app</li> <li>S7: Code, data and results</li> </ul> <p>Usage instructions for the code and app can be found in S1 of the paper's <a href="https://www.frontiersin.org/articles/10.3389/fpsyg.2020.570895/full#supplementary-material">Supplementary Materials</a> and in the readme file contained here.</p>
Jingju Phoneme Classification Features for EUSIPCO 2017 paper
<p>This dataset contains the pre-computed Mel-bands features from the training part of </p> <blockquote> <p>Rong Gong, Rafael Caro Repetto, & Yile Yang. (2017). Jingju a cappella singing dataset part1 [Data set]. Zenodo. http://doi.org/10.5281/zenodo.344932</p> </blockquote> <p>This dataset is used for reproducing the singing voice phoneme classification experiment described in the following paper:</p> <blockquote> <p>Timbre Analysis of Music Audio Signals with Convolutional Neural Networks </p> </blockquote> <p>For the usage and code, please refer to https://github.com/ronggong/EUSIPCO2017 </p>
Phoneme Based Embedded Segmental K-Means for ZeroSpeech2017 Track 2
<p>Our submission applying the phoneme based embedded segmental k-means model to the ZeroSpeech2017 challenge track 2. </p> <p>This is a preliminary version. More details can be found here: <a href="https://www.kamperh.com/papers/bhati+kamper+murty_icassp2018.pdf">https://www.kamperh.com/papers/bhati+kamper+murty_icassp2018.pdf</a> and https://github.com/Saurabhbhati/recipe_zs2017_track2</p> <p> </p>
Dataset for Interspeech 2018 submission: Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions
<p>This dataset contains the materials for training, testing the joint and HSMM models mentioned in the paper "<em>Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions"</em>.</p> <p>The filename list of this dataset can be found in the function <em>get_train_test_recordings_joint()</em> of <em>./general/trainTestSeparation.py</em> file. The dataset contains the Praat TextGrids and .wavs of the variables: <em>train_primary_school, val_primary_school</em> and <em>test_primary_school</em>. For accessing other datasets such as <em>train_nacta_2017, train_nacta</em> and <em>train_sepa</em>, please download them from the links:</p> <p>jingju dataset part1: <a href="https://zenodo.org/record/1185154">https://zenodo.org/record/1185154</a></p> <p>jingju dataset part2: <a href="https://doi.org/10.5281/zenodo.842229">https://doi.org/10.5281/zenodo.842229</a></p> <p>Once you have downloaded these three datasets, you need to set the paths in <em>./general/filePathShared.py</em>.</p> <p>Set <em>path_jingju_dataset</em> to the parent path of these three datasets.</p> <p>Set <em>primarySchool_dataset_root_path</em> to the path of the interspeech2018 dataset (the current dataset).</p> <p>Set <em>nacta_dataset_root_path</em> to the path of the jingju dataset part1.</p> <p>Set <em>nacta2017_dataset_root_path</em> to the path the jingju dataset part2.</p> <p>For more information on this paper, please refer to the Github page: <a href="https://github.com/ronggong/interspeech2018_submission01">https://github.com/ronggong/interspeech2018_submission01</a></p> <p> </p>
Repository of speech features from speakers with and without Parkinson's Disease. Neurovoz - Rasta PLP - V2 - Scientific Reports Publication: Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease
<p>This repository contains the Rasta-PLP features of six different speech recordings (sentences) from Neurovoz corpus (47 parkinsonian and 32 control speakers whose mother tongue is Spanish Castillian.)<br> Number of PLP coefficients: [6, 8, 10, 12, 14, 16, 18, 20].<br> Delta coefficients: Yes<br> Delta Delta coefficients: Yes<br> Sampling rate: 16 kHz<br> Frame size: 15 ms<br> Frame overlapping: 50%</p> <p>This subset of the Neurovoz corpus was recorded between 2015 and 2017 by Universidad Politécncia de Madrid and Hospital General Universitario Gregorio Marañón.</p> <p>This version includes the same files as the previous version and information about UPDRS, H&Y, years since diagnosis and age of each participant.</p> <p>The sentences were:</p> <p>BARBAS: "Cuando las barbas de tu vecino veas pelar, pon las tuyas a remojar"</p> <p>CALLE: "De la calle vendrá quien de tu casa te echará"</p> <p>DIABLO: " Cuando el diablo no sabe qué hacer, con el rabo mata moscas "</p> <p>PETACA BLANCA: " La petaca blanca es mía"</p> <p>PIDIO: "No pidas a quien pidió ni sirvas a quien sirvió"</p> <p>SOMBRA: " El que a buen árbol se arrima, buena sombra le cobija "</p> <p> </p> <p>How to cite:<br> [1] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Grandas-Perez, F., Shattuck-Hufnagel, S. Yagüe-Jimenez, V., and Dehak, N. (2019). Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson’s disease.Scientific reports 9, 19066.</p> <p><br> [2] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Villalba, J., Rusz, J., Shattuck-Hufnagel, S. and Dehak, N. (2019). A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing. Biomedical Signal Processing and Control, 48, 205-220.</p> <p>BibTeX:</p> <pre><code>@article{moro2019phonetic, title={Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge A. and Godino-Llorente, Juan I. and Grandas-Perez, Francisco and Shattuck-Hufnagel, Stefanie and Yague-Jimenez, Virginia and Dehak, Najim}, journal={Scientific Reports}, volume={9}, pages={19066}, year={2019}, publisher={Nature Research Publishing} } @article{moro2019forced, title={A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge Andres and Godino-Llorente, Juan Ignacio and Dehak, Najim}, journal={Biomedical Signal Processing and Control}, pages={205--220}, volume={48}, year={2019}, publisher={Elsevier} } </code></pre> <p> </p>
VeLePor: European Portuguese Verbal Paradigms in Phonemic Notation
<p>This is a collection of European Portuguese verbal paradigms, in phonemic notation. They are suited for both computational and manual analysis.</p>
Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns II: Various Groups of Phonemes"
<p>The .pdf file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> <li>additional plots</li> </ul> <p> </p> <p>The .zip files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p> </p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots. </li> </ul>
Data for HPN-DREAM analysis using OBRWR and PHONEMeS
<p>Data for the HPN-DREAM analysis.</p>
Data from: Experimental evidence for phonemic contrasts in a nonhuman vocal system
Open the record for dataset details and reuse information.
Correlation Between Respiratory Impairment and Phonemes Alteration in Dystrophinopathy Patients With Respiratory Failure
ClinicalTrials.gov study NCT02411370. IPD Sharing: Not stated. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.