Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

24

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

24 results for “phoneme”

Learn how ShareScore rates datasets ↗
zenodo48/100

Transfer of sensorimotor learning reveals phoneme representations in preliterate children - Dataset

<p>This file provides formants values in each speaker and for each trial of the experiment described in the article : Transfer of sensorimotor learning reveals phoneme representations in preliterate children.</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

CLDF dataset with phoneme inventories from the "Journal of the IPA", aggregated by Baird et al. (2021)

<p>Cite the source of the dataset as:</p> <blockquote> <p>Baird, L., Evans, N., &amp; Greenhill, S. J. (2021). Blowing in the wind: Using &#x27;North Wind and the Sun&#x27; texts to sample phoneme inventories. Journal of the International Phonetic Association, 1–42. doi:10.1017/s002510032000033x</p> </blockquote>

opencc-zeroApr 2024View details →
zenodo44/100

Raw and post-processed data for the microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception

<p>The current dataset consists of three main folders:</p> <ul> <li><strong>01-Stimuli/</strong>: Contains the three sets of noises (white noise, bump noise, MPS noise) for the 12 study participants (S01 to S12).</li> <li><strong>02-Raw-data/fastACI/</strong>: Contains the raw data as obtained for each participant, which are also available within the GitHub repository of the fastACI toolbox, using the same directory tree. The results for each (anonymised) participant (under: <strong>publ_osses2022b/data_SXX/1-experimental_results/</strong>) include their audiometric thresholds (folder: <strong>audiometry</strong>), the results for the Intellitest speech test (folder: <strong>intellitest</strong>), and for the phoneme-in-noise test /aba/-/ada/ for the three noises (savegame files in MAT format).</li> <li><strong>02-Raw-data/ACI_sim/</strong>: Contains the raw data as obtained for the artificial listener, i.e., the model osses2022a.m (available within the fastACI toolbox). Twelve sets of simulations (using the waveforms of participants S01 to S12) were run for the three types of test noises. The results of the simulations of the phoneme-in-noise test are stored in the savegame MAT files. The template derived from 100 repetitions of /aba/ and /aba/ at an SNR=-6 dB in white noise is also included (template-osses2022a-speechACI_Logatome-abda-S43M-trial-1-v1-white-2022-7-15-N-0100.mat). The same template was used in all simulations.</li> <li><strong>03-Post-proc-data/ACI_exp/</strong>: Auditory classification images (ACIs) derived from the participants&#39; data (folder: <strong>ACI_exp</strong>) and from the simulations (folder: <strong>ACI_sim</strong>). For each participant (or artificial listener) there are three ACIs (MAT files) for each of the corresponding noises. Cross predictions are also included with performance predictions across &#39;participants&#39; (Crosspred.mat, 12 cross predictions for each noise) or across &#39;noises&#39; (Crosspred-noise.mat, 3 cross predictions for each participant). The cross predictions all have the same names but are stored in dedicated directories.</li> </ul> <p><strong>Use these data:</strong></p> <ol> <li>Download all these data, place them in a local directory of your computer. If you have MATLAB and you downloaded a local copy of the fastACI toolbox (open access at: <a href="http://github.com/aosses-tue/fastACI">GitHub</a>) you can recreate the figures of our paper.</li> <li>After initialising the toolbox (type &#39;startup_fastACI;&#39;, without quotation marks in MATLAB) and then type either of the following commands, to recreate the figure you want. To recreate the figures in the main text:</li> </ol> <pre><code class="language-javascript">publ_osses2022b_JASA_figs('fig1','zenodo'); publ_osses2022b_JASA_figs('fig2a','zenodo'); publ_osses2022b_JASA_figs('fig2b','zenodo'); publ_osses2022b_JASA_figs('fig3','zenodo'); publ_osses2022b_JASA_figs('fig4','zenodo'); publ_osses2022b_JASA_figs('fig5','zenodo'); publ_osses2022b_JASA_figs('fig6','zenodo'); publ_osses2022b_JASA_figs('fig7','zenodo'); publ_osses2022b_JASA_figs('fig8','zenodo'); publ_osses2022b_JASA_figs('fig8b','zenodo'); publ_osses2022b_JASA_figs('fig9','zenodo'); publ_osses2022b_JASA_figs('fig9b','zenodo'); publ_osses2022b_JASA_figs('fig10','zenodo');</code></pre> <p>To generate the figures of the supplementary materials (Appendix in the BioRxiv preprint):</p> <pre><code class="language-javascript">publ_osses2022b_JASA_figs('fig1_suppl','zenodo'); publ_osses2022b_JASA_figs('fig2_suppl','zenodo'); publ_osses2022b_JASA_figs('fig3_suppl','zenodo'); publ_osses2022b_JASA_figs('fig3b_suppl','zenodo'); publ_osses2022b_JASA_figs('fig4_suppl','zenodo'); publ_osses2022b_JASA_figs('fig4b_suppl','zenodo'); publ_osses2022b_JASA_figs('fig5_suppl','zenodo'); publ_osses2022b_JASA_figs('fig5b_suppl','zenodo');</code></pre> <p><strong>References:</strong></p> <ul> <li><strong>Preprint</strong>: Alejandro Osses, L&eacute;o Varnet. &quot;A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception.&quot; BioRxiv.</li> <li><strong>fastACI toolbox</strong>: Alejandro Osses, L&eacute;o Varnet. fastACI toolbox: the MATLAB toolbox for investigating auditory perception using reverse correlation (v1.2). Zenodo. doi:<a href="https://doi.org/10.5281/zenodo.7314014">10.5281/zenodo.7314014</a>. Supplement to: <a href="http://github.com/aosses-tue/fastACI/tree/v1.2">https://github.com/aosses-tue/fastACI/tree/v1.2</a></li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo40/100

Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns"

<p>The .pdf&nbsp;file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> </ul> <p>&nbsp;</p> <p>The .zip&nbsp;files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p>&nbsp;</p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots.&nbsp;</li> </ul>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Duhumbi Phonology: Origin of non-native and marginal phonemes

<p>These files present the origins of the Duhumbi non-native and marginal consonant and vowel phonemes a supplementary material to section 2.2.1, 2.2.2 and 4.10.1 and 4.10.2 of the Duhumbi grammar.</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p> <p>Tim Bodt: bodttim&nbsp;(at) gmail (dot) com</p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

Changes in neuronal representations of phonemes in the ascending auditory system and their role speech recognition

<p>This dataset comprises neural responses to a set of speech sounds from several brain regions. Auditory nerve data was simulated using a computational model of the auditory nerve. Also included are multi-unit extracellular recordings or responses to the same stimuli in the inferior colliculus and auditory cortex of anaethetised guinea pigs.</p>

opencc-by-4.0Aug 2018View details →
zenodo40/100

Eesthetic: Estonian Paradigms in Phonemic Notation

<p>Eesthetic is a collection of Estonian verbal and nominal paradigms, in phonemic and orthographic notation. They are suited for both computational and manual analysis. The dataset conforms to the Paralex standard</p>

opencc-by-4.0Feb 2024View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Main Study Data)

<p>Datasets from Main&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
zenodo40/100

Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore (Pilot Study Data)

<p>Datasets from Pilot&nbsp;study of&nbsp;<strong>Effect of unattended distributional training on phoneme category discrimination in English-Mandarin bilingual adult participants in Singapore</strong>&nbsp;(project linked here:&nbsp;https://doi.org/10.17605/OSF.IO/S6VDN).</p>

opencc-by-4.0Nov 2022View details →
dryad40/100

Data from: Neural tracking of syllabic and phonemic time scale

Open the record for dataset details and reuse information.

publicAug 2024View details →
zenodo36/100

Re-evaluating phoneme frequencies: Supplementary materials S6–S7

<p>Supplementary materials (Sections S6&ndash;S7) for the paper <em>Re-evaluating phoneme frequencies</em><em>&nbsp;</em>(<a href="https://www.frontiersin.org/articles/10.3389/fpsyg.2020.570895/full">Macklin-Cordes &amp; Round, 2020</a>)</p> <p>These materials include:</p> <ul> <li>S6: Data viewer app</li> <li>S7: Code, data and results</li> </ul> <p>Usage instructions for the code and app can be found in S1 of the paper&#39;s <a href="https://www.frontiersin.org/articles/10.3389/fpsyg.2020.570895/full#supplementary-material">Supplementary Materials</a>&nbsp;and in the readme file contained here.</p>

opencc-by-4.0Sep 2020View details →
zenodo36/100

Jingju Phoneme Classification Features for EUSIPCO 2017 paper

<p>This dataset contains the pre-computed Mel-bands features from the training part of&nbsp;</p> <blockquote> <p>Rong Gong, Rafael Caro Repetto, &amp; Yile Yang. (2017). Jingju a cappella singing dataset part1 [Data set]. Zenodo. http://doi.org/10.5281/zenodo.344932</p> </blockquote> <p>This dataset is used for reproducing the singing voice phoneme classification experiment described in the following paper:</p> <blockquote> <p>Timbre Analysis of Music Audio Signals with Convolutional Neural Networks&nbsp;&nbsp;</p> </blockquote> <p>For the usage and code, please refer to https://github.com/ronggong/EUSIPCO2017&nbsp;</p>

opencc-by-nc-4.0Feb 2017View details →
zenodo36/100

Phoneme Based Embedded Segmental K-Means for ZeroSpeech2017 Track 2

<p>Our submission applying the phoneme based embedded segmental k-means model to the&nbsp;ZeroSpeech2017 challenge track 2.&nbsp;</p> <p>This is a preliminary version. More details can be found here:&nbsp;<a href="https://www.kamperh.com/papers/bhati+kamper+murty_icassp2018.pdf">https://www.kamperh.com/papers/bhati+kamper+murty_icassp2018.pdf</a>&nbsp;and&nbsp; https://github.com/Saurabhbhati/recipe_zs2017_track2</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2018View details →
zenodo36/100

Dataset for Interspeech 2018 submission: Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions

<p>This dataset contains the materials for training, testing the joint and HSMM models mentioned in the paper &quot;<em>Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions&quot;</em>.</p> <p>The filename list of this dataset can be found in the function <em>get_train_test_recordings_joint()</em> of <em>./general/trainTestSeparation.py</em> file. The dataset contains the Praat TextGrids and .wavs of the variables: <em>train_primary_school, val_primary_school</em> and <em>test_primary_school</em>. For accessing other datasets such as <em>train_nacta_2017, train_nacta</em> and <em>train_sepa</em>, please download them from the links:</p> <p>jingju dataset part1:&nbsp;<a href="https://zenodo.org/record/1185154">https://zenodo.org/record/1185154</a></p> <p>jingju dataset part2:&nbsp;<a href="https://doi.org/10.5281/zenodo.842229">https://doi.org/10.5281/zenodo.842229</a></p> <p>Once you have downloaded these three datasets, you need to set the paths in <em>./general/filePathShared.py</em>.</p> <p>Set <em>path_jingju_dataset</em> to the parent path of these three datasets.</p> <p>Set <em>primarySchool_dataset_root_path</em> to the path of the interspeech2018 dataset (the current dataset).</p> <p>Set <em>nacta_dataset_root_path</em> to the path of the jingju&nbsp;dataset part1.</p> <p>Set <em>nacta2017_dataset_root_path</em> to the path the jingju&nbsp;dataset part2.</p> <p>For more information on this paper, please refer to the Github page:&nbsp;<a href="https://github.com/ronggong/interspeech2018_submission01">https://github.com/ronggong/interspeech2018_submission01</a></p> <p>&nbsp;</p>

opencc-by-nc-4.0Feb 2018View details →
zenodo36/100

Repository of speech features from speakers with and without Parkinson's Disease. Neurovoz - Rasta PLP - V2 - Scientific Reports Publication: Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease

<p>This repository contains the Rasta-PLP features of six different speech recordings (sentences) from Neurovoz corpus (47 parkinsonian and 32 control speakers whose mother tongue is Spanish Castillian.)<br> Number of PLP coefficients: [6, 8, 10, 12, 14, 16, 18, 20].<br> Delta coefficients: Yes<br> Delta Delta coefficients: Yes<br> Sampling rate: 16 kHz<br> Frame size: 15 ms<br> Frame overlapping: 50%</p> <p>This subset of the Neurovoz corpus was recorded between 2015 and 2017 by Universidad Polit&eacute;cncia de Madrid and Hospital General Universitario Gregorio Mara&ntilde;&oacute;n.</p> <p>This version includes the same files as the previous version and information about UPDRS, H&amp;Y, years since diagnosis and age of each participant.</p> <p>The sentences were:</p> <p>BARBAS: &quot;Cuando las barbas de tu vecino veas pelar, pon las tuyas a remojar&quot;</p> <p>CALLE: &quot;De la calle vendr&aacute; quien de tu casa te echar&aacute;&quot;</p> <p>DIABLO: &quot; Cuando el diablo no sabe qu&eacute; hacer, con el rabo mata moscas &quot;</p> <p>PETACA BLANCA: &quot; La petaca blanca es m&iacute;a&quot;</p> <p>PIDIO: &quot;No pidas a quien pidi&oacute; ni sirvas a quien sirvi&oacute;&quot;</p> <p>SOMBRA: &quot; El que a buen &aacute;rbol se arrima, buena sombra le cobija &quot;</p> <p>&nbsp;</p> <p>How to cite:<br> [1] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Grandas-Perez, F., Shattuck-Hufnagel, S. Yag&uuml;e-Jimenez, V., and Dehak, N. (2019).&nbsp;Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson&rsquo;s disease.Scientific reports&nbsp;9,&nbsp;19066.</p> <p><br> [2] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Villalba, J., Rusz,&nbsp;J.,&nbsp;Shattuck-Hufnagel, S. and Dehak, N. (2019).&nbsp;A forced Gaussians based methodology for the differential evaluation of Parkinson&#39;s Disease by means of speech processing. Biomedical Signal Processing and Control, 48, 205-220.</p> <p>BibTeX:</p> <pre><code>@article{moro2019phonetic, title={Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge A. and Godino-Llorente, Juan I. and Grandas-Perez, Francisco and Shattuck-Hufnagel, Stefanie and Yague-Jimenez, Virginia and Dehak, Najim}, journal={Scientific Reports}, volume={9}, pages={19066}, year={2019}, publisher={Nature Research Publishing} } @article{moro2019forced, title={A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge Andres and Godino-Llorente, Juan Ignacio and Dehak, Najim}, journal={Biomedical Signal Processing and Control}, pages={205--220}, volume={48}, year={2019}, publisher={Elsevier} } </code></pre> <p>&nbsp;</p>

opencc-by-4.0Sep 2019View details →
zenodo36/100

VeLePor: European Portuguese Verbal Paradigms in Phonemic Notation

<p>This is a collection of European Portuguese verbal paradigms, in phonemic notation. They are suited for both computational and manual analysis.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Supplementary material for "Investigating phoneme-dependencies of spherical voice directivity patterns II: Various Groups of Phonemes"

<p>The .pdf&nbsp;file contains</p> <ul> <li>general information on the voice directivity files in the SOFA format</li> <li>information on the indices and names of the SOFA-files</li> <li>additional plots</li> </ul> <p>&nbsp;</p> <p>The .zip&nbsp;files contain</p> <ul> <li>voice directivities in the SOFA format sampled on the sparse measuring grid</li> <li>voice directivities in the SOFA format upsampled to a dense grid</li> </ul> <p>&nbsp;</p> <p>The Matlab script provides</p> <ul> <li>an example reading a dataset, performing spatial upsampling if required, and creating some basic plots.&nbsp;</li> </ul>

opencc-by-4.0Jan 2023View details →
zenodo32/100

Data for HPN-DREAM analysis using OBRWR and PHONEMeS

<p>Data for the HPN-DREAM analysis.</p>

opencc-by-4.0Nov 2024View details →
dryad32/100

Data from: Experimental evidence for phonemic contrasts in a nonhuman vocal system

Open the record for dataset details and reuse information.

publicJun 2016View details →
ClinicalTrials.gov24/100

Correlation Between Respiratory Impairment and Phonemes Alteration in Dystrophinopathy Patients With Respiratory Failure

ClinicalTrials.gov study NCT02411370. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record