Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
23
datasets available to search
ShareScore release 0.9.0
Dataset results
23 results for “Phonetics”
Supplementary material accompanying "Factoring lexical and phonetic phylogenetic characters from word lists"
<p>This repository contains the scripts and the data that were used to run the analyses for the paper "Factoring lexical and phonetic phylogenetic characters from word lists". For details, please refer to the README.md file provided along with the dataset. If you run into problems replicating the analysis, please do not hesitate to contact the authors.</p>
CLDF dataset accompanying Yang's "Phonetic Tone Change" from 2022
<p>Cite the source of the dataset as:</p> <blockquote> <p>Yang, Cathryn (2022): The phonetic tone change *high > rising: Evidence from the Ngwi dialect laboratory. Diachronica. DOI: https://doi.org/10.1075/dia.19062.yan</p> </blockquote>
CLDF dataset derived from List and Prokić's "Benchmark Database of Phonetic Alignments" from 2014
<p>Cite the source of the dataset as:</p> <blockquote> <p>List, Johann-Mattis and Jelena Prokić. (2014). A benchmark database of phonetic alignments in historical linguistics and dialectology. In: Proceedings of the International Conference on Language Resources and Evaluation (LREC), 26 — 31 May 2014, Reykjavik. 288-294.</p> </blockquote>
Data and Analyses for Defining Filler Particles: A Phonetic Account of the Terminology, Form, and Grammatical Classification of Filled Pauses.
<p>Aggregated data and analyses for the article Belz, Malte (2023): Defining Filler Particles: A Phonetic Account of the Terminology,<br> Form, and Grammatical Classification of "Filled Pauses". Languages. <a href="https://www.mdpi.com/journal/languages/special_issues/Pauses_in_Speech">https://www.mdpi.com/journal/languages/special_issues/Pauses_in_Speech</a></p>
Benchmark Database for Phonetic Alignments
<p>In the last two decades, alignment analyses have become an important technique in quantitative historical linguistics and dialectology. Phonetic alignment plays a crucial role in the identification of regular sound correspondences and deeper genealogical relations between and within languages and language families. Surprisingly, up to today, there are no easily accessible benchmark data sets for phonetic alignment analyses. Here we present a publicly available database of manually edited phonetic alignments which can serve as a platform for testing and improving the performance of automatic alignment algorithms. The database consists of a great variety of alignments drawn from a large number of different sources. The data is arranged in a such way that typical problems encountered in phonetic alignment analyses (metathesis, diversity of phonetic sequences) are represented and can be directly tested.</p>
Dargwa: basic information on phonetics
<p>This lecture is part of the lecture series: Glottothèque: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online (electronic resource). Bamberg, Cambridge, Göttingen, Moscow, Nicosia, Paris: LACIM network, at https://spw.unigoettingen.de/projects/lacim/, edited by Christiane Bulut, Anaïd Donabédian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>
Western Thrace Turkish: Phonology - Phonetic Features, Morphology and Syntax
<p>These snippets present linguistic features of Western Thrace Turkish, a Balkan Turkish dialect spoken in Northeastern Greece. This lecture is part of the lecture series: <em>Glottothèque: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online </em>(electronic resource). Bamberg, Cambridge, Göttingen, Moskow, Nicosia, Paris: LACIM network, at https://spw.uni-goettingen.de/projects/lacim/, edited by Christiane Bulut, Anaïd Donabédian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>
Feature Extraction Using Hidden Markov Model for a Phonetic Process
<p>Speech is one of the primary forms of communication among humans. In real life, a dictionary is used to seek the pronunciation of a complex word; but, for computers, this look-up table is called a phonetic dictionary. A speech recognition process tags a word-utterance to its phoneme structure, thereby returning the grapheme representation. However, the speech recognition process is challenging because of the contextual relationship between words and sentences, dependent on speakers’ intentions. Further, factors influencing time, accents, noisy environment, and data security impose accuracy threats. The present research study proposes a new hybrid speech recognition model by considering three significant aspects: sound generation through phonetic representation, sound acoustics for transmission, and sound reception on how the sound is received. These steps are achieved through a speech-to-text model divided into various stages such as noise removal, speech-pause detection, feature extraction through framing, and windowing by adopting Hidden Markov Model (HMM). The implementation is performed on a phonetic tool, Praat. The robustness of the model is estimated using evaluation metrics such as f-measure and accuracy, resulting in 98% and 99% scores, respectively. Thus, the proposed approach efficiently transforms the spoken words into their corresponding text.</p>
Summary data to support Hall (2019), '(e) in Normandy: The sociolinguistics, phonology and phonetics of the "Loi de Position"'
<p>Article abstract:</p> <p>This article uses the pronunciation of stressed Intonational Phrase-final /ε/ and /e/ in two communities in Normandy, France, to illustrate the convergence of two sociolinguistic processes on the same phonological result: increasing application of the <em>Loi de Position</em>. In both communities (one rural and further from Paris, one urban and closer to Paris), there is now no consistent community-wide phonetic distinction between the two phonemes in that environment. It is suggested that the <em>Loi de Position</em> is already widely applied in the rural site, but speakers are still conscious of the formal norm whereby it is not applied; for the urban site, apparent-time changes for this variable reflect changes in Parisian speech. The theoretical implications of the study concerning speakers’ organisation of their vowel-space, and concerning the increasing application of the <em>Loi de Position</em> in the French of France, are examined. These conclusions are reached by per-speaker analysis of F1 and F2 separately from each other (rare in French linguistics). As a measure of community cohesion, the article introduces to linguistics the coefficient of variation (more common in biology and medicine).</p>
Improving phonetic alignment by handling secondary sequence structures
<p>Supplementary material accompanying the paper "Improving phonetic alignment by handling secondary sequence structures".</p> <p>The data consists of 5 files:</p> <ul> <li>gold_standard.psa : the gold standard used in the analysis in PSA format</li> <li> sca-secondary.psa : the output of the algorithm with the secondary extension</li> <li>sca-traditional.psa : the output of the traditional algorithm</li> <li>sca-secondary-diff.psa : the differences of the secondary extension compared to the GS</li> <li>sca-traditional-diff.psa : the differences of the traditional algorithm compared to the GS</li> </ul> <p>For a description of the file-format used in this dataset, please refer to the LingPy tutorial under http://lingpy.org.</p>
Supplementary materials for "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study"
<div> <pre>This is the supplementary data for the article "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study" by Anastasia Bauer, Anna Kuder, Marc Schulder and Job Schepens.<br><br>The supplementary data consists of three components, stored in separate directories:<br>- <code>annotations/</code>: The manual annotations of head nod categories, produced by Anna Kuder and Anastasia Bauer.<br>- <code>pose_analysis/</code>: Code and input/output files for the pose-based automatic analysis of phonetic attributes head nods, produced by Marc Schulder.<br>- <code>statistical_analysis/</code>: Code for the statistical analysis of the other two components and for the creation of related figures, produced by Job Schepens.<br><br>For further details, see the README files of the respective directories.</pre> </div>
Repository of speech features from speakers with and without Parkinson's Disease. Neurovoz - Rasta PLP - V2 - Scientific Reports Publication: Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease
<p>This repository contains the Rasta-PLP features of six different speech recordings (sentences) from Neurovoz corpus (47 parkinsonian and 32 control speakers whose mother tongue is Spanish Castillian.)<br> Number of PLP coefficients: [6, 8, 10, 12, 14, 16, 18, 20].<br> Delta coefficients: Yes<br> Delta Delta coefficients: Yes<br> Sampling rate: 16 kHz<br> Frame size: 15 ms<br> Frame overlapping: 50%</p> <p>This subset of the Neurovoz corpus was recorded between 2015 and 2017 by Universidad Politécncia de Madrid and Hospital General Universitario Gregorio Marañón.</p> <p>This version includes the same files as the previous version and information about UPDRS, H&Y, years since diagnosis and age of each participant.</p> <p>The sentences were:</p> <p>BARBAS: "Cuando las barbas de tu vecino veas pelar, pon las tuyas a remojar"</p> <p>CALLE: "De la calle vendrá quien de tu casa te echará"</p> <p>DIABLO: " Cuando el diablo no sabe qué hacer, con el rabo mata moscas "</p> <p>PETACA BLANCA: " La petaca blanca es mía"</p> <p>PIDIO: "No pidas a quien pidió ni sirvas a quien sirvió"</p> <p>SOMBRA: " El que a buen árbol se arrima, buena sombra le cobija "</p> <p> </p> <p>How to cite:<br> [1] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Grandas-Perez, F., Shattuck-Hufnagel, S. Yagüe-Jimenez, V., and Dehak, N. (2019). Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson’s disease.Scientific reports 9, 19066.</p> <p><br> [2] Moro-Velazquez, L., Gomez-Garcia, J. A., Godino-Llorente, J. I., Villalba, J., Rusz, J., Shattuck-Hufnagel, S. and Dehak, N. (2019). A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing. Biomedical Signal Processing and Control, 48, 205-220.</p> <p>BibTeX:</p> <pre><code>@article{moro2019phonetic, title={Phonetic relevance and phonemic grouping of speech in the automatic detection of Parkinson's Disease}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge A. and Godino-Llorente, Juan I. and Grandas-Perez, Francisco and Shattuck-Hufnagel, Stefanie and Yague-Jimenez, Virginia and Dehak, Najim}, journal={Scientific Reports}, volume={9}, pages={19066}, year={2019}, publisher={Nature Research Publishing} } @article{moro2019forced, title={A forced Gaussians based methodology for the differential evaluation of Parkinson's Disease by means of speech processing}, author={Moro-Velazquez, Laureano and Gomez-Garcia, Jorge Andres and Godino-Llorente, Juan Ignacio and Dehak, Najim}, journal={Biomedical Signal Processing and Control}, pages={205--220}, volume={48}, year={2019}, publisher={Elsevier} } </code></pre> <p> </p>
This Study Will Assess Whether the Treatment Provided by Dentist is Successful in Meeting the Expectations of Patient by Asking Questions Related to Aesthetics, Chewing Ability, Comfort and Phonetics
ClinicalTrials.gov study NCT07100678. IPD Sharing: NO. Countries: 1. Publications: 0.
Task of Acoustic-phonetic Decoding on Anatomic Deficits in Paramedical Assessment of Speech Disorders for Patients Treated for Oral or Oropharyngeal Cancer
ClinicalTrials.gov study NCT04742998. IPD Sharing: NO. Countries: 1. Publications: 3.
IMPORTANCE OF MODERN PEDAGOGICAL TECHNOLOGIES IN TEACHING PHONETICS IN HIGHER EDUCATIONAL INSTITUTIONS
Open the record for dataset details and reuse information.
Supplementary Materials for the article entitled 'Anti-hiatus tendencies in Spanish: Rate of occurrence and phonetic identification', published in Linguistics
<p>see the publication for details</p>
THE PHONETIC SYSTEM OF UZBEK, FINNISH, ENGLISH AND DIFFERENCES BETWEEN THE VOWEL SOUNDS OF THIS THREE LANGUAGES.
Open the record for dataset details and reuse information.
Phonetic Richness for Improved Automatic Speaker Verification: Aplawd-Based Speaker Verification Protocols
<p>This data was created as part of the work entitled "Phonetic Richness for Improved Automatic Speaker Verification", published in EUSIPCO 2024.</p> <p>In doing our experimental evaluation, Pindrop used FLAC files from the APLAWD Markings Dataset. You can obtain copies of these same FLAC files directly from the developers here: <a href="https://github.com/serwy/aplawdw" target="_blank" rel="noopener">https://github.com/serwy/aplawdw</a></p> <p>This data contains two automatic speaker verification protocols: "Aplawd" and "Aplawd-Repetitive". See below for details regarding the structure of the data for each of these protocols.</p> <p><strong>Files defining the Aplawd protocol:</strong><br>> <code>Pindrop_aplawd_protocol/models.csv</code></p> <ul> <li>This file defines what files should be used to create each enrollment model used in this protocol.</li> <li>Columns: <ul> <li><code>enroll_id</code>: unique string identifying the enrollment model that this row corresponds to</li> <li><code>filename</code>: filename specifying the file from the APLAWD Markings Dataset to be used in creating the enrollment model specified for this row by enroll_id.</li> </ul> </li> <li>Note: a single model is created for each enroll_id, using all of the files specified for the given enroll_id</li> </ul> <p>> <code>Pindrop_aplawd_protocol/trials_same_gender.csv</code></p> <ul> <li>This file defines the trials that make up this protocol, specifying pairs of enrollment model and probe audios to test against eachother.</li> <li>Columns: <ul> <li><code>enroll_id</code>: unique string identifying the enrollment model whose identity will be compared against in the trial defined by this row.</li> <li><code>probe_filename</code>: filename specifying the file from the APLAWD Markings Dataset for which the identity of the speaker should be compared against the identity of the corresponding enrollment model for this row.</li> <li><code>probe_identity</code>: unique string identifying the true speaker identity for this row's probe.</li> </ul> </li> </ul> <p> </p> <p><strong>Files defining the Aplawd-Repetitive protocol:</strong></p> <p>(<strong>NOTE</strong>: the key difference from the above Aplawd protocol is in the probe_filename column of the trials_same_gender.csv file)</p> <p>> <code>Pindrop_aplawd_repetitive_protocol/models.csv</code></p> <ul> <li>This file defines what files should be used to create each enrollment model used in this protocol.</li> <li>Columns: <ul> <li><code>enroll_id</code>: unique string identifying the enrollment model that this row corresponds to</li> <li><code>filename</code>: filename specifying the file from the APLAWD Markings Dataset to be used in creating the enrollment model specified for this row by enroll_id.</li> </ul> </li> <li> Note: a single model is created for each enroll_id, using all of the files specified for the given enroll_id</li> </ul> <p>> <code>Pindrop_aplawd_repetitive_protocol/trials_same_gender.csv</code></p> <ul> <li>This file defines the trials that make up this protocol, specifying pairs of enrollment model and probe audios to test against eachother.</li> <li>Columns: <ul> <li><code>enroll_id</code>: unique string identifying the enrollment model whose identity will be compared against in the trial defined by this row.</li> <li><code>probe_filename</code>: an underscore ("_") delimited string of filename(s) specifying the file(s) from the APLAWD Markings Dataset that should be concatenated to form the probe audio for this row's trial. The resulting probe's speaker identity should be compared against the identity of the corresponding enrollment model for this row.</li> <li><code>probe_identity</code>: unique string identifying the true speaker identity for this row's probe.</li> </ul> </li> </ul>
Data for article entitled 'Anti-hiatus tendencies in Spanish: Rate of occurrence and phonetic identification'', published in Linguistics
<p>see the article</p>
A Systematic Investigation of Phonetic Complexity Effects on Articulatory Motor Performance in Progressive Dysarthria
ClinicalTrials.gov study NCT03613038. IPD Sharing: YES. Countries: 1. Publications: 0.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.