Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
37
datasets available to search
ShareScore release 0.7.1
Dataset results
37 results for “IPA”
THCHS-30 - Aligned IPA transcriptions
<p>This upload contains aligned IPA transcriptions for the <a href="https://www.openslr.org/18/">THCHS-30 dataset from OpenSLR</a>. Thereby, punctuation is added, silence marked and duration markers for each phoneme are assigned. Furthermore, the silence on the beginning and ending of each file is marked.</p> <p>The words were transcribed using <a href="https://pypi.org/project/pypinyin/">pypinyin</a> (v0.47.1) via <a href="https://pypi.org/project/dict-from-pypinyin/">dict-from-pypinyin</a> (v0.0.1) and mapped to IPA using the <code>pinyin-ipa-map-TONE3-all.json</code> mapping <a href="https://zenodo.org/record/7525638">from here</a>. The alignment was done using <a href="https://zenodo.org/record/6796264">Montreal Forced Aligner</a> (v2.0.5) and the acoustic model <a href="https://mfa-models.readthedocs.io/en/latest/acoustic/Mandarin/Mandarin%20MFA%20acoustic%20model%20v2_0_0a.html">Mandarin MFA</a> (v2.0.0a).</p> <p>Phoneme duration markers:</p> <ul> <li><code>˘</code> -> [0, 20) percentile (speaker-wise), e.g., <code>a˥˩˘</code></li> <li>(none) -> [20, 80) percentile (speaker-wise), e.g., <code>a˥˩</code></li> <li><code>ˑ</code> -> [80, 90) percentile (speaker-wise), e.g., <code>a˥˩ˑ</code></li> <li><code>ː</code> -> [90, inf) percentile (speaker-wise), e.g., <code>a˥˩ː</code></li> </ul> <p>Thereby each phoneme (including tones) was considered on its own, i.e., phonemes with different tones were not considered together for the percentile calculation.</p> <p>Silence markers:</p> <ul> <li><code>SILX</code> -> silence at start/end of a recording (aligned on all tiers)</li> <li><code>SIL0</code> -> no silence</li> <li><code>SIL1</code> -> [0, 33.33333333) percentile of all silences (speaker-wise)</li> <li><code>SIL2</code> -> [33.33333333, 66.66666666) percentile of all silences (speaker-wise)</li> <li><code>SIL3</code> -> [66.66666666, inf) percentile of all silences (speaker-wise)</li> </ul> <p>Files:</p> <ul> <li><code>grids.zip</code> <ul> <li>contains TextGrids for all audio files containing three tiers <code>words</code>, <code>phonemes</code> and <code>transcription</code> <ul> <li><code>words</code> contains the aligned Chinese words</li> <li><code>phonemes</code> contains the IPA pronunciations including silence markers at start and end (<code>SILX</code>)</li> <li><code>transcription</code> contains unaligned phonemes including punctuation and word boundary labels (<code>SIL0</code>)</li> </ul> </li> <li>the folder structure is equal to one from the dataset</li> </ul> </li> <li><code>grids-sdp.zip</code> <ul> <li>same as <code>grids.zip</code> except the folder structure is the one from <a href="https://pypi.org/project/speech-dataset-parser/">speech-dataset-parser</a> (v0.0.4)</li> </ul> </li> <li><code>preview-start/middle/end.png</code> <ul> <li>preview of the first TextGrid from speaker <code>A2</code> opened in Praat in different positions</li> </ul> </li> <li><code>words-vocabulary.txt</code> <ul> <li>contains all Chinese words from tier <code>words</code></li> </ul> </li> <li><code>phonemes-vocabulary.txt</code> <ul> <li>contains all phonemes from tier <code>phonemes</code></li> </ul> </li> <li><code>transcription-vocabulary.txt</code> <ul> <li>contains all phonemes/punctuation from tier <code>transcription</code></li> </ul> </li> <li><code>phonemes-durations.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code></li> </ul> </li> <li><code>phonemes-durations-simple.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers are ignored</li> </ul> </li> <li><code>phonemes-durations-simple-toneless.pdf</code> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers and tones are ignored</li> </ul> </li> <li><code>pronunciations-broad.dict</code> <ul> <li>contains the broad pronunciations for each word including punctuation but not duration markers</li> <li>e.g., <code>一下。 i˥ ɕ j a˥˩ 。</code></li> </ul> </li> <li><code>pronunciations-narrow.dict</code> <ul> <li>contains the narrow pronunciations for each word including punctuation, duration markers and weights (= occurrence) over all speakers</li> <li>e.g., <code>一下。 3 i˥ ɕ j a˥˩ː 。</code></li> </ul> </li> <li><code>pronunciations-narrow-speakers.zip</code> <ul> <li>contains the narrow pronunciations separated for each speaker</li> </ul> </li> <li><code>script.sh</code> <ul> <li>contains the script to reproduce all results</li> <li>error in line 32: replace with <code>speech-dataset-parser==0.0.4</code></li> </ul> </li> </ul>
CLDF dataset with phoneme inventories from the "Journal of the IPA", aggregated by Baird et al. (2021)
<p>Cite the source of the dataset as:</p> <blockquote> <p>Baird, L., Evans, N., & Greenhill, S. J. (2021). Blowing in the wind: Using 'North Wind and the Sun' texts to sample phoneme inventories. Journal of the International Phonetic Association, 1–42. doi:10.1017/s002510032000033x</p> </blockquote>
Templates for BWS IPA and NPS project evaluations
<p>Templates with example data to facilitate the completion of project evaluations using Excel to undertake three analyses:</p> <p>1. Importance-Performance Analysis (IPA) with Gap Analysis</p> <p>2. Net Promoter Score</p> <p>3. Best-Worst-Scaling template in seperate file in version 1</p>
Proteomic data sets after selecting mitochondrial proteins from the scaffold software for Ingenuity Pathway analysis (IPA Qiagen)
<p>List of fold change proteomic data sets of dFCM- 39 vs. 12Day and105 vs. 12Day, cFCM- 40 vs. 12Day and115 vs. 12Day , mouse heart 90 vs. 1 day after selecting mitochondrial proteins from the scaffold software for Ingenuity Pathway Analysis (IPA Qiagen)</p>
Dataset: ImmunoPrecise Antibodies Ltd. (IPA) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Linguistic Survey of India (in IPA)
<p>This dataset is lexical lists from the Linguistic Survey of India, normalized into IPA.</p> <p> </p> <p><em>Linguistic survey of India</em> / [compiled and edited] by George Abraham Grierson. Calcutta : Office of the Superintendent of Government Printing, India, 1903-1928.</p> <p> </p> <p> </p>
Inflected lexicon of Russian Nouns in IPA notation
<p>This inflected lexicon of Russian Nouns is based on data generated by a DATR fragment for the nominal system of Russian (Dunstan Brown et al, 2011), which was used in Brown and Hippisley (2012). The files were automatically transcribed to an IPA-based notation by Sacha Beniamine (2018), following rules provided by Dunstan Brown. We also corrected a few errors in the original data found manually. We thank Sasha krasovitsky for providing comments on the phonological feature definitions.</p>
IPA Transcription and Recordings
<p>This file provides transcription of dialects data using IPA found in ten survey sites in West Manggarai, East Nusa Tenggara, Indonesia.</p>
IPA_dataset
<p>The Intellectual Property Abstracts dataset (IPA dataset) consists of English relevant and non-relevant pairs of patent abstracts. This dataset is a processed extract from the MAREC dataset (http://www.ifs.tuwien.ac.at/imp/marec.shtml).</p>
Founders All Day IPA - Session Ale
One of my favorite beers from one of my favorite Michigan breweries. - [Founders Brewing Co.](https://foundersbrewing.com/our-beer/all-day-ipa/) - [BeerAdvocate](https://www.beeradvocate.com/beer/profile/1199/58914/) ----- It's not an amazing scan, but I thought pretty good for a shiny / transparent object. - 49 Photos from an iPhone 6 - Reconstructed in Agisoft PhotoScan - Some cleanup in MeshLab ----- Download includes: - OBJ + MTL + 4k texture - Original photos - PhotoScan package Source: Objaverse 1.0 / Sketchfab
Integrated Probabilistic Annotation (IPA): A Bayesian-based annotation method for metabolomic profiles integrating biochemical connections, isotope patterns and adduct relationships - Supplementary Data
<ol> <li>Supplementary_data_1: data and code for standards analysis and database update</li> <li>Supplementary_data_2.zip: data and code used for the generation of the synthetic experiment</li> <li>Supplementary_data_3.zip: data, code, and results of the <em>E. coli</em> dataset analysis</li> <li>Supplementary_data_4.zip: data, code, and results of the beer dataset analysis</li> <li>Supplementary_data_5.zip: data, code, and results of the comparison with xMSannotator</li> </ol>
LJ Speech - Aligned IPA transcriptions
<p>Files:</p> <ul> <li> <p><code>grids.zip</code></p> <ul> <li>contains TextGrids for all audio files containing three tiers <code>words</code>, <code>phonemes</code> and <code>transcription</code> <ul> <li><code>words</code> contains the aligned normalized English words</li> <li><code>phonemes</code> contains IPA pronunciations transcribed using CMU dictionary which then were aligned with Montreal Forced Aligner. The pronunciations were then mapped from ARPAbet to IPA and duration marks were applied (without punctuation)</li> <li><code>transcription</code> contains unaligned phonemes including punctuation and word boundary labels (SIL0)</li> </ul> </li> </ul> </li> <li> <p><code>preview.png</code></p> <ul> <li>preview of the first TextGrid opened in Praat</li> </ul> </li> <li> <p><code>words-vocabulary.txt</code></p> <ul> <li>contains all words from tier <code>words</code></li> </ul> </li> <li> <p><code>phonemes-vocabulary.txt</code></p> <ul> <li>contains all phonemes from tier <code>phonemes</code></li> </ul> </li> <li> <p><code>transcription-vocabulary.txt</code></p> <ul> <li>contains all phonemes/punctuation from tier <code>transcription</code></li> </ul> </li> <li> <p><code>phonemes-durations.pdf</code></p> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code></li> </ul> </li> <li> <p><code>phonemes-durations-simple.pdf</code></p> <ul> <li>contains the plotted phoneme duration distribution of tier <code>phonemes</code> if all duration markers are ignored</li> </ul> </li> <li> <p><code>pronunciations.dict</code></p> <ul> <li>contains the pronunciations for each word including punctuation and weights (occurrence)</li> </ul> </li> <li> <p><code>script.sh</code></p> <ul> <li>contains the script to reproduce all results</li> </ul> </li> </ul> <p>Phoneme duration marker:</p> <ul> <li><code>˘</code> -> [0, 20) percentile</li> <li><code>ˑ</code> -> [80, 90) percentile</li> <li><code>ː</code> -> [90, inf) percentile</li> </ul> <p>Silence marker:</p> <ul> <li><code>SIL0</code> -> no silence</li> <li><code>SIL1</code> -> [0, 33.33) percentile</li> <li><code>SIL2</code> -> [33.33, 66.66) percentile</li> <li><code>SIL3</code> -> [66.66, inf) percentile</li> </ul> <p> </p>
source IPA files of IOS app with on-device models
<p>To meet the requirements of the double-blind review, we choose to anoymously share here our collected dataset of ios apps with on-device models. For some reason unknown to us, we are also unable to upload individual files that are over 5 g. After many attempts, we can only share some of the files here, hope it is understood. We will release all ios app files and models after the anonymity period is over.</p>
IPA Targeted Adoptive Immunotherapy vs Adult Haplo-identical Cell Infusion During Induction of High Risk Leukemia
ClinicalTrials.gov study NCT02508324. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Targeting the IPA and Matching for the Non-Inherited Maternal Antigen for Haplo-Cord Transplantation
ClinicalTrials.gov study NCT01810588. IPD Sharing: Not stated. Countries: 1. Publications: 2.
OCR model for lexical lists in Chinese-IPA Glossing, Ground Truth
<p>The ground truth dataset for the OCR model consisted of jpg, pdf, and xml files. The training process was conducted using Transkribus, employing a PyLaia model constructed using ground truth data derived from lexical lists encompassing ten literary works that document Burmish languages. These languages include Achang, Bola, Chashan, Langsu, Leqi, and Zaiwa. The lexical lists utilized for the training phase were predominantly composed in both Chinese characters and International Phonetic Alphabet (IPA) symbols.</p> <p>The training was done on 311 pages and validation on 34 pages of ten lexical lists of Burmish languages on Transkribus with the default PyLaia model:</p> <p>(1) Achang<br> • adapted by Hill & Cooper (2020) from Dai & Cui (1985)<br> (2) Bola<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (3) Bola<br> • adapted by Hill & Cooper (2020) from Dai et al. (2007)<br> (4) Chashan<br> • adapted by Hill & Cooper (2020) from Dai et al. (2010)<br> (5) Langsu<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (6) Langsu<br> • adapted by Hill & Cooper (2020) from Dai (2005)<br> (7) Leqi<br> • adapted by Hill & Cooper (2020) from He & Chen (2004)<br> (8) Leqi<br> • adapted by Hill & Cooper (2020) from Dai & Li (2006)<br> (9) Leqi<br> • adapted by Hill & Cooper (2020) from Dai & Jie (2007)<br> (10) Zaiwa<br> • adapted by Hill & Cooper (2020) from Xu & Xu (1984)</p> <p>The material utilized for assessing the performance of the trained models is enclosed within the dataset, comprising a lexical list of the Tujia language as documented by Tian in 1986. The primary objective of the trained model was the recognition of printed lexical lists employing Chinese-IPA glossing.</p>
Impact of the Introduction of a Performance Improvement Program on the Initial Management of Sepsis and Septic Shock in Adults in the Emergency Department: a Before-and-after Study (IPA-SOS)
ClinicalTrials.gov study NCT06657625. IPD Sharing: YES. Countries: 1. Publications: 5.
Diagnosis of Invasive Pulmonary Aspergillosis (IPA) in Critically Ill Patients
ClinicalTrials.gov study NCT01866020. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Voriconazole for IPA in Chinese Patients With COPD
ClinicalTrials.gov study NCT02234739. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Supplementary Spreadsheet S1. Canonical pathway using IPA analysis
<p>Supplementary Spreadsheet S1. Canonical pathway using IPA analysis</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.