Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
17
datasets available to search
ShareScore release 0.9.0
Dataset results
17 results for “accentism”
The Accent Audio Files of Koshikijima Accent Database
<p>This is the accent audio files of Koshikijima Accent Database, an accent database of the endangered dialect of Koshikijima Japanese spoken on the Koshikijima Islands, Kagoshima Prefecture, Japan. It has been released for academic purposes,for both research and education. It contains the accent data of two speakers each from eight villages on the Islands, with the consent of the individual speakers to make their speech publicly available. The original data of the database comes from the ten-year-long collaborative fieldwork done by six editors. The database file can be downloaded from <a href="https://doi.org/10.15084/0002000053">https://doi.org/10.15084/0002000053</a>.</p> <p>Those who wish to use the data of this database to write an academic article or an essay are required to cite the following as a reference: Kubozono, Haruo, Zendo Uwano, Nobuko Kibe, Akiko Matsumori & Tetsuo Nitta (eds.) (2016) Koshikijima Accent Database.<a href="https://www2.ninjal.ac.jp/koshikijima/"><https://www2.ninjal.ac.jp/koshikijima/</a>></p> <p> </p>
Accent Classification Dataset
<p>This dataset contains audio recordings of 12 different accents across the UK: Northern Ireland (NI), Scotland, Wales (SW), North East England (NE), North West England (NW), Yorkshire and Humber (YAH), East Midlands (EM), West Midlands (WM), East of England (EE), Greater London (GL), South East England (SE), South West England (SW). We split the data into a Male: Female ratio of 1:1, this is labelled with either '_M' for male or '_F' for female within the dataset. The audio dataset was compiled using opensource YouTube videos and it a collation of different accents, the audio files were trimmed for uniformity. The Audio files are of length 30 seconds, with the first 5 seconds and last 5 seconds of the signal being blank. We also resample the audio signals at 8 kHz, again for uniformity and to remove any noise present in the audio signals whilst retaining the underlying characteristics. The intended application of this dataset was to be used in conjunction with a deep neural network for accent and gender classification tasks.</p> <p>The dataset also contains an unseen dataset of the Google opensource digit dataset, which contains audio files of the digits 1-9. This is included to test any models developed using the original dataset to confirm model performance to data variations. </p>
The SIAEW Corpus of Spanish Iso-Accented English Words
<p>The SIAEW Corpus (Spanish Iso-Accented English Words) is a collection of monosyllabic English words in which one segment (the 'target') is replaced with its Spanish-accented counterpart, at one of 5 gradations of accentedness. Steps are equally-spaced in accentedness as judged by native listeners. The procedure used to generate the Corpus is described in detail in Pérez Ramón, García Lecumberri and Cooke (submitted to Interspeech 2022); for a preprint contact rperez.ram@gmail.com. </p> <p>Please refer to the document SIAEW.pdf for a longer description of the SIAEW Corpus contents.</p>
Effects of rhythm and accent patterns on tempo-keeping property of finger tapping
<p><span>Tempo of music performance is often accelerated irrespective of players' intention. Though the characteristics of the tempo deviation phenomenon have been investigated using a finger-tapping task, most such studies dealt with tapping with a fixed interval; few studies considered the effects of rhythm and accent, both important factors of music performance. Here, we asked how different rhythm and accent patterns affected the tempo-keeping property using a synchronization-continuation task paradigm: Participants were asked to keep tapping while reproducing the given rhythm/accent patterns designated by the target tones. Tapping tempo was significantly deviated depending on rhythm/accent patterns, but their magnitudes were only several percent in 150 seconds, much smaller than those in real music performance. We also ran experiments under the conditions that participants need to reproduce the accent patterns, but only feedback tones were modulated. These auditory modulations affected the tempo deviation, implying that not only motor process for producing the accents but also perceptual process induced by the accented auditory feedback influence the tempo maintenance of finger tapping. In sum, the present finding shows that sensorimotor processing for un-uniform finger tapping can disturb the long-term tempo-keeping process. We also discussed related topics based on the experimental findings.</span></p>
A Study of UCB and MSCs in Children With CP: ACCeNT-CP
ClinicalTrials.gov study NCT03473301. IPD Sharing: NO. Countries: 1. Publications: 2.
Effects of rhythm and accent patterns on tempo-keeping property of finger tapping
Open the record for dataset details and reuse information.
ESCorpus-PE: A speech emotional dataset in Spanish with Peruvian accent
<p>ESCorpus-PE dataset contains emotional utterances of Spanish peruvian speech gathered from Spanish interviews, TV reports, political debate and testimonials. It contains 3749 utterances of three emotional dimensions: Valence, Arousal and Dominance. There are 80 speakers (44 male and 36 female). This data was created from Youtube audios. These audios were selected following a specific criteria specified in the paper: ESCorpus-PE: A speech emotional database in Spanish with Peruvian accent, in the section Methods/Audio/Video Selection. Anyone can use this data only for research purposes.</p> <p>More details on <br> https://github.com/Alessandra-UNSA/Peruvian_Spanish_Corpus</p>
Supplementary figures for paper "Should robots have accents?" published at IEEE RO-MAN 2020
<p>These plots show the preference towards a robot's accent (displayed on the x-axis), broken down by participants' region of origin in the UK. For example, "plot_Wales" shows that around 25% of the 23 participants from Wales indicated that they would like a robot to have an SSBE accent, around 15% a Welsh accent, etc.</p>
Anuran accents: Continental-scale citizen science data reveal spatial and temporal patterns of call variability
<p>Data and code associated with the 2020 publication in Ecology and Evolution (doi:10.1002/ece3.6833).</p>
Supplementary materials for "Topic affects perception of degree of foreign accent in a non-dominant language", published in Linguistics 59.1
<p>The file Rcode.R contains the code; the files EnglishData.csv and RussianData.csv contain the data; see the publication for a description of the data: "Topic affects perception of degree of foreign accent in a non-dominant language", Linguistics 59.1 .</p>
ACCENTING ON AFFECTIVE VARIABLES AND CONSIDERING PSYCHOLOGICAL CHARACTERISTICS OF STUDENTS WHILE TEACHING THE SECOND LANGUAGE.
Open the record for dataset details and reuse information.
Safety and Efficacy of the Accent Magnetic Resonance Imaging™ (MRI) Pacemaker and Tendril MRI™ Lead
ClinicalTrials.gov study NCT01576016. IPD Sharing: Not stated. Countries: 5. Publications: 0.
Accent Cardiac MRI Study
ClinicalTrials.gov study NCT02041702. IPD Sharing: UNDECIDED. Countries: 5. Publications: 0.
ACCENT: AMP945 in Combination with Nab-paclitaxel and Gemcitabine for Treatment of Pancreatic Cancer
ClinicalTrials.gov study NCT05355298. IPD Sharing: NO. Countries: 2. Publications: 0.
Accent MRI Pacemaker and Tendril MRI Lead New Technology Assessment
ClinicalTrials.gov study NCT01258218. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Black dress with lace accents
Source: Objaverse 1.0 / Sketchfab
SautiDB: Nigerian Accent Dataset Collection
<p>The SautiDB dataset collection project is an ongoing effort to collect datasets of various Nigerian accents. The dataset was collected in an uncontrolled manner, users who visit our <a href="https://sautidb.web.app/)">webapp</a> can record their voice and contribute to the dataset. The webapp uses the audio webapi to collect voice samples. We hope this dataset will be useful to people interested in developing voice technology in Nigeria. We will continuously collect more datasets and publish updated versions as we have them. This work grew out of our project <a href="https://www.k4all.org/project/accent-transfer/">Improving Online Experience using Accent Transfer</a>.<br> </p> <p>The filename is of the form nativeLanguage_fluentLanguage_speakerID_gender_sentenceID.wav, where</p> <ul> <li><strong>nativeLanguage:</strong> language spoken by the speaker's tribe. Native (mother) language of the speaker</li> <li><strong>fluentLanguage:</strong> language that the speaker feels best describes their accents</li> <li><strong>speakerID:</strong> ID, assigned to the speaker. It is possible for a speaker to have multiple IDs assigned since we are not authenticating users, we simply cached their browser sessions. </li> <li><strong>gender:</strong> gender of the speaker. We did not explicitly collect this information from users, we hand-labeled it. </li> <li><strong>sentenceID:</strong> the sentence ID for the sentences read. We used the <a href="http://www.festvox.org/cmu_arctic/cmuarctic.data">CMU Arctic sentences</a>.</li> </ul> <p><br> ===========================<br> Before Postprocessing<br> ===========================<br> Number of Samples: 1615<br> Size Webm: 59MB<br> Size Wav: 847MB<br> Sampling Rate: 48000Hz<br> Total Time: 2hrs 30min 21sec</p> <p><br> ============================<br> After Postprocessing<br> ============================<br> Number of Samples: 919<br> Size Wav: 336MB<br> Sampling Rate: 48000Hz<br> Total Time: 0hrs 59min 08sec<br> </p> <p>============================<br> Version 1.1<br> ============================<br> This version has two updates:</p> <p>1. In version 1.0, the naming convention for each language was to space each language with an underscore and uppercased, e.g., "Efik Ibibio" -> "EFIK_IBIBIO". We have changed "EFIK_IBIBIO" -> "EFIKIBIBIO". i.e. the file name, which was previously 'EFIK_IBIBIO_EFIK_IBIBIO_0014_M_A0138.wav', has now been changed to 'EFIKIBIBIO_EFIKIBIBIO_0014_M_A0138.wav'. This change applies only to languages that contain spaces. The rest of the filenames, therefore, remain unchanged, i.e. 'EDO_YORUBA_0053_M_B0389.wav' is still 'EDO_YORUBA_0053_M_B0389.wav'.</p> <p>2. We include an audio_metadata.csv file containing 'filename', 'nativeLanguage', 'fluentLanguage', 'speakerID', 'gender', 'sentenceID' and 'sentence', 'duration'. We hope this will make it easier for users to use our dataset for their work. The duration was calculated using the function 'librosa.get_duration()'.</p> <p>============================<br> Version 1.2<br> ============================<br> This version includes Hausa Langauge.</p> <p>============================<br> After Preprocessing<br> ============================<br> Number of Samples: 1137<br> Size Wav: 426MB<br> Sampling Rate: 48000Hz<br> Total Time: 1hrs 15min 24sec</p> <p><br> <br> The associated Github repository used for post-processing can also be found <a href="https://github.com/AISaturdaysLagos/sautidb_postprocessing_scripts">linked</a>. We are grateful for funding from AI4D-IndabaX with IDRC Grant Number: 109187-002.<br> <br> This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.