Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

17

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

17 results for “accentism”

Learn how ShareScore rates datasets ↗
zenodo40/100

The Accent Audio Files of Koshikijima Accent Database

<p>This is the accent audio files of&nbsp;Koshikijima Accent Database,&nbsp;an accent database of the endangered dialect of Koshikijima Japanese spoken on the Koshikijima Islands, Kagoshima Prefecture, Japan. It has been released for academic purposes,for both research and education. It contains the accent data of two speakers each from eight villages on the Islands, with the consent of the individual speakers to make their speech publicly available.&nbsp;The original data of the database comes from the ten-year-long collaborative fieldwork done by six editors. The database file can be downloaded from <a href="https://doi.org/10.15084/0002000053">https://doi.org/10.15084/0002000053</a>.</p> <p>Those who wish to use the data of this database to write an academic article or an essay are required to cite the following as a reference: &nbsp;Kubozono, Haruo, Zendo Uwano, Nobuko Kibe, Akiko Matsumori &amp; Tetsuo Nitta (eds.) (2016) Koshikijima Accent Database.<a href="https://www2.ninjal.ac.jp/koshikijima/">&lt;https://www2.ninjal.ac.jp/koshikijima/</a>&gt;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Accent Classification Dataset

<p>This dataset contains audio recordings of 12 different accents across the UK: Northern Ireland (NI), Scotland, Wales (SW), North East England (NE), North West England (NW), Yorkshire and Humber (YAH), East Midlands (EM), West Midlands (WM), East of England (EE), Greater London (GL), South East England (SE), South West England (SW). We split the data into a Male: Female ratio of 1:1, this is labelled with either &#39;_M&#39; for male or &#39;_F&#39; for female within the dataset. The audio dataset was compiled using opensource YouTube videos and it a collation of different accents, the audio files were trimmed for uniformity. The Audio files are of length 30 seconds, with the first 5 seconds and last 5 seconds of the signal being blank. We also resample the audio signals at 8 kHz, again for uniformity and to remove any noise present in the audio signals whilst retaining the underlying characteristics. The intended application of this dataset was to be used in conjunction with a deep neural network for accent and gender classification tasks.</p> <p>The dataset also contains an unseen dataset of the Google opensource digit dataset, which contains audio files of the digits 1-9. This is included to test any models developed using the original dataset to confirm model performance to data variations.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

The SIAEW Corpus of Spanish Iso-Accented English Words

<p>The SIAEW Corpus (Spanish Iso-Accented English Words) is a collection of monosyllabic English words in which one segment (the &#39;target&#39;) is replaced with its Spanish-accented counterpart, at one of 5 gradations of accentedness. Steps are equally-spaced in accentedness as judged by native listeners. The procedure used to generate the Corpus is described in detail in P&eacute;rez Ram&oacute;n, Garc&iacute;a Lecumberri and Cooke (submitted to Interspeech 2022); for a preprint contact rperez.ram@gmail.com.&nbsp;</p> <p>Please refer to the document SIAEW.pdf for a longer description of the SIAEW Corpus contents.</p>

opencc-by-4.0Mar 2022View details →
dryad36/100

Effects of rhythm and accent patterns on tempo-keeping property of finger tapping

<p><span>Tempo of music performance is often accelerated irrespective of players' intention. Though the characteristics of the tempo deviation phenomenon have been investigated using a finger-tapping task, most such studies dealt with tapping with a fixed interval; few studies considered the effects of rhythm and accent, both important factors of music performance. Here, we asked how different rhythm and accent patterns affected the tempo-keeping property using a synchronization-continuation task paradigm: Participants were asked to keep tapping while reproducing the given rhythm/accent patterns designated by the target tones. Tapping tempo was significantly deviated depending on rhythm/accent patterns, but their magnitudes were only several percent in 150 seconds, much smaller than those in real music performance. We also ran experiments under the conditions that participants need to reproduce the accent patterns, but only feedback tones were modulated. These auditory modulations affected the tempo deviation, implying that not only motor process for producing the accents but also perceptual process induced by the accented auditory feedback influence the tempo maintenance of finger tapping. In sum, the present finding shows that sensorimotor processing for un-uniform finger tapping can disturb the long-term tempo-keeping process. We also discussed related topics based on the experimental findings.</span></p>

opencc-zeroSep 2023View details →
ClinicalTrials.gov36/100

A Study of UCB and MSCs in Children With CP: ACCeNT-CP

ClinicalTrials.gov study NCT03473301. IPD Sharing: NO. Countries: 1. Publications: 2.

closedIPD-NOFeb 2026View details →
dryad36/100

Effects of rhythm and accent patterns on tempo-keeping property of finger tapping

Open the record for dataset details and reuse information.

publicAug 2024View details →
zenodo32/100

ESCorpus-PE: A speech emotional dataset in Spanish with Peruvian accent

<p>ESCorpus-PE dataset contains emotional utterances of Spanish peruvian speech gathered from Spanish interviews, TV reports, political debate and testimonials. It contains 3749 utterances of three emotional dimensions: Valence, Arousal and Dominance. There are 80 speakers (44 male and 36 female). This data was created from Youtube audios. These audios were selected following a specific criteria specified in the paper: ESCorpus-PE: A speech emotional database in Spanish with Peruvian accent, in the section Methods/Audio/Video Selection. Anyone can use this data only for research purposes.</p> <p>More details on&nbsp;<br> https://github.com/Alessandra-UNSA/Peruvian_Spanish_Corpus</p>

opencc-by-4.0Dec 2021View details →
zenodo28/100

Supplementary figures for paper "Should robots have accents?" published at IEEE RO-MAN 2020

<p>These plots show the preference towards a robot&#39;s accent (displayed on the x-axis), broken down by participants&#39; region of origin in the UK. For example, &quot;plot_Wales&quot; shows that around 25% of the 23 participants from Wales indicated that they would like a robot to have an SSBE accent, around 15% a Welsh accent, etc.</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Anuran accents: Continental-scale citizen science data reveal spatial and temporal patterns of call variability

<p>Data and code associated with the 2020 publication in Ecology and Evolution (doi:10.1002/ece3.6833).</p>

openother-openSep 2020View details →
zenodo28/100

Supplementary materials for "Topic affects perception of degree of foreign accent in a non-dominant language", published in Linguistics 59.1

<p>The file Rcode.R contains the code; the files EnglishData.csv and RussianData.csv contain the data; see the publication for a description of the data: &quot;Topic affects perception of degree of foreign accent in a non-dominant language&quot;, Linguistics 59.1 .</p>

opencc-by-4.0Nov 2020View details →
zenodo28/100

ACCENTING ON AFFECTIVE VARIABLES AND CONSIDERING PSYCHOLOGICAL CHARACTERISTICS OF STUDENTS WHILE TEACHING THE SECOND LANGUAGE.

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
ClinicalTrials.gov28/100

Safety and Efficacy of the Accent Magnetic Resonance Imaging™ (MRI) Pacemaker and Tendril MRI™ Lead

ClinicalTrials.gov study NCT01576016. IPD Sharing: Not stated. Countries: 5. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov28/100

Accent Cardiac MRI Study

ClinicalTrials.gov study NCT02041702. IPD Sharing: UNDECIDED. Countries: 5. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov24/100

ACCENT: AMP945 in Combination with Nab-paclitaxel and Gemcitabine for Treatment of Pancreatic Cancer

ClinicalTrials.gov study NCT05355298. IPD Sharing: NO. Countries: 2. Publications: 0.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov24/100

Accent MRI Pacemaker and Tendril MRI Lead New Technology Assessment

ClinicalTrials.gov study NCT01258218. IPD Sharing: Not stated. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo20/100

Black dress with lace accents

Source: Objaverse 1.0 / Sketchfab

opencc-byAug 2015View details →
zenodo16/100

SautiDB: Nigerian Accent Dataset Collection

<p>The SautiDB dataset collection project&nbsp;is an ongoing effort to collect datasets of various Nigerian accents. The dataset was collected in an uncontrolled manner, users who visit our <a href="https://sautidb.web.app/)">webapp</a>&nbsp;can record&nbsp;their voice and contribute to the dataset. The webapp uses the audio webapi to collect voice samples. We hope this dataset will be useful to people interested in developing voice technology in Nigeria. We will continuously collect more datasets and publish updated versions as we have them. This work grew out of our project <a href="https://www.k4all.org/project/accent-transfer/">Improving Online Experience using Accent Transfer</a>.<br> &nbsp;</p> <p>The filename is of the form nativeLanguage_fluentLanguage_speakerID_gender_sentenceID.wav, where</p> <ul> <li><strong>nativeLanguage:</strong> language spoken by the speaker&#39;s tribe. Native (mother) language of the speaker</li> <li><strong>fluentLanguage:</strong> language that the speaker feels best describes their accents</li> <li><strong>speakerID:</strong> ID, assigned to the speaker. It is possible for a speaker to have multiple IDs assigned since we are not authenticating users, we simply cached their browser sessions.&nbsp;</li> <li><strong>gender:</strong> gender of the speaker. We did not explicitly collect this information from users, we&nbsp;hand-labeled it.&nbsp;</li> <li><strong>sentenceID:</strong> the sentence ID for the sentences read. We used the <a href="http://www.festvox.org/cmu_arctic/cmuarctic.data">CMU Arctic sentences</a>.</li> </ul> <p><br> ===========================<br> Before Postprocessing<br> ===========================<br> Number of Samples: 1615<br> Size Webm: 59MB<br> Size Wav: 847MB<br> Sampling Rate: 48000Hz<br> Total Time: 2hrs 30min 21sec</p> <p><br> ============================<br> After Postprocessing<br> ============================<br> Number of Samples: 919<br> Size Wav: 336MB<br> Sampling Rate: 48000Hz<br> Total Time: 0hrs 59min 08sec<br> &nbsp;</p> <p>============================<br> Version 1.1<br> ============================<br> This version has two updates:</p> <p>1. In version 1.0, the naming convention for each language was to space each language with an underscore and uppercased, e.g., &quot;Efik Ibibio&quot; -&gt; &quot;EFIK_IBIBIO&quot;.&nbsp;We have changed &quot;EFIK_IBIBIO&quot; -&gt; &quot;EFIKIBIBIO&quot;. i.e. the file name, which was previously &#39;EFIK_IBIBIO_EFIK_IBIBIO_0014_M_A0138.wav&#39;, has now been changed to &#39;EFIKIBIBIO_EFIKIBIBIO_0014_M_A0138.wav&#39;.&nbsp;This change applies only to languages that contain spaces. The rest of the filenames, therefore, remain unchanged, i.e. &#39;EDO_YORUBA_0053_M_B0389.wav&#39; is still &#39;EDO_YORUBA_0053_M_B0389.wav&#39;.</p> <p>2. We include an audio_metadata.csv file containing &#39;filename&#39;, &#39;nativeLanguage&#39;, &#39;fluentLanguage&#39;, &#39;speakerID&#39;, &#39;gender&#39;, &#39;sentenceID&#39; and &#39;sentence&#39;, &#39;duration&#39;.&nbsp;We hope this will make it easier for users to use our dataset for their work. The duration was calculated using the function &#39;librosa.get_duration()&#39;.</p> <p>============================<br> Version 1.2<br> ============================<br> This version includes Hausa Langauge.</p> <p>============================<br> After Preprocessing<br> ============================<br> Number of Samples: 1137<br> Size Wav: 426MB<br> Sampling Rate: 48000Hz<br> Total Time: 1hrs 15min 24sec</p> <p><br> <br> The associated Github repository used for post-processing can also be found <a href="https://github.com/AISaturdaysLagos/sautidb_postprocessing_scripts">linked</a>. We are grateful for funding from AI4D-IndabaX with IDRC Grant Number: 109187-002.<br> <br> This work is licensed under a <a href="http://creativecommons.org/licenses/by-nc/4.0/">Creative Commons Attribution-NonCommercial 4.0 International License</a>.</p>

restrictedFeb 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record