Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

166

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

166 results for “dialect”

Learn how ShareScore rates datasets ↗
zenodo40/100

The Albanian Dialect Corpus

<p>This dataset provides the following information:</p> <ul> <li> <p><strong><code>data</code></strong>&nbsp;folder: contains three folders that group tweets for each region</p> <ul> <li><strong><code>al</code></strong>&nbsp;folder: tweets posted by users in Albania;</li> <li><strong><code>mk</code></strong>&nbsp;folder: tweets posted by users in North Macedonia;</li> <li><strong><code>ks</code></strong>&nbsp;folder: tweets posted by users in Kosovo;</li> </ul> <p><code>Each file in any of these folders represents tweets of a specific user and each line of any file represents a single tweet.</code></p> </li> </ul> <p>In order to maintain confidentiality and protect user privacy, all identifying information about the user, including their username, has been hidden. This is done to ensure that no personally identifiable information is revealed and to respect the user&#39;s right to privacy. Users are represented in this format&nbsp;<code>{no}</code>.txt where&nbsp;<code>no</code>&nbsp;correspond to the ordered usernames of the users in their respective folders.</p> <p>Finally, this corpus was used in our work&nbsp;<code><strong>Using</strong> <strong>Twitter to Collect a Multi-Dialectal Corpus of Albanian using advanced geotagging and dialect</strong> <strong>modeling</strong></code>.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Dialectal Finnish Generators

<p>H&auml;m&auml;l&auml;inen, M., Partanen, N., Alnajjar, K., Rueter J. &amp; Poibeau T. (2020). Automatic Dialect Adaptation in Finnish and its Effect on Perceived Creativity. In <em>Proceedings of the 11th International Conference on Computational Creativity</em></p> <p>&nbsp;</p> <p>OpenNMT-tf models for generating dialectal Finnish from standard Finnish. The input should be a chunk of 3 words, split into characters and separated by an underscore (<em>koira juoksee kovaa</em> should be&nbsp;<em>k o i r a _ j u o k s e e _ k o v a a</em>). The flag model also expects a dialect name to be appended to the input&nbsp;(<em>Kainuu&nbsp;k o i r a _ j u o k s e e _ k o v a a).</em></p> <p>Take a look at the README.md for more information.</p>

opencc-by-4.0Aug 2020View details →
dryad36/100

Familiarity, homogeneity, and discrimination of song dialects: Data and playback study stimuli

<p>Male songbirds of many species sing local song dialects that are restricted to defined geographical areas. In most tests of responses to local versus foreign dialects, males respond more aggressively to songs from their own dialect, presumably because local males represent more of a threat to their success. We asked how hearing foreign songs during development and territory establishment affects discrimination of the local dialect in wild Savannah sparrows, <em>Passerculus sandwichensis</em>. After foreign songs had been heard from loudspeakers in the study area in at least two consecutive breeding seasons, males reduced the intensity of their responses to the local version of population-specific buzz segment of the song. Four years after the foreign songs were last broadcast on the study area, males again responded more aggressively to the local version of the buzz. As for the basis of these responses, we found no evidence that birds discriminated among dialects by comparing them to their own songs. However, auditory experience with a foreign song, whether during song development (from speaker-simulated song tutors) or during the current breeding season (from neighbours' songs), reduced the intensity of birds' responses to the local buzz type. Both familiarity, in the form of auditory experience with a song type, and homogeneity, when a song type is sung by all or nearly all of the population, appear to contribute to heightened aggressive responses to a local song dialect.</p>

opencc-zeroJan 2024View details →
zenodo36/100

Dvoice : An open source dataset for Automatic Speech Recognition on African Languages and Dialects

<p>DVoice is a community initiative that aims to provide African languages and dialects with data and models to facilitate their use of voice technologies. The lack of data on these languages makes it necessary to collect data using methods that are specific to each language. Two different approaches are currently used: the DVoice platform, which is based on Mozilla Common Voice, for collecting authentic recordings from the community, and transfer learning techniques for automatically labeling the recordings. The DVoice platform currently manages 7 languages including Darija (Moroccan Arabic dialect) whose dataset appears on this version, Wolof, Mandingo, Serere, Pular, Diola and Soninke. The Swahili-labeled data present in this version was obtained after automatic labeling via the learning transfer of the Voxlingua107 dataset. For a first time, we also advocate for the increase of data given their small size that we currently have. Thus this version of the dataset contains easily identifiable augmented data.</p>

opencc-by-4.0Mar 2022View details →
zenodo36/100

Understanding SQL Dialects and Testing Their Implementations

<p>The artifact contains:</p> <ul> <li>The SQL test suite, which is the main contribution of our paper.</li> <li>A DuckDB database that contains all the raw data based on which we made our conclusions.</li> <li>The scripts used to generate the results as well as the tool to extract SQL statements from Java projects in the folder Artifact Files.</li> </ul>

opencc-by-4.0Mar 2022View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p><strong>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</strong></p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Training a Text-to-Speech System for Dialectal Arabic with a Focus on the Iraqi Dialect

<p>This research introduces a novel approach to Text-to-Speech (TTS) synthesis, focusing on the phonetic complexities of Arabic dialects, with particular emphasis on the Iraqi dialect. While existing Arabic speech corpora provide substantial coverage of Modern Standard Arabic (MSA), they fall short in capturing the phonetic richness of regional dialects. To address this gap, we utilized Nawar Halabi's Arabic Speech Corpus as a base dataset and enriched it with custom-recorded samples of the Iraqi dialect, incorporating distinctive phonemes such as گ ,ڤ ,پ ,چ ,ۆ ,ڵ ,ێ, and using the Tatweel character (ـ) as a vowel. Our approach, powered by the FastPitch model and a customized phonetiser, successfully synthesized the Iraqi dialect while also demonstrating adaptability to other Arabic dialects, including Egyptian, Khaliji, Syrian, and more. The results of this research signify a promising advancement in Arabic TTS technology, expanding its scope to authentically represent the diverse linguistic landscape of the Arabic-speaking world.</p>

opencc-by-4.0May 2024View details →
zenodo36/100

DB3V: A Dialect Dominated Dataset of Bird Vocalisation for Cross-corpus Bird Species Recognition

<p>The first cross-corpus dataset that focuses on dialects in bird vocalisations. The DB3V comprises more than 25 hours of audio recordings from 10 bird species distributed across three distinct regions in the contiguous United States (CONUS).</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

CLDF dataset derived from the LSDO's "Chin Dialect Data" from 2019

<p>Cite the source of the dataset as:</p> <blockquote> <p>Language and Social Development Organization (2019): Chin dialect data collection. Yangon: LSDO.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Tesseract OCR models for the Alsatian dialects

<p>This dataset provides trained Tesseract (<a href="https://github.com/tesseract-ocr/tesseract">https://github.com/tesseract-ocr/tesseract</a>) OCR models for the Alsatian dialects. These models were developed in the context of the RESTAURE project, funded by the French ANR.&nbsp;</p> <p>Two models are provided :</p> <p>The first model, ISKO_2015, has been presented in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01252241">https://hal.archives-ouvertes.fr/hal-01252241</a>. The Tesseract model has been trained using the jTessBoxEditor tool (<a href="http://vietocr.sourceforge.net/training.html">http://vietocr.sourceforge.net/training.html</a>), Version 1.4 (2 May 2015), based on images automatically generated from the training texts (excerpts from 7 different printed works, totalling about 9,000 words). The generation of the images used a 36pt font size, and two fonts were used (Arial and Times New Roman), with their normal and italic variants.<br> The Tesseract model (gsw.traineddata) can be used with Tesseract 3.0x.</p> <p>The second model, 2018, has been trained for Tesseract 4.0x, using jTessBoxEditor version 2.0.1 (28 July 2018). Again, images were automatically generated from the training text. The training text is different from the one used for the ISKO_2015 model and is &quot;artificial&quot;, in the sense that it has been built by appending word n-grams extracted from a large variety of published texts in Alsatian, for a time period spanning 2 centuries and for different text genres. The images corresponding to this training text have been automatically generated with the Tesseract text2image tool, using the following parameters: --ptsize=36 --leading=20. The fonts used are listed in the gsw.font_properties file.</p> <p>Dictionary data has also been used for training. We conflated Alsatian words found in several lexicons and corpora:</p> <ul> <li>Lexicons produced by the OLCA (Office pour la Langue et les Cultures d&#39;Alsace et de Moselle): <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Lexicon from a Wiktionary user page: <a href="http://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais">https://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais</a></li> <li>Lexicon from the ACPA association: <a href="http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm">http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm</a></li> <li>Chronicles published by Raymond Matzen in the local newspaper &quot;Les Derni&egrave;res Nouvelles d&#39;Alsace&quot;</li> <li>Transcriptions of television shows found in Erhart, P. (2012). <em>Les dialectes dans les m&eacute;dias: quelle image de l&rsquo;Alsace v&eacute;hiculent-ils dans les &eacute;missions de la t&eacute;l&eacute;vision r&eacute;gionale?</em>, Universit&eacute; de Strasbourg, <a href="http://www.theses.fr/167563386">http://www.theses.fr/167563386</a></li> <li>French-Alsatian parallel corpus provided by the OLCA</li> <li>Excerpts from Adolf, P. (2006). <em>Dictionnaire comparatif multilingue: fran&ccedil;ais-allemand-alsacien-anglais.</em>, Strasbourg, France, Midgard, 2006, 373 p.</li> </ul> <p>The Tesseract models can be used&nbsp; for instance using the gImageReader tool (<a href="https://github.com/manisandro/gImageReader">https://github.com/manisandro/gImageReader</a>), which provides a graphical user interface for the Tesseract tool.&nbsp;</p> <p>When evaluated against the same test corpus (prose by Marie Hart, theater and poetry by Gustave Stokopf and prose by Charles Zumstein, totalling about 4,900 words), both models achieve roughly the same performance levels. Usually, even better performance levels can be achieved by combining the Alsatian-specific model with the French and German models available for Tesseract (available from <a href="https://github.com/tesseract-ocr/tessdata">https://github.com/tesseract-ocr/tessdata</a>)</p>

opencc-by-sa-4.0Aug 2018View details →
zenodo36/100

Citation Tone of Lugang dialect in Chaonan District of Shantou City in Guangdong Province, China

<p>Citation Tone of Lugang dialect in Chaonan District of Shantou City in Guangdong Province, China&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

cldf-datasets/normansinitic: Structural and lexical data for the paper by Norman (2013) on Chinese dialect classification

<p>Original source of the data:</p> <blockquote> <p>Norman, J. (2003): Chinese dialects. Phonology. In: Thurgood, G. &amp; LaPolla, R.: The Sino-Tibetan Languages. Routledge: London and New York. 72-83.</p> </blockquote>

opencc-by-4.0Nov 2019View details →
zenodo36/100

CLDF dataset accompanying Auderset et al.'s "Subgrouping in a dialect continuum" from 2023

<p>Cite the source of the dataset as:</p> <blockquote> <p>Auderset, Sandra, Simon J. Greenhill, Christian T. DiCanio, Eric W. Campbell. (2023) &quot;Subgrouping in a `dialect continuum&#x27;: A Bayesian phylogenetic analysis of the Mixtecan language family&quot;. Journal of Language Evolution 8 (1). 33–-63. DOI: https://doi.org/10.1093/jole/lzad004.</p> </blockquote>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Data for Finnish Dialect Detection

<p>The data used in the paper &quot;Finnish Dialect Identification: The Effect of Audio and Text&quot;.</p> <p><strong>If you use the data, please cite:</strong></p> <p>H&auml;m&auml;l&auml;inen, Mika; Alnajjar, Khalid; Partanen, Niko &amp; Rueter, Jack (2021).&nbsp;<a href="https://aclanthology.org/2021.emnlp-main.692/">Finnish Dialect Identification: The Effect of Audio and Text</a>. In&nbsp;<em>Proceedings of the 2021&nbsp;Conference on Empirical Methods in Natural Language Processing (EMNLP)</em>.</p> <p>The <strong>metadata.json </strong>contains the dialectal and normalized transcriptions, the length of the wav files in milliseconds, the path to the wav file, the role of the speaker and the dialect. The <strong>data.zip*</strong> files contain the wav files. They are partial zip files to make Zenodo upload easier. <strong>Code_models.zip</strong> contains the code for training the bimodal model and the final trained model presented in the paper. There is also an OpenNMT model that is the text only model described in the paper. You can use it by running:&nbsp;<em>python3 -m onmt.bin.translate -model nmt-model_step_100000.pt -src source_test.txt -output pred.txt -replace_unk</em></p> <p>Our dataset is licensed under CC BY NC ND 4.0. Academic use only.</p> <p>&nbsp;</p> <p>The corpus is based on <a href="http://urn.fi/urn:nbn:fi:lb-2017020201">Suomenkielen n&auml;ytteit&auml;</a>&nbsp;(CC BY &nbsp;<a href="http://www.kotus.fi/">Kotimaisten kielten keskus</a>).</p>

opencc-by-nc-nd-4.0Aug 2021View details →
dryad36/100

Familiarity, homogeneity, and discrimination of song dialects: Data and playback study stimuli

Open the record for dataset details and reuse information.

publicJan 2024View details →
dryad32/100

Data from: Nestling and adult sparrows respond differently to conspecific dialects

Understanding the causes and consequences of divergence in mate recognition traits has long been a fundamental question in evolutionary biology. In songbirds, songs are culturally transmitted, and cultural divergence can generate discrete geographic variation in song (i.e., dialects). Understanding how responses to within- versus across-species variation in songs changes across life stages may shed light on the functional significance of population divergence in learned traits. Here, we use a novel combination of song playbacks to adult and nestling golden-crowned sparrows to compare responses to local conspecific, foreign conspecific and heterospecific songs prior to and after song learning. We found that nestlings respond equally little to both foreign conspecific and heterospecific songs. By contrast, the response of adult males to foreign conspecific songs was stronger than their response to heterospecific song, but weaker than their response to local conspecific song. Our study suggests that early local experience may interact with conspecific biases prior to song learning, in a way that has not been previously documented. Our results illustrate the importance of studying behavior at multiple life stages in order to better understand the effect of early experience on cultural and biological evolution.

opencc-zeroDec 2017View details →
dryad32/100

Data from: Female, but not male, tropical sparrows respond more strongly to the local song dialect: implications for population divergence.

In addition to the observed high diversity of species in the tropics, divergence among populations of the same species exists over short geographic distances in both phenotypic traits and neutral genetic markers. Divergence among populations suggests great potential for the evolution of reproductive isolation and eventual speciation. In birds, song can evolve quickly through cultural transmission resulting in regional dialects, which can be a critical component of reproductive isolation through variation in female preference. We examined female and male behavioral responses to local and non-local dialects in two allopatric populations of rufous-collared sparrows (Zonotrichia capensis) in the Andes of Ecuador. Here we show that female sparrows prefer their natal song dialect to the dialect from an allopatric population just 25 km away and separated by unsuitable higher elevation habitat (pass 4200 m), thus providing evidence of prezygotic reproductive isolation among populations. Males showed similar territorial responses to all conspecific dialects, with no consistent difference with respect to distance, making male territoriality uninformative for estimating reproductive isolation. This study provides novel evidence for culturally-based prezygotic isolation over very short distances in a tropical bird.

opencc-zeroDec 2010View details →
zenodo32/100

FIGURES 4a–o in Song dialects as diagnostic characters - acoustic differentiation of the Canary Island Goldcrest subspecies Regulus regulus teneriffae Seebohm 1883 and R. r. ellenthalerae Päckert et al. 2006 (Aves: Passeriformes: Regulidae)

FIGURES 4a–o. Territorial song of Regulus regulus teneriffae, song types indicated on the sonagrams, lower left. La Gomera: a–d, song type A, strophes of four different males; Tenerife: e, g–h, song type A1 withfixed motiv I, strophes of three different males; f, rare song type F; i, song type B, male from Esperanza Forest; k–m, song type B, three different males from Anaga Mts; n, song type C, male from Anaga Mts; o, song type D, male from Anaga Mts.

opennotspecifiedSep 2006View details →
zenodo32/100

FIGURE 2 in Song dialects as diagnostic characters - acoustic differentiation of the Canary Island Goldcrest subspecies Regulus regulus teneriffae Seebohm 1883 and R. r. ellenthalerae Päckert et al. 2006 (Aves: Passeriformes: Regulidae)

FIGURE 2. Distribution of Gold­ and Firecrests on the Atlantic islands. Distribution of dialect song types from the Canary Islands is indicated by symbols on the according island.

opennotspecifiedSep 2006View details →
zenodo32/100

FIGURES 1a–d in Song dialects as diagnostic characters - acoustic differentiation of the Canary Island Goldcrest subspecies Regulus regulus teneriffae Seebohm 1883 and R. r. ellenthalerae Päckert et al. 2006 (Aves: Passeriformes: Regulidae)

FIGURES 1a–d. Vocalization pattern of Goldcrests; a, bipartite song of European Goldcrests, R. r. regulus, Czech Republik; b, tripartite song of R. r. azoricus, São Miguel; syntax: introduction (unit 1) + mid part (ascending phrase, unit 2) + final part (terminal flourish, unit 3); c, parameters measured in sonagraphic analysis indicated by horizontal and vertical lines; d, subsong of R. r. japonensis, Far East Siberia; from a continuous performance of about two minutes.

opennotspecifiedSep 2006View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record