Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.9.0
Dataset results
166 results for “dialect”
Comparative wordlist of Tibetan dialects
<p>This is a collection of Tibetan dialect material compiled by Roland Bielmeier and his collaborators, mostly from existing sources, but also making use of original fieldwork.</p>
Annotated Corpus for the Alsatian Dialects
<p>This corpus contains a collection of texts in the Alsatian dialects which were manually annotated with parts-of-speech, lemmas, translations into French and location entities.</p><p>The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 21 documents and 12,907 syntactic words. The annotation process is detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704806">http://hal.archives-ouvertes.fr/hal-01704806</a></p><p><strong>Information about version 3</strong></p><p>Version 3 corrects some minor errors in the CONLL-U files: wrong token indexes after multiword tokens and missing _ in glosses. In addition, all files are concatenated into a single CONLL-U file.</p><p><strong>Information about version 2</strong></p><p>Version 2 contains the same annotated documents as version 1, but some errors have been corrected and the annotated corpus is provided in the <a href="http://universaldependencies.org/format.html">CoNLL-U format</a></p><p>The untokenised and unannotated versions of the documents are found in the "txt" folder. The annotated versions of the documents are found in the "ud" folder (<a href="http://universaldependencies.org/format.html">CoNLL-U format</a>).</p><p>In addition to the form, the lemma and the part-of-speech additional information is also provided:</p><ul><li>translation of the lemma into French (Gloss field)</li><li>annotation of location names (NamedType field)</li></ul>
Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying
<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p> </p> <p>* To cite this dataset refer to <a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., & Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p> </p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>
Sentiment dataset of Algerian dialect
<p>* This sentiment dataset of Algerian dialect consists of 11760 comments (6111 positive/ 5649 negative comments)) collected from (Facebook, YouTube and Twitter) during Hirak 2019.<br>* Comments concern the Algerian spoken language, written in Arabic and/or Latin characters and/or Arabizi, which could be either Modern Standard Arabic, French or local dialect.<br>* Value ‘1’ is attributed for Positive review / value ‘0’ attributed for Negative review.<br>* Due to the nature of this Dataset, some comments contain offensive language. This does not reflect author values, however the aim is to providing a resource to help in analysing positive and negative sentiments (that probably containing harmful content).<br>* For more information please contact (@Ahmed Cherif Mazari) : <a href="mailto:mazari.ac@gmail.com" target="_blank" rel="noopener">mazari.ac@gmail.com</a></p>
Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic
<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>
Fig. 3 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations
Fig. 3. Dendrogram of territorial complexes Lower-Dnipro Area (localities are described in table 1).
Fig. 1 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations
Fig. 1. Components of Chaffinch song structure: 1 — phrase; 2 — inserted element; 3 — pre-flourish; 4 — flourish. a — element; b — sub-element.
Fig. 2 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations
Fig. 2. Study area and Lower-Dnipro dialect territory:> 50 — localities with more than 50 % specific dialectal Lower Dnipro song types in song complex; 26–50 — specific dialectal types constitute 26–50 % of song complex; 5–25 — dialectal types constitute 5–25 %; 0 — specific dialectal types were not found.
Fig. 5 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations
Fig. 5. Specific "southern" elements in Lower-Dnipro dialect. Elements found in song complexes of South Ukraine: LD — Lower-Dnipro dialect; SE — South-East sub-dialect of Left-bank dialect; Cr — Crimean dialect; Da — Danube dialect; Cp — Carpathian dialect.
Estonian dialect models
<p>Dialectal data and normalization models presented in the following paper:</p> <p>Hämäläinen, M., Alnajjar, K., & Tuisk, T. (2022) <em>Help from the Neighbors: Estonian Dialect Normalization Using a Finnish<br> Dialect Generator</em>. In <em>The Proceedings of The Third Workshop for Deep Learning for Low Resource NLP.</em></p>
New version of the digitized Dialect Atlas of Finnish by Lauri Kettunen
<p>This package includes alternative representations of the digitized version of Lauri Kettunen's Dialect Atlas of Finnish (Kettunen 1940; <a href="http://urn.fi/urn:nbn:fi:csc-kata20151130145346403821" rel="nofollow">http://urn.fi/urn:nbn:fi:csc-kata20151130145346403821</a>). The first version of the digitized data was prepared by profs. Sheila Embleton and Eric Wheeler (York University, Canada) (Embleton & Wheeler 1997, 2000), and further refined for publication by Jyri Lehtinen under the BEDLAN project (Biological Evolution and the Diversification of Languages).</p> <p>The alternative representation provided here includes the data formatted as (1) multistate character format and (2) binary character format. Both formats resemble the representation of population genetic data, and are thus easier to apply to population genetic analyses than the representation provided through the Fairdata repository. Also the linguistic classification of the collected linguistic traits is provided.</p> <p>The data has been prepared by the BEDLAN research project (Biological Evolution and the Diversification of Languages), which has conducted research of Kettunen's dialect atlas using population genetic techniques (Syrjänen et al. 2016, Honkola et al. 2018, Santaharju et al. in revision).</p> <p>Citation: Santaharju, Jenni, Kaj Syrjänen, Terhi Honkola, Perttu, Seppä, Outi Vesakoski and Unni Leino (submitted revision): New version of the digitized Dialect Atlas of Finnish by Lauri Kettunen. In Digital Humanities in the Nordic and Baltic Countries Publications. Zenodo. <a href="https://doi.org/10.5281/zenodo.10078078" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10078078</a></p> <div> <h2>kettunen_multistate.csv</h2> </div> <p>This file contains the multistate representation of the data. Here different states are coded as integers representing the different map symbols from Kettunen's original dialect atlas (for an example, see <a href="http://kettunen.fnhost.org/" rel="nofollow">http://kettunen.fnhost.org/</a>, which provides scans of Kettunen's original atlas map pages by Juha Kuokkala).</p> <ol> <li> <p><code>Municipality_number</code></p> <p>Unique identifier for each data point (municipality).</p> </li> <li> <p><code>Municipality_name</code></p> <p>Municipality name. These may not be unique, as several municipalities may have the same name.</p> </li> <li> <p><code>lon_WGS84</code></p> <p>WGS 84 longitude of the municipality centroid.</p> </li> <li> <p><code>lat_WGS84</code></p> <p>WGS 84 latitude of the municipality centroid.</p> </li> <li> <p><code>1a - 213c</code></p> <p>Dialect features present on each map page. The integer at the beginning matches the map page in Kettunen's atlas, so e.g. page 1 starts with 1. The letter following the integer represents different overlapping variants within a single map page. Each page number is suffixed with as many different letters as there are overlapping variants on that page; for instance page 1 has at most 3 overlapping variants for any datapoint, so the table includes three columns (1a-1c). The values are either integers (representing different map symbols from Kettunen's map page legends, with 1 being the topmost variant, 2 being the second variant, and so on), "-" (for absent data) or "NA" (for missing data).</p> </li> </ol> <div> <h2>kettunen_binary.csv</h2> </div> <ol> <li> <p><code>Municipality_number</code></p> <p>Unique identifier for each data point (municipality).</p> </li> <li> <p><code>Municipality_name</code></p> <p>Municipality name. These may not be unique, as several municipalities may have the same name.</p> </li> <li> <p><code>lon_WGS84</code></p> <p>WGS 84 longitude of the municipality centroid.</p> </li> <li> <p><code>lat_WGS84</code></p> <p>WGS 84 latitude of the municipality centroid.</p> </li> <li> <p><code>1_2 - 213_16</code></p> <p>Dialect features present on each map page. The leftmost integer matches the map page in Kettunen's atlas, so e.g. page 1 starts with 1. The second integer matches the multistate characters found in the cells of kettunen_multistate.csv, and thus reflect the different symbols in Kettunen's map page legends. Again here, 1 represents the topmost box in Kettunen's map page's legend, 2 the second box, and so on. The actual data field contains either 0 (absent), 1 (present) or "NA" (missing).</p> </li> </ol> <div> <h2>kettunen_map_explanations.csv</h2> </div> <ol> <li> <p><code>Map number</code></p> <p>Unique identifier for each map.</p> </li> <li> <p><code>The explanation of the map</code></p> <p>Map name and description of the dialect feature.</p> </li> <li> <p><code>Level1 - Level5</code></p> <p>Hierarchical classification of the dialect feature.</p> </li> </ol> <div> <h2>References</h2> </div> <p>Embleton, Sheila M. and Eric Wheeler, S. 1997. Finnish dialect atlas for quantitative studies. <em>Journal of Quantitative Linguistics</em> 4.99-102. <a href="https://doi.org/10.1080/0929617970859008" rel="nofollow">DOI: https://doi.org/10.1080/0929617970859008</a></p> <p>Embleton, Sheila, M. and Eric Wheeler, S. 2000. Computerized dialect atlas of Finnish: Dealing with ambiguity. <em>Journal of Quantitative Linguistics</em> 7.227-31. <a href="https://doi.org/10.1076/jqul.7.3.227.4109" rel="nofollow">DOI: https://doi.org/10.1076/jqul.7.3.227.4109</a></p> <p>Honkola, Terhi, Kalle Ruokolainen, Kaj Syrjänen, Unni-Päivä Leino, Ilpo Tammi, Niklas Wahlberg, and Outi Vesakoski. 2018. Evolution within a Language: Environmental differences contribute to divergence of dialect groups. <em>BMC Evolutionary Biology</em> 18, no. 132. <a href="https://doi.org/10.1186/s12862-018-1238-6" rel="nofollow">DOI: https://doi.org/10.1186/s12862-018-1238-6</a></p> <p>Kettunen, Lauri. 1940. <em>Suomen Murteet III A. Murrekartasto.</em> Helsinki: Suomalaisen kirjallisuuden seura.</p> <p>Santaharju, Jenni, Terhi Honkola, Perttu, Seppä, Kaj Syrjänen, Unni Leino and Outi Vesakoski (in revision): Linguistic convergence and its drivers in Finnish dialects.</p> <p>Syrjänen, Kaj, Terhi Honkola, Jyri Lehtinen, Antti Leino and Outi Vesakoski. 2016. Applying population genetic approaches within languages: Finnish dialects as linguistic populations. <em>Language Dynamics and Change</em> 6.235-83. <a href="https://doi.org/10.1163/22105832-00602002" rel="nofollow">DOI: https://doi.org/10.1163/22105832-00602002</a></p>
A Monologue Narrative Text of the Itoman Dialect of Okinawan: How to Worship in Yukkanuhii
<p>This paper provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. He describes the way worship is conducted in Yukkanuhi. Japanese translations, morphological analyses, and interlinear glosses are provided. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>
Data supplement to: "Refractory depression - Mechanisms & Efficacy of Radically Open Dialectical Behaviour Therapy (RefraMED): findings of randomised trial on benefits and harms"
<p>Data and code to support primary analyses reported in "Refractory depression – mechanisms and efficacy of radically open dialectical behaviour therapy (RefraMED): findings of a randomised trial on benefits and harms". British Journal of Psychiatry. doi: 10.1192/bjp.2019.53</p>
Data set on linguistic similarity of German dialects
<p>- The data provide information on pairwise dialect similarities for 439 NUTS 3 regions in Germany</p> <p><em>- </em>Data come from the maps and questionnaires of the Sprachatlas des Deutschen Reichs, digitized using ArcGIS software</p> <p>- The data source is a questionnaire with translations of standardized German sentences into local dialects between 1879 and 1888</p> <p>- The measure is defined as the number of co-occurrences for all pairs of sites (z-scaled)</p>
A Chin dialect survey (Part 1 of 2)
<p>This dataset was compiled from 2005-2014 and includes a survey of Chin dialects, primarily southern Chin, spoken in Burma. The aim of the survey was to investigate mutual intelligibility and the diversity of the Chin family.</p>
A Chin dialect survey (Part 2 of 2)
<p>This dataset was compiled from 2005-2014 and includes a survey of Chin dialects, primarily southern Chin, spoken in Burma. The aim of the survey was to investigate mutual intelligibility and the diversity of the Chin family.</p>
Yucatec Maya Dialect Atlas
<p>The "Yucatec Maya Dialect Atlas" is a dataset created at the Universidad Autónoma de Yucatán, Mérida, México. The data was collected between 2000 and 2007; funded by the CONACyT (funding ID: 36387-H). The data was collected in 80 different locations in the peninsula of Yucatán (n speakers = 157). Elicitation was based on a questionnaire with Spanish prompts (n = 665).</p> <p> </p>
Sudanese dialect speech dataset
<p>This is speech dataset for the Sudanese dialect data been collected from YouTube videos represent the characteristics of the Sudanese dialect, mainly the middle of Sudan dialect -Khartoum in particular- and have some northern tendency, primarily two programs Hajj Muzakir and Dukkan Wad Elbaseer.</p> <p>Transcription is done manually by listening to the audio files repeatedly to write the captions for the collected conversations to make sure that every word is written as said by the speakers. Transcription is written without diacritics on the Arabic alphabet, in a manner that reflects the Sudanese way of speaking, therefore, any correction to the noticeable mistakes was not applied to get rid of any biases and make the data representative.</p> <p>The 'Dataset' subdirectory contains all the audio and text files for the corpus, the files organized based on program name 'hm_' for Hajj Muzakir program and 'wb_' Dukkan Wad Elbaseer, each filename follows three categories first two litters for the program name 'hm' or 'wb', second the number of the episode third the number of the clip, hm_01_0001.wav and wb_01_0001.wav represent first episode of each program and the first clip.</p>
Data and Code Supplement to: "Processes of change in a randomized clinical trial of Radically Open Dialectical Behavior Therapy (RO DBT) for adults with treatment refractory depression"
<p>Dataset to support secondary analyses reported in "Processes of change in a randomized clinical trial of Radically Open Dialectical Behavior Therapy (RO DBT) for adults with treatment refractory depression" in the Journal of Consulting and Clinical Psychology</p>
*-koatē/-škoatē inchoative verbs in Skolt and Kola Saami dialects
<p>The file contains two sets of data concerning *<em>-koatē/-škoatē</em> inchoative verbs in Skolt and Kola Saami dialects. The first one is a schematic representation of the verbs found in KKLS [1] (alphabetical span of A–K) with a focus on the derivational variants *<em>-koatē</em> vs. *<em>-škoatē</em>. Two versions are presented on separate sheets, one with a single alphabetical listing and another with grouping according to the (original) syllable count of the verb base. The columns P...V code the variants occurring in different dialects (codes: k = *-koatē, s = *-škoatē, ks = both).</p> <p>[1] KKLS = T. I. Itkonen 1958: <em>Koltan- ja kuolanlapin sanakirja. Wörterbuch des Kolta- und Kolalappischen</em>. I–II. Helsinki: Société Finno-Ougrienne.</p> <p>The second data set contains the *<em>-koatē/-škoatē</em> verb lexemes found in Szabó 1968 and 1987 [2, 3]. The data is grouped according to (original) syllable count, suffix variant (*<em>-koatē</em> vs. *<em>-škoatē</em>) and language area (Kildin vs. Ter [Turja]). Only one occurrence of each derivative lexeme is quoted.</p> <p>[2] Szabó, László 1968: <em>Kolalappische Volksdichtung (Texte aus den Dialekten in Kildin und Ter)</em>. Zweiter Teil nebst grammatischen Aufzeichnungen. Göttingen: Vandenhoeck & Ruprecht.</p> <p>[3] Szabó, László 1987: The use of the inchoative in Kola-Sami sentences. – <em>Nordlyd </em>13: 70–103.</p> <p>This is a data set supplement of: Kuokkala, J. (2019). Saami -<em>(š)goahti</em> inchoatives: their variation, history, and suggested cognates in Veps and Mordvin. <em>Suomalais-Ugrilaisen Seuran Aikakauskirja</em> 97, 153–181. <a href="https://doi.org/10.33340/susa.75718">https://doi.org/10.33340/susa.75718</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.