Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

166

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

166 results for “dialect”

Learn how ShareScore rates datasets ↗
zenodo40/100

Comparative wordlist of Tibetan dialects

<p>This is a collection of Tibetan dialect material compiled by Roland Bielmeier and his collaborators, mostly from existing sources, but also making use of original fieldwork.</p>

opencc-by-4.0Jun 2017View details →
zenodo40/100

Annotated Corpus for the Alsatian Dialects

<p>This corpus contains a collection of texts in the Alsatian dialects which were manually annotated with parts-of-speech, lemmas, translations into French and location entities.</p><p>The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 21 documents and 12,907 syntactic words. The annotation process is detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704806">http://hal.archives-ouvertes.fr/hal-01704806</a></p><p><strong>Information about version 3</strong></p><p>Version 3 corrects some minor errors in the CONLL-U files: wrong token indexes after multiword tokens and missing _ in glosses. In addition, all files are concatenated into a single CONLL-U file.</p><p><strong>Information about version 2</strong></p><p>Version 2 contains the same annotated documents as version 1, but some errors have been corrected and the annotated corpus is provided in the <a href="http://universaldependencies.org/format.html">CoNLL-U format</a></p><p>The untokenised and unannotated versions of the documents are found in the "txt" folder. The annotated versions of the documents are found in the "ud" folder (<a href="http://universaldependencies.org/format.html">CoNLL-U format</a>).</p><p>In addition to the form, the lemma and the part-of-speech additional information is also provided:</p><ul><li>translation of the lemma into French (Gloss field)</li><li>annotation of location names (NamedType field)</li></ul>

opencc-by-sa-4.0Feb 2018View details →
zenodo40/100

Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying

<p>Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying.</p> <p>&nbsp;</p> <p>* To cite this dataset refer to&nbsp;<a href="http://dx.doi.org/10.12785/ijcds/130177" target="_blank" rel="nofollow noopener">http://dx.doi.org/10.12785/ijcds/130177</a><br>Mazari, A. C., &amp; Kheddar, H. (2023). "Deep Learning-based Analysis of Algerian Dialect Dataset Targeted Hate Speech, Offensive Language and Cyberbullying." IJCDS, 13(1).</p> <p>&nbsp;</p> <div> <p>* Due to the nature of this Dataset, comments contain offensiveness and hate speech. This does not reflect author values, however the aim is to providing a resource to help in detecting and preventing spread of such harmful content.</p> </div> <div> <h3>Features</h3> <ul> <li>Algerian Dialect</li> <li>Cyberbullying</li> <li>Hate speech</li> <li>Offensive Language</li> <li>Dialect Dataset</li> </ul> </div>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Sentiment dataset of Algerian dialect

<p>* This sentiment dataset of Algerian dialect consists of 11760 comments (6111 positive/ 5649 negative comments)) collected from (Facebook, YouTube and Twitter) during Hirak 2019.<br>* Comments concern the Algerian spoken language, written in Arabic and/or Latin characters and/or Arabizi, which could be either Modern Standard Arabic, French or local dialect.<br>* Value &lsquo;1&rsquo; is attributed for Positive review / value &lsquo;0&rsquo; attributed for Negative review.<br>* Due to the nature of this Dataset, some comments contain offensive language. This does not reflect author values, however the aim is to providing a resource to help in analysing positive and negative sentiments (that probably containing harmful content).<br>* For more information please contact (@Ahmed Cherif Mazari) :&nbsp;<a href="mailto:mazari.ac@gmail.com" target="_blank" rel="noopener">mazari.ac@gmail.com</a></p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Dvoice : An open source dataset for Automatic Speech Recognition on Moroccan dialectal Arabic

<p>Dialectal Voice is a community project initiated by AIOX Labs to facilitate voice recognition by Intelligent Systems. Today, the need for AI systems capable of recognizing the human voice is increasingly expressed within communities. However, we note that for some languages such as Darija, there are not enough voice technology solutions. To meet this need, we then proposed to establish this program of iterative and interactive construction of a dialectal database open to all in order to help improve models of voice recognition and generation.</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Fig. 3 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations

Fig. 3. Dendrogram of territorial complexes Lower-Dnipro Area (localities are described in table 1).

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 1 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations

Fig. 1. Components of Chaffinch song structure: 1 — phrase; 2 — inserted element; 3 — pre-flourish; 4 — flourish. a — element; b — sub-element.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 2 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations

Fig. 2. Study area and Lower-Dnipro dialect territory:&gt; 50 — localities with more than 50 % specific dialectal Lower Dnipro song types in song complex; 26–50 — specific dialectal types constitute 26–50 % of song complex; 5–25 — dialectal types constitute 5–25 %; 0 — specific dialectal types were not found.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Fig. 5 in Archaic Dialect Of Chaffinch, Fringilla Coelebs (Passeriformes, Fringillidae), Song In The Lower-Dnipro Area (South Ukraine) And Its Territorial Relations

Fig. 5. Specific "southern" elements in Lower-Dnipro dialect. Elements found in song complexes of South Ukraine: LD — Lower-Dnipro dialect; SE — South-East sub-dialect of Left-bank dialect; Cr — Crimean dialect; Da — Danube dialect; Cp — Carpathian dialect.

opencc-by-4.0Dec 2021View details →
zenodo40/100

Estonian dialect models

<p>Dialectal data and normalization models presented in the following paper:</p> <p>H&auml;m&auml;l&auml;inen, M., Alnajjar, K., &amp; Tuisk, T. (2022)&nbsp;<em>Help from the Neighbors: Estonian Dialect Normalization Using a Finnish<br> Dialect Generator</em>. In <em>The Proceedings of The Third Workshop for Deep Learning for Low Resource NLP.</em></p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

New version of the digitized Dialect Atlas of Finnish by Lauri Kettunen

<p>This package includes alternative representations of the digitized version of Lauri Kettunen's Dialect Atlas of Finnish (Kettunen 1940; <a href="http://urn.fi/urn:nbn:fi:csc-kata20151130145346403821" rel="nofollow">http://urn.fi/urn:nbn:fi:csc-kata20151130145346403821</a>). The first version of the digitized data was prepared by profs. Sheila Embleton and Eric Wheeler (York University, Canada) (Embleton &amp; Wheeler 1997, 2000), and further refined for publication by Jyri Lehtinen under the BEDLAN project (Biological Evolution and the Diversification of Languages).</p> <p>The alternative representation provided here includes the data formatted as (1) multistate character format and (2) binary character format. Both formats resemble the representation of population genetic data, and are thus easier to apply to population genetic analyses than the representation provided through the Fairdata repository. Also the linguistic classification of the collected linguistic traits is provided.</p> <p>The data has been prepared by the BEDLAN research project (Biological Evolution and the Diversification of Languages), which has conducted research of Kettunen's dialect atlas using population genetic techniques (Syrj&auml;nen et al. 2016, Honkola et al. 2018, Santaharju et al. in revision).</p> <p>Citation: Santaharju, Jenni, Kaj Syrj&auml;nen, Terhi Honkola, Perttu, Sepp&auml;, Outi Vesakoski and Unni Leino (submitted revision): New version of the digitized Dialect Atlas of Finnish by Lauri Kettunen. In Digital Humanities in the Nordic and Baltic Countries Publications. Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.10078078" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10078078</a></p> <div> <h2>kettunen_multistate.csv</h2> </div> <p>This file contains the multistate representation of the data. Here different states are coded as integers representing the different map symbols from Kettunen's original dialect atlas (for an example, see <a href="http://kettunen.fnhost.org/" rel="nofollow">http://kettunen.fnhost.org/</a>, which provides scans of Kettunen's original atlas map pages by Juha Kuokkala).</p> <ol> <li> <p><code>Municipality_number</code></p> <p>Unique identifier for each data point (municipality).</p> </li> <li> <p><code>Municipality_name</code></p> <p>Municipality name. These may not be unique, as several municipalities may have the same name.</p> </li> <li> <p><code>lon_WGS84</code></p> <p>WGS 84 longitude of the municipality centroid.</p> </li> <li> <p><code>lat_WGS84</code></p> <p>WGS 84 latitude of the municipality centroid.</p> </li> <li> <p><code>1a - 213c</code></p> <p>Dialect features present on each map page. The integer at the beginning matches the map page in Kettunen's atlas, so e.g. page 1 starts with 1. The letter following the integer represents different overlapping variants within a single map page. Each page number is suffixed with as many different letters as there are overlapping variants on that page; for instance page 1 has at most 3 overlapping variants for any datapoint, so the table includes three columns (1a-1c). The values are either integers (representing different map symbols from Kettunen's map page legends, with 1 being the topmost variant, 2 being the second variant, and so on), "-" (for absent data) or "NA" (for missing data).</p> </li> </ol> <div> <h2>kettunen_binary.csv</h2> </div> <ol> <li> <p><code>Municipality_number</code></p> <p>Unique identifier for each data point (municipality).</p> </li> <li> <p><code>Municipality_name</code></p> <p>Municipality name. These may not be unique, as several municipalities may have the same name.</p> </li> <li> <p><code>lon_WGS84</code></p> <p>WGS 84 longitude of the municipality centroid.</p> </li> <li> <p><code>lat_WGS84</code></p> <p>WGS 84 latitude of the municipality centroid.</p> </li> <li> <p><code>1_2 - 213_16</code></p> <p>Dialect features present on each map page. The leftmost integer matches the map page in Kettunen's atlas, so e.g. page 1 starts with 1. The second integer matches the multistate characters found in the cells of kettunen_multistate.csv, and thus reflect the different symbols in Kettunen's map page legends. Again here, 1 represents the topmost box in Kettunen's map page's legend, 2 the second box, and so on. The actual data field contains either 0 (absent), 1 (present) or "NA" (missing).</p> </li> </ol> <div> <h2>kettunen_map_explanations.csv</h2> </div> <ol> <li> <p><code>Map number</code></p> <p>Unique identifier for each map.</p> </li> <li> <p><code>The explanation of the map</code></p> <p>Map name and description of the dialect feature.</p> </li> <li> <p><code>Level1 - Level5</code></p> <p>Hierarchical classification of the dialect feature.</p> </li> </ol> <div> <h2>References</h2> </div> <p>Embleton, Sheila M. and Eric Wheeler, S. 1997. Finnish dialect atlas for quantitative studies. <em>Journal of Quantitative Linguistics</em> 4.99-102. <a href="https://doi.org/10.1080/0929617970859008" rel="nofollow">DOI: https://doi.org/10.1080/0929617970859008</a></p> <p>Embleton, Sheila, M. and Eric Wheeler, S. 2000. Computerized dialect atlas of Finnish: Dealing with ambiguity. <em>Journal of Quantitative Linguistics</em> 7.227-31. <a href="https://doi.org/10.1076/jqul.7.3.227.4109" rel="nofollow">DOI: https://doi.org/10.1076/jqul.7.3.227.4109</a></p> <p>Honkola, Terhi, Kalle Ruokolainen, Kaj Syrj&auml;nen, Unni-P&auml;iv&auml; Leino, Ilpo Tammi, Niklas Wahlberg, and Outi Vesakoski. 2018. Evolution within a Language: Environmental differences contribute to divergence of dialect groups. <em>BMC Evolutionary Biology</em> 18, no. 132. <a href="https://doi.org/10.1186/s12862-018-1238-6" rel="nofollow">DOI: https://doi.org/10.1186/s12862-018-1238-6</a></p> <p>Kettunen, Lauri. 1940. <em>Suomen Murteet III A. Murrekartasto.</em> Helsinki: Suomalaisen kirjallisuuden seura.</p> <p>Santaharju, Jenni, Terhi Honkola, Perttu, Sepp&auml;, Kaj Syrj&auml;nen, Unni Leino and Outi Vesakoski (in revision): Linguistic convergence and its drivers in Finnish dialects.</p> <p>Syrj&auml;nen, Kaj, Terhi Honkola, Jyri Lehtinen, Antti Leino and Outi Vesakoski. 2016. Applying population genetic approaches within languages: Finnish dialects as linguistic populations. <em>Language Dynamics and Change</em> 6.235-83. <a href="https://doi.org/10.1163/22105832-00602002" rel="nofollow">DOI: https://doi.org/10.1163/22105832-00602002</a></p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

A Monologue Narrative Text of the Itoman Dialect of Okinawan: How to Worship in Yukkanuhii

<p>This paper provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. He describes the way worship is conducted in Yukkanuhi. Japanese translations, morphological analyses, and interlinear glosses are provided. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Data supplement to: "Refractory depression - Mechanisms & Efficacy of Radically Open Dialectical Behaviour Therapy (RefraMED): findings of randomised trial on benefits and harms"

<p>Data and code to support primary analyses reported in&nbsp;&quot;Refractory depression &ndash; mechanisms and efficacy of radically open dialectical behaviour therapy (RefraMED): findings of a randomised trial on benefits and harms&quot;. British Journal of Psychiatry.&nbsp;doi: 10.1192/bjp.2019.53</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Data set on linguistic similarity of German dialects

<p>- The data&nbsp;provide information on pairwise dialect similarities for 439 NUTS 3 regions in Germany</p> <p><em>- </em>Data come from the maps and questionnaires of the&nbsp;Sprachatlas des Deutschen Reichs, digitized using ArcGIS software</p> <p>-&nbsp;The data source is a questionnaire with translations of standardized German sentences into local dialects between 1879 and 1888</p> <p>-&nbsp;The measure is defined as the number of co-occurrences for all pairs of sites (z-scaled)</p>

opencc-by-4.0Apr 2019View details →
zenodo40/100

A Chin dialect survey (Part 1 of 2)

<p>This dataset was compiled from 2005-2014 and includes a survey of Chin dialects, primarily southern Chin, spoken in Burma. The aim of the survey was to investigate mutual intelligibility and the diversity of the Chin family.</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

A Chin dialect survey (Part 2 of 2)

<p>This dataset was compiled from 2005-2014 and includes a survey of Chin dialects, primarily southern Chin, spoken in Burma. The aim of the survey was to investigate mutual intelligibility and the diversity of the Chin family.</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Yucatec Maya Dialect Atlas

<p>The &quot;Yucatec Maya Dialect Atlas&quot; is a dataset created at the Universidad Aut&oacute;noma de Yucat&aacute;n, M&eacute;rida, M&eacute;xico. The data was collected between 2000 and 2007; funded by the CONACyT (funding ID: 36387-H). The data was collected in 80 different locations in the peninsula of Yucat&aacute;n (n speakers = 157). Elicitation was based on a questionnaire with Spanish prompts (n = 665).</p> <p>&nbsp;</p>

opencc-by-nc-4.0Jul 2021View details →
zenodo40/100

Sudanese dialect speech dataset

<p>This is speech dataset for the Sudanese dialect data been collected from YouTube videos represent the characteristics of the Sudanese dialect, mainly the middle of Sudan dialect -Khartoum in particular- and have some northern tendency, primarily two programs Hajj Muzakir and Dukkan Wad Elbaseer.</p> <p>Transcription is done manually by listening to the audio files repeatedly to write the captions for the collected conversations to make sure that every word is written as said by the speakers. Transcription is written without diacritics on the Arabic alphabet, in a manner that reflects the Sudanese way of speaking, therefore, any correction to the noticeable mistakes was not applied to get rid of any biases and make the data representative.</p> <p>The &#39;Dataset&#39; subdirectory contains all the audio and text files for the corpus, the files organized based on program name &#39;hm_&#39; for Hajj Muzakir program and &#39;wb_&#39; Dukkan Wad Elbaseer, each filename follows three categories first two litters for the program name &#39;hm&#39; or &#39;wb&#39;, second the number of the episode third the number of the clip, hm_01_0001.wav and wb_01_0001.wav represent first episode of each program and the first clip.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

Data and Code Supplement to: "Processes of change in a randomized clinical trial of Radically Open Dialectical Behavior Therapy (RO DBT) for adults with treatment refractory depression"

<p>Dataset to support secondary&nbsp;analyses reported in&nbsp;&quot;Processes of change in a randomized clinical trial of Radically Open Dialectical&nbsp;Behavior Therapy (RO DBT) for adults with treatment refractory depression&quot; in the&nbsp;Journal of Consulting and Clinical Psychology</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

*-koatē/-škoatē inchoative verbs in Skolt and Kola Saami dialects

<p>The file contains two sets of data concerning *<em>-koatē/-&scaron;koatē</em> inchoative verbs in Skolt and Kola Saami dialects. The first one is a schematic representation of the verbs found in KKLS [1] (alphabetical span of A&ndash;K) with a focus on the derivational variants *<em>-koatē</em> vs. *<em>-&scaron;koatē</em>. Two versions are presented on separate sheets, one with a single alphabetical listing and another with grouping according to the (original) syllable count of the verb base. The columns P...V code the variants occurring in different dialects (codes: k = *-koatē, s = *-&scaron;koatē, ks = both).</p> <p>[1] KKLS = T. I. Itkonen 1958: <em>Koltan- ja kuolanlapin sanakirja. W&ouml;rterbuch des Kolta- und Kolalappischen</em>. I&ndash;II. Helsinki: Soci&eacute;t&eacute; Finno-Ougrienne.</p> <p>The second data set contains the *<em>-koatē/-&scaron;koatē</em> verb lexemes found in Szab&oacute; 1968 and 1987 [2, 3]. The data is grouped according to (original) syllable count, suffix variant (*<em>-koatē</em> vs. *<em>-&scaron;koatē</em>) and language area (Kildin vs. Ter [Turja]). Only one occurrence of each derivative lexeme is quoted.</p> <p>[2] Szab&oacute;, L&aacute;szl&oacute; 1968: <em>Kolalappische Volksdichtung (Texte aus den Dialekten in Kildin und Ter)</em>. Zweiter Teil nebst grammatischen Aufzeichnungen. G&ouml;ttingen: Vandenhoeck &amp; Ruprecht.</p> <p>[3] Szab&oacute;, L&aacute;szl&oacute; 1987: The use of the inchoative in Kola-Sami sentences. &ndash; <em>Nordlyd </em>13: 70&ndash;103.</p> <p>This is a data set supplement of: Kuokkala, J. (2019). Saami -<em>(&scaron;)goahti</em> inchoatives: their variation, history, and suggested cognates in Veps and Mordvin. <em>Suomalais-Ugrilaisen Seuran Aikakauskirja</em> 97, 153&ndash;181. <a href="https://doi.org/10.33340/susa.75718">https://doi.org/10.33340/susa.75718</a></p>

opencc-by-4.0Jun 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record