Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

166

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

166 results for “Dialect”

Learn how ShareScore rates datasets ↗
zenodo52/100

Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets

<p>Murreviikko (literally &#39;Dialect week&#39;) is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko (&#39;dialect week&#39;) and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, H&auml;me, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni &#39;my book-my&#39;) are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata &gt; pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id&#39;s, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union&rsquo;s Digital Single Market directive (2019/790).</p>

opencc-by-4.0May 2023View details →
zenodo48/100

A Monologue Narrative Text of the Itoman Dialect of Okinawan: Yukkanuhii in My Childhood

<p>This dataset provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. The speaker recounts his childhood memories of the event Yukkanuhii, which is held on the fourth day of the fifth lunar month. The event involves races in small boats called Haaree. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Supplement for "Using Phylogenetic Networks to Model Chinese Dialect History"

<p>This is the supplementary material accompanying the paper &quot;Using Phylogenetic Networks to Model Chinese Dialect History&quot;, which appeared in 2014 in &quot;Language Dynamics and Change&quot; (volume 4, issue 2).</p>

opencc-zeroAug 2014View details →
zenodo44/100

CLDF dataset derived from Deepadung et al.'s "Lexical Comparison of Palaung Dialects" from 2015

<p>Cite the source of the dataset as:</p> <blockquote> <p>Deepadung, Sujaritlak; Buakaw, Supakit; and Rattanapitak, Ampica (2015): A lexical comparison of the Palaung dialects spoken in China, Myanmar, and Thailand. Mon-Khmer Studies 44. 19-38.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Dataset for Number agreement and dependency length in Finnish dialects

<p>This material contains the dataset and the scripts from the <a href="https://version.helsinki.fi/gramadapt/depling2021-number-agreement">gitlab repository</a>&nbsp;of the following article. Please cite the article when using the data.</p> <p>Sinnem&auml;ki, Kaius &amp; Akira Takaki 2021. Number agreement, dependency length, and word order in Finnish traditional dialects. In <em>Proceedings of the Sixth International Conference on Dependency Linguistics (Depling, SyntaxFest 2021)</em>. Stroudsburg, PA: The Association for Computational Linguistics.</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

CLDF dataset derived from Castro's "Sui Dialect Research" from 2015

<p>Cite the source of the dataset as:</p> <blockquote> <p>Castro, Andy and Pan, Xingwen (2015): Sui dialect research. SIL: Guiyang.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Castro and Hansen's "Zhuang Dialects in Hongshui He" from 2010

<p>Cite the source of the dataset as:</p> <blockquote> <p>Castro, Andy; Hansen, Bruce (2010): Hongshui He Zhuang dialect intelligibility survey. Dallas: SIL International.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Beijing University's "Chinese Dialect Vocabularies" from 1964

<p>Cite the source of the dataset as:</p> <blockquote> <p>Běijīng Dàxué 北京大学 (1964): Hànyǔ fāngyán cíhuì 汉语方言词汇 [Chinese dialect vocabularies]. Beijing: Wenzi Gaige.</p> </blockquote>

opencc-by-4.0Nov 2019View details →
zenodo44/100

CLDF dataset derived from Allen's "Bai Dialect Survey" from 2007

<p>Cite the source of the dataset as:</p> <blockquote> <p>Allen, Bryan (2007): Bai Dialect Survey. Dallas: SIL International.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Serva's "Dialects of Madagascar" from 2020

<p>Cite the source of the dataset as:</p> <blockquote> <p>Serva M., Pasquini M. (2020): Dialects of Madagascar, PLoS ONE 15(10).</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Pronunciation Dictionaries for the Alsatian Dialects

<p>This dataset contains a collection of pronunciation dictionaries which were manually transcribed using the X-SAMPA transcription system. The transcriptions were performed based on audio recordings available on the following websites :</p> <ul> <li>OLCA: <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Els&auml;ssich Web diktionnair: <a href="http://www.ami-hebdo.com/elsadico/index.php">http://www.ami-hebdo.com/elsadico/index.php</a></li> </ul> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The transcription process is described in the following research report : 10.5281/zenodo.1174219. It is also detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704814">http://hal.archives-ouvertes.fr/hal-01704814</a></p> <p>Three pronunciation dictionaries are available :</p> <ul> <li>elsassich_dico.csv : 702 transcriptions from the &quot;Els&auml;ssich Web diktionnair&quot;</li> <li>olca67.csv : 1,458 transcriptions from the &quot;OLCA&quot; lexicons for the northern part of the Alsace region (Bas-Rhin)</li> <li>olca68.csv : 1,401 transcriptions from the &quot;OLCA&quot; lexicons for the southern part of the Alsace region (Haut-Rhin)</li> </ul>

opencc-by-sa-4.0Feb 2018View details →
zenodo44/100

Lexicon of Place Names in the Alsatian Dialects

<p>This dataset contains a lexicon of place names in the Alsatian dialects. These place names were collected from several resources and manually categorised according to location types defined in the QUAERO project:</p> <ul> <li>loc.fac &ndash; Facility</li> <li>loc.phys.astro &ndash; Astronym</li> <li>loc.phys.geo &ndash; Geonym</li> <li>loc.phys.hydro &ndash; Hydronym</li> <li>loc.adm.nat &ndash; Country</li> <li>loc.adm.reg &ndash; Region</li> <li>loc.adm.sup &ndash; Supranational</li> <li>loc.adm.town &ndash; City</li> <li>loc.oro &ndash; Odonym</li> </ul> <p>The CSV file contains 4 columns:<br> 1. Place name in Alsatian<br> 2. Place name in French<br> 3. Quaero category<br> 4. Source(s): WikiAls (articles from the Alemannic Wikipedia), WikiFr (articles from the French Wikipedia), corpus (Wikipedia articles from the Alemannic Wikipedia and chronicles from an information magazine published by the Haut-Rhin department (southern Alsace) General Council). This field also indicates whether the same spelling can be found in other lexicons of place names for Alsatian: AlsaDico (Edmond Jung. <em>L&rsquo;alsadico : 22 000 mots et expressions fran&ccedil;ais-alsacien</em>. La<br> Nu&eacute;e bleue, Strasbourg, 2006.) and&nbsp;Els&agrave;sser (Marc Hug. <em>Toponymes d&rsquo;Alsace</em>. Online, <a href="http://elsasser.free.fr/NomCommu/ecrantot.html">http://elsasser.free.fr/<br> NomCommu/ecrantot.html</a>, 2007.)</p> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The lexicon is also decribed in the following article: <a href="https://hal.archives-ouvertes.fr/hal-01702656">https://hal.archives-ouvertes.fr/hal-01702656</a>.</p>

opencc-by-sa-4.0Aug 2018View details →
zenodo44/100

CLDF Dataset derived from Hattori's "Japanese Dialects" from 1973

<p>Cite the source of the dataset as:</p> <blockquote> <p>Hattori, S. (1973): Japanese dialects. In: Diachronic, areal and typological linguistics. Edited by H. M. Hoenigswald and R. H. Langacre. 368-400.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Mitterhofer's "Dialect Survey of Bena" from 2013

<p>Cite the source of the dataset as:</p> <blockquote> <p>Mitterhofer, Bernadette. 2013. Lessons from a dialect survey of Bena: Analyzing wordlists. SIL International.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Hóu's "Phonological Database of Chinese Dialects" from 2004

<p>Cite the source of the dataset as:</p> <blockquote> <p>Hóu, J. (2004): Xiàndài Hànyǔ fāngyán yīnkù 现代汉语方言音库 [Phonological database of Chinese dialects]. Shànghǎi: Shànghǎi Jiàoyù.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

CLDF dataset derived from Liú et al.'s "Collection of Basic Words in Chinese Dialects" from 2007

<p>Cite the source of the dataset as:</p> <blockquote> <p>Líu, L.; Wáng, H.; Bǎi, Y. (2007): Xiàndài Hànyǔ fāngyán héxīncí, tèzhēng cíjí 现代汉语方言核心词·特征词集 [Collection of basic vocabulary words and characteristic dialect words in modern Chinese dialects]. Nánjīng: Fènghuáng.</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Dataset for Variation in number agreement in Finnish dialects

<p>This material contains the dataset and the scripts from the <a href="https://version.helsinki.fi/gramadapt/verbinlukukongruenssi2021">gitlab repository</a> of the following article. Please cite the article when using the data.</p> <p>Sinnem&auml;ki, Kaius &amp; Viljami Haakana 2021. <a href="https://researchportal.helsinki.fi/files/164044959/SinnemakiHaakana_2021_Kongruenssi.pdf">Variationistinen korpustutkimus predikaatin differentiaalisesta lukukongruenssista ja substantiiviluokasta suomen murteissa</a>. In Leena Maria Heikkola, Geda Paulsen, Katarzyna Wojciechowicz &amp; Jutta Rosenberg (eds.), <a href="https://www.doria.fi/handle/10024/181090"><em>Spr&aring;kets funktion: Juhlakirja Urpo Nikanteen 60-vuotisp&auml;iv&auml;n kunniaksi - Festskrift till Urpo Nikanne p&aring; 60-&aring;rsdagen - Festschrift for Urpo Nikanne in honor of his 60th birthday</em></a>, 83&ndash;112. &Aring;bo: &Aring;bo Akademi F&ouml;rlag.</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

CLDF dataset derived from Wang's "Basic Words in Chinese Dialects" from 2004

<p>Cite the source of the dataset as:</p> <blockquote> <p>Wang, F. 2004. BCD: basic words of Chinese dialects. Unpublished dataset. [Digital version in: List, J.-M. (2015): Network perspectives on Chinese dialect history. Bulletin of Chinese Linguistics 8. 42-67.]</p> </blockquote>

opencc-by-4.0Jul 2021View details →
zenodo44/100

Tigrinya Dialect Identification (TDI)

<p><strong>The Tigrinya Dialect Identification (TDI)</strong> dataset contains text on three Tigrinya dialects or varieties namely: Z, D, and L. The purpose of this dataset is to study dialect identification for Tigrinya using machine learning.</p> <p>For the Z variety, we used snippets from the book ኽልተ ዛንታት (Kilte Zantatat). For the L variety, we used book chapters from ፋቶ (Fato) and ዕርቂ እንደርታ (Erqi Enderta). For the D variant, we could not find a book. Instead, we collected data from two Facebook users, Akeza Awalom and Guraya Asadi Raya that consistently write in that variety.&nbsp; Sentences collected for each dialect were translated to the other dialect with expert native speakers in the target dialect.</p> <p><br> <strong>Source&nbsp;by Dialect</strong></p> <table align="left"> <tbody> <tr> <td> <p><strong>Dialect</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>No. sentences</strong></p> </td> </tr> <tr> <td> <p>Z</p> </td> <td> <p>Kilte Zantatat</p> </td> <td> <p>1041</p> </td> </tr> <tr> <td> <p>L</p> </td> <td> <p>Fato&nbsp;</p> </td> <td> <p>764</p> </td> </tr> <tr> <td>&nbsp;</td> <td>Erqi nderta</td> <td>405</td> </tr> <tr> <td> <p>D</p> </td> <td> <p><a href="https://www.facebook.com/akeza.awealom">Akeza Awalom</a></p> </td> <td> <p>224</p> </td> </tr> <tr> <td>&nbsp;</td> <td><a href="https://www.facebook.com/guraya.asadiraya">GualRaya</a></td> <td>530</td> </tr> </tbody> </table> <p><br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgements</strong></p> <p>&nbsp;</p> <p>Special thanks to Meles Solomon, who provided us with his book, ዕርቂ እንደርታ (Erqi Enderta).&nbsp; He also helped with translations to the L dialect. Thanks also goes to Tesfay Gebreegziabher&nbsp; and Gidey Gebrekidan for allowing us to use their books Fato and Kilte Zantatat respectively. Many thanks to Teklay Berhane, Abeba Asemu, Haftu Abadi, Tsegay Kinfe, Moges Bekru, Kahsay Berhe Adhana, Kibrom Mulugeta, Tsegazeab Kidanu, Tsgab Weldemariam, Abu W Debay, Solomon Shibabaw, Hagos Hiete for their valuable contributions as translators.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Supplementary Materials to "Subgrouping in a `dialect continuum': A Bayesian phylogenetic analysis of the Mixtecan language family"

<p>SM0: metadata on the languages of the sample</p> <p>SM 1: custom word list</p> <p>SM2: prose explanation of cognate coding and IPA conversion</p> <p>SM3: annotated cognate sets</p> <p>SM4: nexus files of the broad and fine grained cognate coding</p> <p>SM5: NeighborNet visualization with coloring by Josserand (1983)&#39;s groupings and by groupings from our analysis</p> <p>SM6: BEAST2 xml files</p> <p>SM7: MCC trees from BEAST2 analysis</p> <p>SM8: DensiTree visualization and visualization of full MCC tree of best performing model</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record