Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
166
datasets available to search
ShareScore release 0.9.0
Dataset results
166 results for “Dialect”
Murreviikko: an Annotated and Normalized Corpus of Dialectal Finnish Tweets
<p>Murreviikko (literally 'Dialect week') is a campaign founded in the University of Eastern Finland to promote the use of Finnish dialects in social media. It started in 2020 and takes place mid-October.</p> <p>The original data was collected from Twitter with the search word murreviikko ('dialect week') and hashtag #murreviikko separately for 2020, 2021 and 2022. The current dataset combines all the original collections.</p> <p>The tweets are dialectologically annotated on two levels: following the East-West division of Finnish dialects, and following a seven-way division of Finnish dialects (South-West, Häme, Southern Ostrobothnia, Central and Northern Ostrobothnia, Far North, Savo, and South-East), appended with the Helsinki slang. There is also a class for dialectal tweets, which are not discernible (NA) because of contrasting or scarce dialectal features.</p> <p>The original tweets are normalized to a phonetic standard, but word order is not altered, or grammar rules of standard Finnish followed otherwise. This means that for instance standard Finnish possessive suffixes (minun kirja-ni 'my book-my') are not added if they are not present in the original tweet (minun kirja). Likewise, dialect words are not corrected to the standard alternative, even if such words would exist (pruukata > pruukata instead of standard tavata).</p> <p>Following the rules of the Twitter API, this repository only includes the tweet id's, dialect annotations and normalizations. The original tweets are available for scientific use by request, as granted by the European Union’s Digital Single Market directive (2019/790).</p>
A Monologue Narrative Text of the Itoman Dialect of Okinawan: Yukkanuhii in My Childhood
<p>This dataset provides a monologue narrative text of the Itoman dialect of the Okinawan language spoken by a male speaker in his 70s. The speaker recounts his childhood memories of the event Yukkanuhii, which is held on the fourth day of the fifth lunar month. The event involves races in small boats called Haaree. The dataset includes an audio file (.wav) and an annotated xml file (.eaf). Japanese translations, morphological analyses, and interlinear glosses are provided in an .eaf file.</p>
Supplement for "Using Phylogenetic Networks to Model Chinese Dialect History"
<p>This is the supplementary material accompanying the paper "Using Phylogenetic Networks to Model Chinese Dialect History", which appeared in 2014 in "Language Dynamics and Change" (volume 4, issue 2).</p>
CLDF dataset derived from Deepadung et al.'s "Lexical Comparison of Palaung Dialects" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Deepadung, Sujaritlak; Buakaw, Supakit; and Rattanapitak, Ampica (2015): A lexical comparison of the Palaung dialects spoken in China, Myanmar, and Thailand. Mon-Khmer Studies 44. 19-38.</p> </blockquote>
Dataset for Number agreement and dependency length in Finnish dialects
<p>This material contains the dataset and the scripts from the <a href="https://version.helsinki.fi/gramadapt/depling2021-number-agreement">gitlab repository</a> of the following article. Please cite the article when using the data.</p> <p>Sinnemäki, Kaius & Akira Takaki 2021. Number agreement, dependency length, and word order in Finnish traditional dialects. In <em>Proceedings of the Sixth International Conference on Dependency Linguistics (Depling, SyntaxFest 2021)</em>. Stroudsburg, PA: The Association for Computational Linguistics.</p> <p> </p>
CLDF dataset derived from Castro's "Sui Dialect Research" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Castro, Andy and Pan, Xingwen (2015): Sui dialect research. SIL: Guiyang.</p> </blockquote>
CLDF dataset derived from Castro and Hansen's "Zhuang Dialects in Hongshui He" from 2010
<p>Cite the source of the dataset as:</p> <blockquote> <p>Castro, Andy; Hansen, Bruce (2010): Hongshui He Zhuang dialect intelligibility survey. Dallas: SIL International.</p> </blockquote>
CLDF dataset derived from Beijing University's "Chinese Dialect Vocabularies" from 1964
<p>Cite the source of the dataset as:</p> <blockquote> <p>Běijīng Dàxué 北京大学 (1964): Hànyǔ fāngyán cíhuì 汉语方言词汇 [Chinese dialect vocabularies]. Beijing: Wenzi Gaige.</p> </blockquote>
CLDF dataset derived from Allen's "Bai Dialect Survey" from 2007
<p>Cite the source of the dataset as:</p> <blockquote> <p>Allen, Bryan (2007): Bai Dialect Survey. Dallas: SIL International.</p> </blockquote>
CLDF dataset derived from Serva's "Dialects of Madagascar" from 2020
<p>Cite the source of the dataset as:</p> <blockquote> <p>Serva M., Pasquini M. (2020): Dialects of Madagascar, PLoS ONE 15(10).</p> </blockquote>
Pronunciation Dictionaries for the Alsatian Dialects
<p>This dataset contains a collection of pronunciation dictionaries which were manually transcribed using the X-SAMPA transcription system. The transcriptions were performed based on audio recordings available on the following websites :</p> <ul> <li>OLCA: <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Elsässich Web diktionnair: <a href="http://www.ami-hebdo.com/elsadico/index.php">http://www.ami-hebdo.com/elsadico/index.php</a></li> </ul> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The transcription process is described in the following research report : 10.5281/zenodo.1174219. It is also detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704814">http://hal.archives-ouvertes.fr/hal-01704814</a></p> <p>Three pronunciation dictionaries are available :</p> <ul> <li>elsassich_dico.csv : 702 transcriptions from the "Elsässich Web diktionnair"</li> <li>olca67.csv : 1,458 transcriptions from the "OLCA" lexicons for the northern part of the Alsace region (Bas-Rhin)</li> <li>olca68.csv : 1,401 transcriptions from the "OLCA" lexicons for the southern part of the Alsace region (Haut-Rhin)</li> </ul>
Lexicon of Place Names in the Alsatian Dialects
<p>This dataset contains a lexicon of place names in the Alsatian dialects. These place names were collected from several resources and manually categorised according to location types defined in the QUAERO project:</p> <ul> <li>loc.fac – Facility</li> <li>loc.phys.astro – Astronym</li> <li>loc.phys.geo – Geonym</li> <li>loc.phys.hydro – Hydronym</li> <li>loc.adm.nat – Country</li> <li>loc.adm.reg – Region</li> <li>loc.adm.sup – Supranational</li> <li>loc.adm.town – City</li> <li>loc.oro – Odonym</li> </ul> <p>The CSV file contains 4 columns:<br> 1. Place name in Alsatian<br> 2. Place name in French<br> 3. Quaero category<br> 4. Source(s): WikiAls (articles from the Alemannic Wikipedia), WikiFr (articles from the French Wikipedia), corpus (Wikipedia articles from the Alemannic Wikipedia and chronicles from an information magazine published by the Haut-Rhin department (southern Alsace) General Council). This field also indicates whether the same spelling can be found in other lexicons of place names for Alsatian: AlsaDico (Edmond Jung. <em>L’alsadico : 22 000 mots et expressions français-alsacien</em>. La<br> Nuée bleue, Strasbourg, 2006.) and Elsàsser (Marc Hug. <em>Toponymes d’Alsace</em>. Online, <a href="http://elsasser.free.fr/NomCommu/ecrantot.html">http://elsasser.free.fr/<br> NomCommu/ecrantot.html</a>, 2007.)</p> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The lexicon is also decribed in the following article: <a href="https://hal.archives-ouvertes.fr/hal-01702656">https://hal.archives-ouvertes.fr/hal-01702656</a>.</p>
CLDF Dataset derived from Hattori's "Japanese Dialects" from 1973
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hattori, S. (1973): Japanese dialects. In: Diachronic, areal and typological linguistics. Edited by H. M. Hoenigswald and R. H. Langacre. 368-400.</p> </blockquote>
CLDF dataset derived from Mitterhofer's "Dialect Survey of Bena" from 2013
<p>Cite the source of the dataset as:</p> <blockquote> <p>Mitterhofer, Bernadette. 2013. Lessons from a dialect survey of Bena: Analyzing wordlists. SIL International.</p> </blockquote>
CLDF dataset derived from Hóu's "Phonological Database of Chinese Dialects" from 2004
<p>Cite the source of the dataset as:</p> <blockquote> <p>Hóu, J. (2004): Xiàndài Hànyǔ fāngyán yīnkù 现代汉语方言音库 [Phonological database of Chinese dialects]. Shànghǎi: Shànghǎi Jiàoyù.</p> </blockquote>
CLDF dataset derived from Liú et al.'s "Collection of Basic Words in Chinese Dialects" from 2007
<p>Cite the source of the dataset as:</p> <blockquote> <p>Líu, L.; Wáng, H.; Bǎi, Y. (2007): Xiàndài Hànyǔ fāngyán héxīncí, tèzhēng cíjí 现代汉语方言核心词·特征词集 [Collection of basic vocabulary words and characteristic dialect words in modern Chinese dialects]. Nánjīng: Fènghuáng.</p> </blockquote>
Dataset for Variation in number agreement in Finnish dialects
<p>This material contains the dataset and the scripts from the <a href="https://version.helsinki.fi/gramadapt/verbinlukukongruenssi2021">gitlab repository</a> of the following article. Please cite the article when using the data.</p> <p>Sinnemäki, Kaius & Viljami Haakana 2021. <a href="https://researchportal.helsinki.fi/files/164044959/SinnemakiHaakana_2021_Kongruenssi.pdf">Variationistinen korpustutkimus predikaatin differentiaalisesta lukukongruenssista ja substantiiviluokasta suomen murteissa</a>. In Leena Maria Heikkola, Geda Paulsen, Katarzyna Wojciechowicz & Jutta Rosenberg (eds.), <a href="https://www.doria.fi/handle/10024/181090"><em>Språkets funktion: Juhlakirja Urpo Nikanteen 60-vuotispäivän kunniaksi - Festskrift till Urpo Nikanne på 60-årsdagen - Festschrift for Urpo Nikanne in honor of his 60th birthday</em></a>, 83–112. Åbo: Åbo Akademi Förlag.</p>
CLDF dataset derived from Wang's "Basic Words in Chinese Dialects" from 2004
<p>Cite the source of the dataset as:</p> <blockquote> <p>Wang, F. 2004. BCD: basic words of Chinese dialects. Unpublished dataset. [Digital version in: List, J.-M. (2015): Network perspectives on Chinese dialect history. Bulletin of Chinese Linguistics 8. 42-67.]</p> </blockquote>
Tigrinya Dialect Identification (TDI)
<p><strong>The Tigrinya Dialect Identification (TDI)</strong> dataset contains text on three Tigrinya dialects or varieties namely: Z, D, and L. The purpose of this dataset is to study dialect identification for Tigrinya using machine learning.</p> <p>For the Z variety, we used snippets from the book ኽልተ ዛንታት (Kilte Zantatat). For the L variety, we used book chapters from ፋቶ (Fato) and ዕርቂ እንደርታ (Erqi Enderta). For the D variant, we could not find a book. Instead, we collected data from two Facebook users, Akeza Awalom and Guraya Asadi Raya that consistently write in that variety. Sentences collected for each dialect were translated to the other dialect with expert native speakers in the target dialect.</p> <p><br> <strong>Source by Dialect</strong></p> <table align="left"> <tbody> <tr> <td> <p><strong>Dialect</strong></p> </td> <td> <p><strong>Source</strong></p> </td> <td> <p><strong>No. sentences</strong></p> </td> </tr> <tr> <td> <p>Z</p> </td> <td> <p>Kilte Zantatat</p> </td> <td> <p>1041</p> </td> </tr> <tr> <td> <p>L</p> </td> <td> <p>Fato </p> </td> <td> <p>764</p> </td> </tr> <tr> <td> </td> <td>Erqi nderta</td> <td>405</td> </tr> <tr> <td> <p>D</p> </td> <td> <p><a href="https://www.facebook.com/akeza.awealom">Akeza Awalom</a></p> </td> <td> <p>224</p> </td> </tr> <tr> <td> </td> <td><a href="https://www.facebook.com/guraya.asadiraya">GualRaya</a></td> <td>530</td> </tr> </tbody> </table> <p><br> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><strong>Acknowledgements</strong></p> <p> </p> <p>Special thanks to Meles Solomon, who provided us with his book, ዕርቂ እንደርታ (Erqi Enderta). He also helped with translations to the L dialect. Thanks also goes to Tesfay Gebreegziabher and Gidey Gebrekidan for allowing us to use their books Fato and Kilte Zantatat respectively. Many thanks to Teklay Berhane, Abeba Asemu, Haftu Abadi, Tsegay Kinfe, Moges Bekru, Kahsay Berhe Adhana, Kibrom Mulugeta, Tsegazeab Kidanu, Tsgab Weldemariam, Abu W Debay, Solomon Shibabaw, Hagos Hiete for their valuable contributions as translators.</p>
Supplementary Materials to "Subgrouping in a `dialect continuum': A Bayesian phylogenetic analysis of the Mixtecan language family"
<p>SM0: metadata on the languages of the sample</p> <p>SM 1: custom word list</p> <p>SM2: prose explanation of cognate coding and IPA conversion</p> <p>SM3: annotated cognate sets</p> <p>SM4: nexus files of the broad and fine grained cognate coding</p> <p>SM5: NeighborNet visualization with coloring by Josserand (1983)'s groupings and by groupings from our analysis</p> <p>SM6: BEAST2 xml files</p> <p>SM7: MCC trees from BEAST2 analysis</p> <p>SM8: DensiTree visualization and visualization of full MCC tree of best performing model</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.