Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8
datasets available to search
ShareScore release 0.9.0
Dataset results
8 results for “Uralic languages”
Collection of spatial information and maps of human past and environment in the Uralic languages speaker area
<p>The collection of spatial information and maps of the past and environment in the Uralic languages speaker area consists excessive amount of multidisciplinary data related to the vast region extending from Eastern Europe to Siberia, encompassing countries like Russia, Finland, and parts of Scandinavia. Uralic speakers are predominantly found in this region, with historical roots in areas around the Ural Mountains and adjacent territories. These datasets can be integrated for multidisciplinary purposes, allowing to explore human-environment interactions, migration patterns, and cultural evolution over time. Datasets are collected initially by the BEDLAN team <a href="https://bedlan.net/">https://bedlan.net/</a> - a research group specialized in various disciplines - linguists, archaeologists, geneticists, and geographers. The data collection and mapmaking have grown beyond the initial stages (publications, applications, exhibitions), hence collaborative effort for data publishing is now crucial. As the data collections and mapmaking continue to evolve dynamically together with ongoing projects, the current repository will be updated accordingly.</p>
Phlorest phylogeny derived from Honkola et al. 2013 'Cultural and climatic changes shape the evolutionary history of the Uralic languages'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Honkola T, Vesakoski O, Korhonen K, Lehtinen J, Syrjänen K & Wahlberg N. 2013. Cultural and climatic changes shape the evolutionary history of the Uralic languages. Journal of Evolutionary Biology, 26(6):1244–1253.</p> </blockquote>
Dataset of CONTAINMENT and SUPPORT in the Uralic languages of the Volga-Kama area
<p>This open access dataset contains examples of the expressions of CONTAINMENT and SUPPORT in the Uralic languages of Volga-Kama area. The exact languages and the sources of data are given in Table 1. The dataset contains data on the relational nouns (RN) and plain spatial cases expressing prototypical CONTAINMENT and SUPPORT in the languages. The RN included in the dataset are listed in Table 2, and case forms in Table 3.</p> <p> </p> <table> <tbody> <tr> <td> <p>language</p> </td> <td> <p>corpora</p> </td> </tr> <tr> <td> <p>Erzya (MdE)</p> </td> <td> <p>Syatko-subcorpus of the MokshEr corpus (MokshEr 2010)</p> </td> </tr> <tr> <td> <p>Moksha (MdM)</p> </td> <td> <p>Subcorpora in Moksha of the MokshEr corpus (MokshEr 2010)</p> </td> </tr> <tr> <td> <p>Meadow Mari (MaM)</p> </td> <td> <p>Marko East (Marko [no year]), Oncyko (Oncyko 2000), Meadow Mari corpus (Arkhangelskiy 2019b), Wanca (Meadow Mari) (Helsingin yliopisto et al. 2019)</p> </td> </tr> <tr> <td> <p>Hill Mari (MaH)</p> </td> <td> <p>Marko West (Marko [no year]), Wanca (Hill Mari) (Helsingin yliopisto et al. 2019)</p> </td> </tr> <tr> <td> <p>Udmurt (Udm)</p> </td> <td> <p>Pilot version of Udmurt corpus (relational nouns; presently included into [Arkhangelskiy 2018])</p> <p>Udmurt corpus (content nouns) (Arkhangelskiy 2018)</p> </td> </tr> <tr> <td> <p>Komi Zyrian (KoZ)</p> </td> <td> <p>Komi Zyrian Web Corpus (Arkhangelskiy 2019a), Коми корпус (Fu-Lab team 2021)</p> </td> </tr> <tr> <td> <p>Komi Permyak (KoP)</p> </td> <td> <p>Komi Permyak text collection from the University of Turku (Permyak 2009)</p> </td> </tr> </tbody> </table> <p><strong>Table 1.</strong> Languages included into the dataset and the sources of the data for each language.</p> <p> </p> <table> <tbody> <tr> <td> <p> </p> </td> <td> <p>MdE</p> </td> <td> <p>MdM</p> </td> <td> <p>MaM</p> </td> <td> <p>MaH</p> </td> <td> <p>Udm</p> </td> <td> <p>KoZ</p> </td> <td> <p>KoP</p> </td> </tr> <tr> <td> <p>containment</p> </td> <td> <p><em>pot</em>(<em>mo</em>)-</p> </td> <td> <p><em>potmə-</em></p> </td> <td> <p><em>kørgø</em>,<em> kørgə-</em></p> </td> <td> <p><em>kørgə̈-</em></p> </td> <td> <p><em>puʃk-</em></p> </td> <td> <p><em>pɨt͡ʃk-</em></p> </td> <td> <p><em>pɨt͡ʃk-</em></p> </td> </tr> <tr> <td> <p>support</p> </td> <td> <p><em>lang-</em></p> </td> <td> <p><em>lang-</em></p> </td> <td> <p><em>ymba-</em></p> </td> <td> <p><em>βə̈(l)-</em></p> </td> <td> <p><em>vɨl-</em></p> </td> <td> <p><em>vɨl-/vɨv-</em></p> </td> <td> <p><em>vɨl-/vɨv-</em></p> </td> </tr> </tbody> </table> <p><strong>Table 2.</strong> RN included in the dataset.</p> <p> </p> <table> <tbody> <tr> <td> <p> </p> </td> <td> <p>MdE</p> </td> <td> <p>MdM</p> </td> <td> <p>MaM</p> </td> <td> <p>MaH</p> </td> <td> <p>Udm</p> </td> <td> <p>KoZ</p> </td> <td> <p>KoP</p> </td> </tr> <tr> <td> <p>location</p> </td> <td> <p><em>-so</em>/<em>-se</em> (inessive)</p> </td> <td> <p><em>-sa</em> (inessive)</p> </td> <td> <p><em>-ʃte/-ʃto/-ʃtø</em> (inessive)</p> </td> <td> <p><em>-ʃtə/-ʃtə̈</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> </tr> <tr> <td> <p>source</p> </td> <td> <p><em>-sto</em>/<em>-ste</em> (elative)</p> </td> <td> <p><em>-sta</em> (elative)</p> </td> <td> <p><em>gət͡ɕ</em> (source postposition)</p> </td> <td> <p><em>gə̈t͡s</em> (source postposition)</p> </td> <td> <p><em>-ɨɕ</em> (elative)</p> </td> <td> <p><em>-ɨɕ</em> (elative)</p> </td> <td> <p><em>-iɕ</em> (elative)</p> </td> </tr> <tr> <td> <p>goal</p> </td> <td> <p><em>-s</em> (illative)</p> </td> <td> <p><em>-s</em>/<em>-t͜s</em> (illative)</p> </td> <td> <p><em>-ʃke/-ʃko/-ʃkø/-ʃ</em>, (illative)</p> </td> <td> <p><em>-ʃkə/-ʃkə̈/-ʃ</em>, (illative)</p> </td> <td> <p><em>-e</em>/<em>-ɨ </em>(illative)</p> </td> <td> <p><em>-ɘ </em>(illative)</p> </td> <td> <p><em>-ɘ </em>(illative)</p> </td> </tr> <tr> <td> <p>path</p> </td> <td> <p><em>-ka</em>/<em>-ga</em>/<em>-va</em> (prolative)</p> </td> <td> <p><em>-ka</em>/<em>-ga</em>/<em>-va</em>/<em>-gæ</em> (prolative)</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> <td> <p><em>-ti/-eti</em>/<em>-jeti/</em><em>-ɨti</em> (prolative)</p> </td> <td> <p><em>-ɘd</em> (prolative); <em>-ti</em> (transitive)</p> </td> <td> <p>-<em>ɘt </em>(prolative); <em>-ti</em> (transitive)</p> </td> </tr> </tbody> </table> <p><strong>Table 3.</strong> Cases that have been included into the dataset. All cases do not necessary show in every set, as for some combinations of RN and case there is no data.</p> <p> </p> <p>The main purpose of the dataset is to enable the study of variation between a plain case and RN inflected in case when expressing CONTAINMENT or SUPPORT. To facilitate this each expression of relation has been given a prototypicality score 4 = most prototypical, 1 = non-prototypical, which tells if the relation between landmark and trajector expressed in the sentence is typical for the entities participating in it. The prototypicality scores are based on the pre-linguistic concepts of containment and support, which are robustly attested and therefore should be independent of any single language. The scoring is based on the authors understanding of the language external relations, and no native consultants are used to verify the results. Therefore, some caution is in order when using the dataset.</p> <p> </p> <p>The dataset contains files with data of CONTAINMENT RN, SUPPORT RN, and plain case on all the included languages. The files are named according to the scheme element_languge (e. g. Containment_Erzya for the containment data on Erzya). In addition, files named element_frequencies show the number of examples divided by case and prototypicality score for each language, and element_summary shows the total number of prototypicality scores for each language. For plain case there are also summary files for the scores of CONTAINMENT and SUPPORT separately.</p> <p> </p> <p>The dataset is annotated for following information:</p> <ol> <li>The case in which the content noun or RN is inflected.</li> <li>The predicate as inflected in the data.</li> <li>The content noun as given in the data.</li> <li>Translations of both (mainly in citation form, but in predicate sometimes with some grammatical information, cf. abbreviations below).</li> <li>The prototypicality score. In the data on plain cases the prototypicality score is given only for the clauses where the relation is either CONTAINMENT or SUPPORT (i. e. the prototypicality score indicates the prototypicality of the relation as CONTAINMENT or SUPPORT according to the type of relation expressed).</li> <li>In the data on plain cases, the relation expressed by the case is marked (CONT = CONTAINMENT, SUP = SUPPORT, N/A = some other relation).</li> <li>The original sentence context.</li> <li>Free translation. Some of the translations are done following the lexical meanings and syntactic structures of the languages, so the English is unidiomatic from time to time.</li> <li>The file name with which the original sentence can be located in the corpus.</li> </ol> <p> </p> <p>The three final columns are partly lacking at the moment from the Mari and Komi languages. The translations in the data are intended only as guidelines, and anyone using the dataset should refer to the original language data in the analysis. The data in the columns is presented according to the following conventions :</p> <ul> <li>If the predicate is in square brackets, it means that the predicate is not present in the clause with the target LM. This can be because of two reasons: 1) The predicate is given in a previous clause, and is elliptically omitted, 2) the “predicate” is copula, which is not obligatory in the present tense in the languages studied.</li> <li>The following abbreviations are used to specify the meaning of the predicate when the English translation is ambiguous (note that the use is not checked, and the abbreviations might be lacking from some predicates):</li> </ul> <p>CAUS causative</p> <p>CONT continuative</p> <p>CVB converb</p> <p>FRQ frequentative</p> <p>INCH inchoative</p> <p>INF infinitive</p> <p>ITR intransitive</p> <p>MOM momentaneous</p> <p>NEG negative</p> <p>NMLZ nominalization</p> <p>PASS passive</p> <p>PTCP participle</p> <p>REFL reflexive</p> <p>TRA transitive</p> <p>The authors of this dataset are Tomi Koivunen and Riku Erkkilä and it is published under CC-BY-NC-ND licence. If used in a publication, please refer to this publication as well as mention the original source(s):</p> <p>This dataset has been used in following publications:</p> <p> </p> <p>References to used corpora:</p> <p>Arkhangelskiy, Timofey. 2018. <em>Udmurt corpus</em>. http://udmurt.web-corpora.net/index.html.</p> <p>Arkhangelskiy, Timofey. 2019a. <em>Komi-Zyrian corpus</em>. http://komi-zyrian.web-corpora.net/index.html.</p> <p>Arkhangelskiy, Timofey. 2019b. <em>Meadow Mari corpus</em>. http://meadow-mari.web-corpora.net/index_en.html.</p> <p>Fu-Lab team. 2021. <em>Корпус коми языка</em>. http://komicorpora.ru/.</p> <p>Helsingin yliopisto, FIN-CLARIN, H. Jauhiainen, T. Jauhiainen & K. Lindén. 2019. <em>Wanca 2016, Korp Version</em>. Kielipankki. http://urn.fi/urn:nbn:fi:lb-2019052401.</p> <p>Marko. (no year). <em>MARKO - Corpus of Mari language</em>. University of Turku.</p> <p>MokshEr, V.3. 2010. <em>Mokšan ja ersän sähköinen korpus</em>. Turun yliopisto.</p> <p>Oncyko. 2000. <em>Oncyko corpus</em>. University of Turku.</p> <p>Permyak. 2009. <em>Turku Komi-Permyak Corpus</em>. University of Turku.</p>
SemUr - Semantic Databases for Uralic Languages
<p>These databases are translated from <a href="http://mikakalevi.com/semfi">SemFi</a> by using Giellatekno XML dictionaries. The included python script can be used to update these databases or to create new ones for other languages.</p> <p>Currently, SemUr has the following languages</p> <ul> <li>SemSms - Skolt Sami</li> <li>SemKpv - Komi Zyrian</li> <li>SemMyv - Erzya</li> <li>SemMdf - Moksha</li> </ul> <p> </p> <p><strong>Cite as</strong></p> <p>Hämäläinen, Mika. (2018). <a href="https://helda.helsinki.fi//bitstream/handle/10138/282733/paper9.pdf?sequence=1">Extracting a Semantic Database with Syntactic Relations for Finnish to Boost Resources for Endangered Uralic Languages</a>. In The Proceedings of Logic and Engineering of Natural Language Semantics 15 (LENLS15)</p> <p> </p>
Kinura: A database of kinship terminology from the Uralic Language family
<p>The data repository for the Kinura dataset </p>
Kinura: A database of kinship terminology from the Uralic Language family
<p><strong>The data repository for the Kinura dataset.</strong></p>
Online dictionary backups of Uralic languages
<p>https://sanat.csc.fi/</p>
Data from: Cultural and climatic changes shape the evolutionary history of the Uralic languages
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.