Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1 result for “Volga-Kama”

Learn how ShareScore rates datasets ↗
zenodo44/100

Dataset of CONTAINMENT and SUPPORT in the Uralic languages of the Volga-Kama area

<p>This open access dataset contains examples of the expressions of CONTAINMENT and SUPPORT in the Uralic languages of Volga-Kama area. The exact languages and the sources of data are given in Table 1. The dataset contains data on the relational nouns (RN) and plain spatial cases expressing prototypical CONTAINMENT and SUPPORT in the languages. The RN included in the dataset are listed in Table 2, and case forms in Table 3.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p>language</p> </td> <td> <p>corpora</p> </td> </tr> <tr> <td> <p>Erzya (MdE)</p> </td> <td> <p>Syatko-subcorpus of the MokshEr corpus (MokshEr 2010)</p> </td> </tr> <tr> <td> <p>Moksha (MdM)</p> </td> <td> <p>Subcorpora in Moksha of the MokshEr corpus (MokshEr 2010)</p> </td> </tr> <tr> <td> <p>Meadow Mari (MaM)</p> </td> <td> <p>Marko East (Marko [no year]), Oncyko (Oncyko 2000), Meadow Mari corpus (Arkhangelskiy 2019b), Wanca (Meadow Mari) (Helsingin yliopisto et al. 2019)</p> </td> </tr> <tr> <td> <p>Hill Mari (MaH)</p> </td> <td> <p>Marko West (Marko [no year]), Wanca (Hill Mari) (Helsingin yliopisto et al. 2019)</p> </td> </tr> <tr> <td> <p>Udmurt (Udm)</p> </td> <td> <p>Pilot version of Udmurt corpus (relational nouns; presently included into [Arkhangelskiy 2018])</p> <p>Udmurt corpus (content nouns) (Arkhangelskiy 2018)</p> </td> </tr> <tr> <td> <p>Komi Zyrian (KoZ)</p> </td> <td> <p>Komi Zyrian Web Corpus (Arkhangelskiy 2019a), Коми корпус (Fu-Lab team 2021)</p> </td> </tr> <tr> <td> <p>Komi Permyak (KoP)</p> </td> <td> <p>Komi Permyak text collection from the University of Turku (Permyak 2009)</p> </td> </tr> </tbody> </table> <p><strong>Table 1.</strong> Languages included into the dataset and the sources of the data for each language.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p>&nbsp;</p> </td> <td> <p>MdE</p> </td> <td> <p>MdM</p> </td> <td> <p>MaM</p> </td> <td> <p>MaH</p> </td> <td> <p>Udm</p> </td> <td> <p>KoZ</p> </td> <td> <p>KoP</p> </td> </tr> <tr> <td> <p>containment</p> </td> <td> <p><em>pot</em>(<em>mo</em>)-</p> </td> <td> <p><em>potmə-</em></p> </td> <td> <p><em>k&oslash;rg&oslash;</em>,<em> k&oslash;rgə-</em></p> </td> <td> <p><em>k&oslash;rgə̈-</em></p> </td> <td> <p><em>puʃk-</em></p> </td> <td> <p><em>pɨt͡ʃk-</em></p> </td> <td> <p><em>pɨt͡ʃk-</em></p> </td> </tr> <tr> <td> <p>support</p> </td> <td> <p><em>lang-</em></p> </td> <td> <p><em>lang-</em></p> </td> <td> <p><em>ymba-</em></p> </td> <td> <p><em>&beta;ə̈(l)-</em></p> </td> <td> <p><em>vɨl-</em></p> </td> <td> <p><em>vɨl-/vɨv-</em></p> </td> <td> <p><em>vɨl-/vɨv-</em></p> </td> </tr> </tbody> </table> <p><strong>Table 2.</strong> RN included in the dataset.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td> <p>&nbsp;</p> </td> <td> <p>MdE</p> </td> <td> <p>MdM</p> </td> <td> <p>MaM</p> </td> <td> <p>MaH</p> </td> <td> <p>Udm</p> </td> <td> <p>KoZ</p> </td> <td> <p>KoP</p> </td> </tr> <tr> <td> <p>location</p> </td> <td> <p><em>-so</em>/<em>-se</em> (inessive)</p> </td> <td> <p><em>-sa</em> (inessive)</p> </td> <td> <p><em>-ʃte/-ʃto/-ʃt&oslash;</em> (inessive)</p> </td> <td> <p><em>-ʃtə/-ʃtə̈</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> <td> <p><em>-ɨn</em> (inessive)</p> </td> </tr> <tr> <td> <p>source</p> </td> <td> <p><em>-sto</em>/<em>-ste</em> (elative)</p> </td> <td> <p><em>-sta</em> (elative)</p> </td> <td> <p><em>gət͡ɕ</em> (source postposition)</p> </td> <td> <p><em>gə̈t͡s</em> (source postposition)</p> </td> <td> <p><em>-ɨɕ</em> (elative)</p> </td> <td> <p><em>-ɨɕ</em> (elative)</p> </td> <td> <p><em>-iɕ</em> (elative)</p> </td> </tr> <tr> <td> <p>goal</p> </td> <td> <p><em>-s</em> (illative)</p> </td> <td> <p><em>-s</em>/<em>-t͜s</em> (illative)</p> </td> <td> <p><em>-ʃke/-ʃko/-ʃk&oslash;/-ʃ</em>, (illative)</p> </td> <td> <p><em>-ʃkə/-ʃkə̈/-ʃ</em>, (illative)</p> </td> <td> <p><em>-e</em>/<em>-ɨ </em>(illative)</p> </td> <td> <p><em>-ɘ </em>(illative)</p> </td> <td> <p><em>-ɘ </em>(illative)</p> </td> </tr> <tr> <td> <p>path</p> </td> <td> <p><em>-ka</em>/<em>-ga</em>/<em>-va</em> (prolative)</p> </td> <td> <p><em>-ka</em>/<em>-ga</em>/<em>-va</em>/<em>-g&aelig;</em> (prolative)</p> </td> <td> <p>-</p> </td> <td> <p>-</p> </td> <td> <p><em>-ti/-eti</em>/<em>-jeti/</em><em>-ɨti</em> (prolative)</p> </td> <td> <p><em>-ɘd</em> (prolative); <em>-ti</em> (transitive)</p> </td> <td> <p>-<em>ɘt </em>(prolative); <em>-ti</em> (transitive)</p> </td> </tr> </tbody> </table> <p><strong>Table 3.</strong> Cases that have been included into the dataset. All cases do not necessary show in every set, as for some combinations of RN and case there is no data.</p> <p>&nbsp;</p> <p>The main purpose of the dataset is to enable the study of variation between a plain case and RN inflected in case when expressing CONTAINMENT or SUPPORT. To facilitate this each expression of relation has been given a prototypicality score 4 = most prototypical, 1 = non-prototypical, which tells if the relation between landmark and trajector expressed in the sentence is typical for the entities participating in it. The prototypicality scores are based on the pre-linguistic concepts of containment and support, which are robustly attested and therefore should be independent of any single language. The scoring is based on the authors understanding of the language external relations, and no native consultants are used to verify the results. Therefore, some caution is in order when using the dataset.</p> <p>&nbsp;</p> <p>The dataset contains files with data of CONTAINMENT&nbsp;RN, SUPPORT RN, and plain case on all the included languages. The files are named according to the scheme element_languge (e. g. Containment_Erzya for the containment data on Erzya). In addition, files named element_frequencies show the number of examples divided by case and prototypicality score for each language, and element_summary shows the total number of prototypicality scores for each language. For plain case there are also summary files for the scores of CONTAINMENT and SUPPORT&nbsp;separately.</p> <p>&nbsp;</p> <p>The dataset is annotated for following information:</p> <ol> <li>The case in which the content noun or RN is inflected.</li> <li>The predicate as inflected in the data.</li> <li>The content noun as given in the data.</li> <li>Translations of both (mainly in citation form, but in predicate sometimes with some grammatical information, cf. abbreviations below).</li> <li>The prototypicality score. In the data on plain cases the prototypicality score is given only for the clauses where the relation is either CONTAINMENT or SUPPORT (i. e. the prototypicality score indicates the prototypicality of the relation as CONTAINMENT or SUPPORT according to the type of relation expressed).</li> <li>In the data on plain cases, the relation expressed by the case is marked (CONT&nbsp;= CONTAINMENT, SUP&nbsp;= SUPPORT, N/A&nbsp;= some other relation).</li> <li>The original sentence context.</li> <li>Free translation. Some of the translations are done following the lexical meanings and syntactic structures of the languages, so the English is unidiomatic from time to time.</li> <li>The file name with which the original sentence can be located in the corpus.</li> </ol> <p>&nbsp;</p> <p>The three final columns are partly lacking at the moment from the Mari and Komi languages. The translations in the data are intended only as guidelines, and anyone using the dataset should refer to the original language data in the analysis. The data in the columns is presented according to the following conventions :</p> <ul> <li>If the predicate is in square brackets, it means that the predicate is not present in the clause with the target LM. This can be because of two reasons: 1) The predicate is given in a previous clause, and is elliptically omitted, 2) the &ldquo;predicate&rdquo; is copula, which is not obligatory in the present tense in the languages studied.</li> <li>The following abbreviations are used to specify the meaning of the predicate when the English translation is ambiguous (note that the use is not checked, and the abbreviations might be lacking from some predicates):</li> </ul> <p>CAUS&nbsp; &nbsp; causative</p> <p>CONT&nbsp; &nbsp;continuative</p> <p>CVB&nbsp; &nbsp; &nbsp; converb</p> <p>FRQ&nbsp; &nbsp; &nbsp; frequentative</p> <p>INCH&nbsp; &nbsp; &nbsp;inchoative</p> <p>INF&nbsp; &nbsp; &nbsp; &nbsp; infinitive</p> <p>ITR&nbsp; &nbsp; &nbsp; &nbsp; intransitive</p> <p>MOM&nbsp; &nbsp; &nbsp;momentaneous</p> <p>NEG&nbsp; &nbsp; &nbsp; negative</p> <p>NMLZ&nbsp; &nbsp; nominalization</p> <p>PASS&nbsp; &nbsp; passive</p> <p>PTCP&nbsp; &nbsp; participle</p> <p>REFL&nbsp; &nbsp; reflexive</p> <p>TRA&nbsp; &nbsp; &nbsp; transitive</p> <p>The authors of this dataset are Tomi Koivunen and Riku Erkkil&auml; and it is published under CC-BY-NC-ND licence. If used in a publication, please refer to this publication as well as mention the original source(s):</p> <p>This dataset has been used in following publications:</p> <p>&nbsp;</p> <p>References to used corpora:</p> <p>Arkhangelskiy, Timofey. 2018. <em>Udmurt corpus</em>. http://udmurt.web-corpora.net/index.html.</p> <p>Arkhangelskiy, Timofey. 2019a. <em>Komi-Zyrian corpus</em>. http://komi-zyrian.web-corpora.net/index.html.</p> <p>Arkhangelskiy, Timofey. 2019b. <em>Meadow Mari corpus</em>. http://meadow-mari.web-corpora.net/index_en.html.</p> <p>Fu-Lab team. 2021. <em>Корпус коми языка</em>. http://komicorpora.ru/.</p> <p>Helsingin yliopisto, FIN-CLARIN, H. Jauhiainen, T. Jauhiainen &amp; K. Lind&eacute;n. 2019. <em>Wanca 2016, Korp Version</em>. Kielipankki. http://urn.fi/urn:nbn:fi:lb-2019052401.</p> <p>Marko. (no year). <em>MARKO - Corpus of Mari language</em>. University of Turku.</p> <p>MokshEr, V.3. 2010. <em>Mok&scaron;an ja ers&auml;n s&auml;hk&ouml;inen korpus</em>. Turun yliopisto.</p> <p>Oncyko. 2000. <em>Oncyko corpus</em>. University of Turku.</p> <p>Permyak. 2009. <em>Turku Komi-Permyak Corpus</em>. University of Turku.</p>

opencc-by-4.0Sep 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record