Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

43

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

43 results for “Nouns”

Learn how ShareScore rates datasets ↗
zenodo48/100

Dataset containing binominal lexemes in Harakmbut (isolate, Peru), for "The derivational use of classifiers in Western Amazonia" and "When the alienability contrast fails to surface in adnominal possession: Bound nouns in Harakmbut"

<p>This is the dataset used, amongst others, in the paper: Van linden, An. Forthcoming. When the alienability contrast fails to surface in adnominal possession: Bound nouns in Harakmbut. Special Issue &ldquo;Re-assessing the explanatory potential of alienability contrasts&rdquo;, guest-edited by Fran&ccedil;oise Rose &amp; An Van linden. <em>Linguistics &ndash; An Interdisciplinary Journal of the Language Sciences</em>. [<a href="https://doi.org/10.1515/ling-2022-0039">https://doi.org/10.1515/ling-2022-0039</a>]</p> <p>For more details, see the ReadMe file.</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Between syntax and morphology: German noun+verb units (Glossa)

<p><strong>This dataset accompanies a paper to be published in Glossa.&nbsp;Under the present DOI, all data generated for this research as well as all&nbsp;scripts used are stored. The paper itself is CC-licensed, refer to glossa-journal.org.</strong></p><p><strong>Abstract</strong></p><p>We show that graphemic variation—at least in some writing systems—can be analysed in terms of grammatical variation given a usage-based probabilistic view of the grammar-graphemics interface. &nbsp; Concretely, we examine a type of noun+verb unit in German, which can be written as one word or two. We argue that the variation in writing is rooted in the units' ambiguous status in between morphology (one word) and syntax (two words). The major influencing factors are shown to be the semantic relation between the noun and the verb (argument or oblique relation) and the morphosyntactic context. In prototypically nominal contexts, a re-interpretation of the unit as a noun+noun compound is facilitated, which favours spelling as one word, while in prototypically verbal contexts, a syntactic realisation and consequently spelling as two words is preferred. We report the results of two large-scale corpus studies and a controlled production experiment to corroborate our analysis.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

On the order of demonstrative, numeral, adjective, and noun

<p>Cite the source of the dataset as:</p> <blockquote> <p>Dryer, M.S. (2018). On the order of demonstrative, numeral, adjective, and noun. Language 94(4), 798-833. doi:10.1353/lan.2018.0054.</p> </blockquote>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Supplementary data for the article "Head and dependent marking and dependency length in possessive noun phrases: a typological study of morphological and syntactic complexity"

<p>This material contains the Supplementary material (including the R-scripts)&nbsp;of the following article. Please cite the article when using the data.</p> <p>Sinnem&auml;ki, Kaius and Haakana, Viljami. 2023. Head and dependent marking and dependency length in possessive noun phrases: a typological study of morphological and syntactic complexity.&nbsp;<em>Linguistics Vanguard</em>&nbsp;9(s1).&nbsp;45-57.&nbsp;<a href="https://doi.org/10.1515/lingvan-2021-0074">https://doi.org/10.1515/lingvan-2021-0074</a></p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Tokens of the noun 'person' in the West Polesian corpus

<p>This dataset contains all the tokens of the noun &#39;person&#39; in free texts in the West Polesian corpus. Data were collected by Kristian Roncero in the Brest region (Belarus) between Jan 2016 and June 2017. Data are represented according to the IPA conventions, although punctuation marks are used and proper names have their first letter in upper case. I have tried to respect all the differences in the pronunciation, which means sometimes stems appear as palatalized (<em>ʧʲelovjek</em>-, <em>lʲud</em>-, as in Contemporary Standard Russian (CSR)); or most often unpalatalised (which is more in line with the general phonological rules of West Polesian) and the vocalism is not very consistent.</p> <p>The first letters of the code represent the village and speaker. The numbers after the speaker code indicate the file and the remaining the minute and second(s) where this sentence appears.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Table 1. The declination of feminine nouns (after [4])-ADX — Agent for Morphologic Analysis of Lexical Entries in a Dictionary

<p>An example of such searching trees for nouns of all genders (masculine, feminine, neuter)<br> together with the entire analysis made by the ADX agent for the word &ldquo;cepelor&rdquo; is presented in<br> figure 2. The pairs of numbers in brackets form lists of lines and columns from the inflections&rsquo;<br> charts where endings from the top of the list are found. We have noted the zero ending with &ndash; (that<br> is always the root of the tree) and the void list with * (Null/Nil).</p>

opencc-by-4.0Jan 2010View details →
zenodo40/100

rueter/gt-xml-lexica-sms: Skolt Sami to X nouns

<p>Skolt Sami nouns xml file for <a href="http://giellatekno.uit.no">http://giellatekno.uit.no</a> infrastructure use and applied in MediaWiki at <a href="https://sanat.csc.fi">https://sanat.csc.fi</a>, summer 2017.</p>

openother-openJan 2018View details →
zenodo40/100

rueter/gt-xml-lexica-myv: Erzya To X nouns

<p>Erzya nouns xml file for <a href="http://giellatekno.uit.no">http://giellatekno.uit.no</a> infrastructure use and applied in MediaWiki at <a href="https://sanat.csc.fi">https://sanat.csc.fi</a>, summer 2017.</p>

openother-openJan 2018View details →
zenodo40/100

rueter/gt-xml-lexica-kpv: Komi-Zyrian to X nouns

<p>Komi-Zyrian nouns xml file for <a href="http://giellatekno.uit.no">http://giellatekno.uit.no</a> infrastructure use and applied in MediaWiki at <a href="https://sanat.csc.fi">https://sanat.csc.fi</a>, summer 2017.</p>

openother-openJan 2018View details →
zenodo40/100

Data and results for "A corpus-based study to triangulating experimental evidence regarding verb-noun association for action verbs"

<p>This repository provides spreadsheets containing the results of corpus-based and experimental studies for my undergraduate thesis titled "A corpus-based study to triangulating experimental evidence regarding verb-noun association for action verbs" (supervised by Gede Primahadi Wijaya Rajeg, PhD [main] and Ketut Santi Indriani, M.Hum. [associate]) in the Bachelor of English Literature (BoEL) program, Faculty of Humanities, Udayana University. The thesis explores convergences/divergences between different methods and data types for a set of verb-noun collocations for several action verbs and their synonyms. The description of the dataset is as follows:</p> <ol> <li>"data-raw": A raw dataset containing the results of an experiment conducted using Gorilla Experiment Builder. This consists of responses regarding verb-noun collocation co-occurrences from 17 participants into one. Link to Gorilla Experiment: (https://app.gorilla.sc/openmaterials/622948).</li> <li>"Corpus Analysis Results": A compiled data containing search results of frequencies found in the Corpus of Contemporary American English (COCA). The frequencies were compiled into tables for the five main verbs showing the number of co-occurrences of specific verb-noun collocations.</li> <li>"Experiment Results (1)": A compiled data containing the experiment results calculated as a total, showing the number of co-occurrences of specific verb-noun collocations across five verbs from the participant responses.</li> </ol> <p>The thesis is part of the pedagogical outcome of the&nbsp;<a title="CompLexico" href="https://www.cirhss.org/complexico/" target="_blank" rel="noopener"><em>CompLexico</em></a> research group at <a title="CIRHSS" href="https://www.cirhss.org/" target="_blank" rel="noopener"><em>CIRHSS</em></a>, and the Psycholinguistics course I took with I Made Sena Darmasetiyawan, PhD at BoEL, both in the Faculty of Humanities, Udayana University.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Abstractions and exemplars: the measure noun phrase alternation in German

<p>This is the data set and scripts used in &quot;Abstractions and exemplars: the measure noun phrase alternation in German&quot; as accepted for publication in Cognitive Linguistics as of 29 may 2018.</p> <p>In this paper, an alternation in German measure noun phrases is examined under a varying-abstraction perspective. In a specific measure NP construction, the embedded kind-denoting noun either agrees in case with the measure noun (&quot;eine Tasse guter Kaffee&quot; &#39;a cup of good coffee&#39;) or it stands in the genitive (&quot;eine Tasse guten Kaffees&quot;). Each of the two alternants is syntactically similar to a non-alternating construction. I propose a prototype model which assigns a common prototypical meaning to each of the alternants and its corresponding non-alternating construction. Based on this, I argue that lexical, morpho-syntactic, and stylistic features help to predict the choice of the alternant. A large corpus study is presented which supports this analysis. However, in addition to the prototype effects, an exemplar effect is also shown to influence the choice, namely the relative frequencies with which lemmas occur in the non-alternating constructions. I argue that allowing both prototype and exemplar effects is more adequate than following radical prototype or exemplar approaches. It is also verified in two experiments that the corpus-derived model corresponds to the behaviour of native speakers. The weak effect size of the experimental validation is discussed in the context of corpus-based cognitive linguistics and the validation of corpus-derived models.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Inflected lexicon of Russian Nouns in IPA notation

<p>This inflected lexicon of Russian Nouns is based on data generated by a DATR fragment for the nominal system of Russian (Dunstan Brown et al, 2011), which was used in Brown and Hippisley (2012). The files were automatically transcribed to an IPA-based notation by Sacha Beniamine (2018), following rules provided by Dunstan Brown. We also corrected a few errors in the original data found manually. We thank Sasha krasovitsky for providing comments on the phonological feature definitions.</p>

opengpl-3.0Sep 2019View details →
zenodo40/100

TYPES: Male holotype from Panama: Panama: Parque Nacional Altos de Campana, 1 hectare PANCODING Inventory, 895 m, 8.68333°, -79.92972°, June 14–19, 2007, M. Arnedo, D. Dimitrov, G. Hormiga, F. Labarque, M. Ramírez, deposited in MIUP, PBI_OON 42313; same data, 1 male paratype deposited in MACN-Ar 29895, PBI_OON 42312. ETYMOLOGY: A noun in apposition; in Greek religion and mythology, Pan is the god of the wild natural world, of shepherds, flocks, and mountains, and of hunting and rustic music. He has hindquarters, legs, and horns of a goat, and the name is here employed to note the large mac- rosetae at the eye region of males that resemble the horns in some illustrations of this god. DIAGNOSIS: This is one of the most autapomor- phic species from the Americas; males have the labium fused with the sternum (fig. 34B), small chelicerae, shorter than the endite length, with anterior blunt projections, and directed backward in lateral view (fig. 34D, E); clypeus directed back- ward (fig. 34D); two light areas on the sternum just below the endites (fig. 34B), carapace almost flat in lateral view and two strong macrosetae at the eye region, pointing forward (fig. 34C–E). Other characters of the male palp, such as the presence of two apophyses, also distinguish this species from others (fig. 38D–F). MALE (PBI_OON 42312): Total length 1.00. Habitus as in figure 34A–C. CEPHALOTHO- RAX: Carapace orange, with brown stripe along in Taxonomic Revision Of The Jumping Goblin Spiders Of The Genus Orchestina Simon, 1882, In The Americas (Araneae: Oonopidae)

TYPES: Male holotype from Panama: Panama: Parque Nacional Altos de Campana, 1 hectare PANCODING Inventory, 895 m, 8.68333°, -79.92972°, June 14–19, 2007, M. Arnedo, D. Dimitrov, G. Hormiga, F. Labarque, M. Ramírez, deposited in MIUP, PBI_OON 42313; same data, 1 male paratype deposited in MACN-Ar 29895, PBI_OON 42312. ETYMOLOGY: A noun in apposition; in Greek religion and mythology, Pan is the god of the wild natural world, of shepherds, flocks, and mountains, and of hunting and rustic music. He has hindquarters, legs, and horns of a goat, and the name is here employed to note the large mac- rosetae at the eye region of males that resemble the horns in some illustrations of this god. DIAGNOSIS: This is one of the most autapomor- phic species from the Americas; males have the labium fused with the sternum (fig. 34B), small chelicerae, shorter than the endite length, with anterior blunt projections, and directed backward in lateral view (fig. 34D, E); clypeus directed back- ward (fig. 34D); two light areas on the sternum just below the endites (fig. 34B), carapace almost flat in lateral view and two strong macrosetae at the eye region, pointing forward (fig. 34C–E). Other characters of the male palp, such as the presence of two apophyses, also distinguish this species from others (fig. 38D–F). MALE (PBI_OON 42312): Total length 1.00. Habitus as in figure 34A–C. CEPHALOTHO- RAX: Carapace orange, with brown stripe along

opencc-by-4.0Feb 2017View details →
zenodo40/100

Animacy and non-animacy denoting nouns

<p>Universit&auml;t Z&uuml;rich<br> Institut f&uuml;r Computerlinguistik<br> Manfred Klenner</p> <p>Die Daten sind im Rahmen eines Projekts, das vom Schweizer Nationalfond gef&ouml;rdert wurde (Nr. 105215-179302), entstanden.<br> ------------------------------------------------------------------------------------------------------------------------<br> **License**: Creative Commons Attribution-ShareAlike 4.0 International Public License</p> <p><br> Repository: animacy data for animcay classification</p> <p>This is the training data for an animacy classifier (see References LREC)</p> <p>1) gold_actor:&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; &nbsp; 7468 nouns denoting animate entities<br> 2) gold_nonactor &nbsp;&nbsp; &nbsp; &nbsp; 5511 nouns denoting non-animate entities</p> <p>subsets of 1:</p> <p>gold_direct&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; &nbsp; 6897 nouns directly denoting animate entities<br> gold_metonym&nbsp;&nbsp; &nbsp;&nbsp;&nbsp; &nbsp; &nbsp; 587 &nbsp;metonymy trigger nouns</p> <p><br> Format: just lists</p> <p>Note: although some person names are in the data, a separate NER for person names should be used .</p> <p>References:</p> <p>@inproceedings{LREC,<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;month = {Juni},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; author = {Manfred Klenner and Anne G{\&quot;o}hring},<br> &nbsp; &nbsp; &nbsp; &nbsp;booktitle = {Proceedings of the Language Resources and Evaluation Conference},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;address = {Marseille, France},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;title = {Animacy Denoting {G}erman Nouns: Annotation and Classification},<br> &nbsp; &nbsp; &nbsp; &nbsp;publisher = {European Language Resources Association},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;pages = {1360--1364},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; year = {2022},<br> &nbsp; &nbsp; &nbsp; &nbsp; language = {english},<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;url = {https://doi.org/10.5167/uzh-219148},<br> &nbsp; &nbsp; &nbsp; &nbsp; abstract = {In this paper, we introduce a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities. W<br> e present inter-annotator agreement of our crowd-sourced seed annotations (9,000 nouns) and discuss the results of machine learning models applied to this data.}<br> }<br> &nbsp;</p>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Supplementary material for "Using a parallel corpus to study patterns of word order variation: Determiners and quantifiers within the noun phrase in European languages"

<p>- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files.</p> <p>Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.</p>

opencc-by-4.0May 2023View details →
zenodo36/100

Dataset and documented R code for "Nouns and verbs in the speech signal"

<p>The files available constitute supplementary material to the following article:</p> <p>Lohmann, Arne. Nouns and verbs in the speech signal: Are there phonetic correlates of grammatical category? <em>Linguistics</em> - <em>An Interdisciplinary Journal of the Language Sciences</em>.</p> <p>The article is to be published online in 2020, and in 2021 in the print version of the journal.</p>

opencc-by-4.0Jan 2020View details →
zenodo36/100

Wikidata Dump nouns_ru_en

<p> RDF dump of wikidata produced with <a href="https://tools.wmflabs.org/wdumps/">wdumps</a>. </p> <p> Nouns Russian, English with grammar gender<br> <a href="https://tools.wmflabs.org/wdumps/dump/229">View on wdumper</a> </p> <p> <b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0 </p>

opencc-zeroApr 2020View details →
zenodo36/100

Event/Nonevent Noun Disambiguation Test Set

<p>The dataset is designed to test classifiers&rsquo; ability to determine whether a nominal, in a context, has an eventive sense or not. The test dataset contains examples of eventive and non-eventive use of 3 ambiguous English nouns - &quot;publication&quot;, &quot;organization&quot; and &quot;construction&quot;. There are ~ 150 sentences per noun: ~75 sentences in which the noun has an eventive sense and ~75 sentences in which this is not the case. The examples are taken from internet&nbsp;news&nbsp;sites.&nbsp;</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

Dataset for 'Beyond bidirectional association: Distinguishing light verb constructions from other conventionalised noun-verb combinations in modern Tibetan'

Open the record for dataset details and reuse information.

opencc-by-4.0Mar 2024View details →
zenodo36/100

ParaKasem: Kasem noun dataset

<p>Paralex complient Kasem dataset. Please cite:</p> <p>Guzm&aacute;n Naranjo, Mat&iacute;as. 2019. Analogical Classification in Formal Grammar. Language Science Press, Berlin. doi:10.5281/zenodo.3191825.</p>

opencc-by-4.0Nov 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record