Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
43
datasets available to search
ShareScore release 0.7.1
Dataset results
43 results for “Nouns”
Armenian: Noun phrase structure and determination
<ul> <li> <p><strong>Basic word order in noun phrase</strong></p> </li> <li> <p><strong>Basic characteristics of definite article</strong></p> </li> <li> <p><strong>Parallels between noun phrase and clause</strong></p> </li> <li> <p><strong>Definite article and specificity</strong></p> </li> <li> <p><strong>Definite article and nominalization</strong></p> </li> <li> <p><strong>Definite article as marker of argumenthood</strong></p> </li> <li> <p><strong>Typological parallels</strong></p> </li> </ul> <p> </p> <p>This lecture is part of the lecture series:</p> <p><em>Glottothèque: Languages of the Anatolia, Caucasus, Iran, Mesopotamia; grammatical snippets online </em>(electronic resource). Bamberg, Cambridge, Göttingen, Moskow, Nicosia, Paris: LACIM network, at https://spw.uni-goettingen.de/projects/lacim/, edited by Christiane Bulut, Anaïd Donabédian-Demopoulos, Geoffrey Haig, Geoffrey Khan, Pollet Samvelian, Stavros Skopeteas, Nina Sumbatova.</p>
cldf-datasets/dryerorder: Dataset of Dryer 2018 "On the order of demonstrative, numeral, adjective, and noun"
<p><strong>Dryer, M.S. (2018). On the order of demonstrative, numeral, adjective, and noun. Language 94(4), 798-833. doi:10.1353/lan.2018.0054.</strong></p>
Russian Event Noun Disambiguation Test Set (6x200)
<p>The dataset is designed to test classifiers’ ability to determine whether a nominal, in a context, has an eventive reading or not. The test dataset contains examples of eventive and non-eventive use of 6 ambiguous Russian nouns. There are 200 sentences per noun: 100 sentences in which the noun has an eventive reading and 100 sentences in which this is not the case. The examples are taken from internet news sites. </p> <p> </p>
Data Outputs: An analysis of proper nouns in Marian Keyes published novels 1995-2020.xlsx
<p>A data set analysing proper nouns in Marian Keyes' published novels 1995-2020 conducted for my Irish Research Council funded PhD thesis: </p> <ol> <li>Watermelon</li> <li>Lucy Sullivan is Getting Married</li> <li>Rachel's Holiday</li> <li>Last Chance Saloon</li> <li>Sushi for Beginners</li> <li>Angels</li> <li>The Other Side of the Story</li> <li>Anybody Out There?</li> <li>This Charming Man</li> <li>The Brightest Star in the Sky</li> <li>The Mystery of Mercy Close</li> <li>The Woman Who Stole My Life</li> <li>The Break</li> <li>Grown Ups</li> </ol> <p>The results are not statistically significant and therefore not incorporated into my final write up. </p> <p>The list of character names may be useful to incoroprate into a STOP list for anyone else seeking to analyse Keyes' novels via distant reading methods.</p>
Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff
<p>Dataset for publication "Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff"</p> <p>To appear in: SKASE Journal of Theoretical Linguistics, ISSN: <a href="https://portal.issn.org/resource/ISSN/1336-782X">1336-782X</a></p> <p>Authors: Olav Mueller-Reichau (University of Leipzig, Germany) and Matthias Irmer (OntoChem GmbH, Halle (Saale), Germany)</p> <p> </p>
Silver Standard of quantified positive, negative and neutral German noun phrases
<p>Repository: silver standard quantified (simple) noun phrases</p> <p>43,842 positive (+), negative (-) or neutral (0) NPs, </p> <p>e.g. "ein notorischer Verführer -5.625" is highly negative (-5.625)</p> <p>see References LREC for a description of the data</p> <p>Format: tsv</p> <p>References:</p> <p>@inproceedings{LREC,<br> month = {Juni},<br> author = {Manfred Klenner and Anne G{\"o}hring},<br> booktitle = {Proceedings of the Language Resources and Evaluation Conference},<br> address = {Marseille, France},<br> title = {Animacy Denoting {G}erman Nouns: Annotation and Classification},<br> publisher = {European Language Resources Association},<br> pages = {1360--1364},<br> year = {2022},<br> language = {english},<br> url = {https://doi.org/10.5167/uzh-219148},<br> abstract = {In this paper, we introduce a gold standard for animacy detection comprising almost 14,500 German nouns that might be used to denote either animate entities or non-animate entities. We present inter-annotator agreement of our crowd-sourced seed annotations (9,000 nouns) and discuss the results of machine learning models applied to this data.}<br> }<br> </p>
Breaking the Mold: A corpus study of numeral+noun phrases in Scottish Gaelic - Corpus Data
<p>Hello! This is the corpus data compiled in the searches done for my master's thesis: Breaking the Mold: A corpus study of numeral + noun phrases in Scottish Gaelic. The data comes from the Digital Archive of Scottish Gaelic's Corpas na Gàidhlig (<a href="https://dasg.ac.uk/corpus/">https://dasg.ac.uk/corpus/</a>). </p> <p>The relevant data are discussed in the thesis, but the full corpus is found here for anyone who wishes to view it. </p> <p>For any questions, please contact me: emma.mckenzie(at)helsinki.fi </p>
Parahungarian : a dataset of Hungarian nouns
Open the record for dataset details and reuse information.
Nominal Anchoring Functions of Porohanon Common Noun Markers
<p>Porohanon, spoken in the Municipality of Poro, Camotes, Cebu, Philippines is a member of the Peripheral group of the Central Bisayan branch of the Bisayan complex. This study proposes a SPECIFIC vs. NONSPECIFIC contrast in the ABSOLUTIVE, ERGATIVE, and GENITIVE forms of the common noun markers, instead of a DEFINITE vs. INDEFINITE contrast stated in previous descriptions of the speech variety.</p>
Supplementary material for "Tweaking UD annotations to investigate the placement of determiners, quantifiers and numerals in the noun phrase"
<p>The supplementary material includes</p> <ul> <li>the Python script used to extract relevant word order patterns from the parsed corpus;</li> <li>the R script used to process the extracted data, also containing the list of lemmata;</li> <li>a spreadsheet featuring the frequency and entropy values.</li> <li>a doc file with comparative concepts for the investigated categories.</li> </ul>
Adjective+noun phrases that readers use in reviews to refer to the style of the text
<p>Adjective+noun phrases that readers use in reviews to refer to the style of the text</p> <table> <tbody> <tr> <td> </td> <td> <p>Stijl</p> </td> <td> <p>Taal</p> </td> <td> <p>Toon</p> </td> <td> <p>Woord</p> </td> <td> <p>Zin</p> </td> <td> <p>Uitdrukking</p> </td> </tr> <tr> <td> <p>Reader's experience</p> </td> <td> <p>29%</p> </td> <td> <p>57%</p> </td> <td> <p>77%</p> </td> <td> <p>39%</p> </td> <td> <p>28%</p> </td> <td> <p>34%</p> </td> </tr> <tr> <td> <p>Qualitative evaluation</p> </td> <td> <p>26%</p> </td> <td> <p>4%</p> </td> <td> <p>6%</p> </td> <td> <p>10%</p> </td> <td> <p>24%</p> </td> <td> <p>7%</p> </td> </tr> <tr> <td> <p>Features of text</p> </td> <td> <p>5%</p> </td> <td> <p>7%</p> </td> <td> <p>1%</p> </td> <td> <p>16%</p> </td> <td> <p>39%</p> </td> <td> <p>4%</p> </td> </tr> <tr> <td> <p>Other references</p> </td> <td> <p>10%</p> </td> <td> <p>6%</p> </td> <td> <p>12%</p> </td> <td> <p>4%</p> </td> <td> <p>4%</p> </td> <td> <p>5%</p> </td> </tr> <tr> <td> <p>Author</p> </td> <td> <p>19%</p> </td> <td> <p>4%</p> </td> <td> <p>3%</p> </td> <td> <p>4%</p> </td> <td> <p>1%</p> </td> <td> <p>1%</p> </td> </tr> <tr> <td> <p>Region</p> </td> <td> <p>1%</p> </td> <td> <p>19%</p> </td> <td> <p>1%</p> </td> <td> <p>10%</p> </td> <td> <p>2%</p> </td> <td> <p>40%</p> </td> </tr> <tr> <td> <p>Time</p> </td> <td> <p>2%</p> </td> <td> <p>3%</p> </td> <td> <p>1%</p> </td> <td> <p>2%</p> </td> <td> <p>0%</p> </td> <td> <p>3%</p> </td> </tr> <tr> <td> <p>Quantitative evaluation</p> </td> <td> <p>0%</p> </td> <td> <p>1%</p> </td> <td> <p>0%</p> </td> <td> <p>12%</p> </td> <td> <p>2%</p> </td> <td> <p>5%</p> </td> </tr> </tbody> </table>
Russian Event/Non-Event Noun Disambiguation Data
<p>The dataset contains training, validation and test sets for classifiers designed for context-based discrimination between eventive and non-eventive reading of Russian nouns. The set also includes supplementary materials (word embedding models and lists of unambiguous nouns that were used for training set generation).</p>
Stress Shift in Noun-Verb Conversion Pairs: Data
Open the record for dataset details and reuse information.
LINGUISTIC AND CULTURAL FEATURES OF THE PROPER NOUNS IN FAIRY TALES AND ITS TRANSLATION INTO RUSSIAN
Open the record for dataset details and reuse information.
THE ROLE OF NOUNS IN ENGLISH LANGUAGE AND ITS IMPROVEMENT
Open the record for dataset details and reuse information.
OBJECT PHRASES INCLUDING NOUNS AND NON-FINITE VERBS IN THE SENTENCES AND THEIR TYPES IN THE CONTEXT
Open the record for dataset details and reuse information.
On the Role of Modifiers in Interpretation of Noun-Noun Compounds in English
<p>Study 1: Corpus analysis of noun-noun compounds </p> <p>Study 2: Compound creation study</p>
Syntactic and Semantic Boundness of Noun Phrases in Burmese
<p>This paper investigates the correlation of form and function in Burmese noun phrases. Burmese has several ways of combining nouns with verbal and other modifiers, which exhibit different degrees of syntactic boundness and semantic integration. The claim is that syntactic and semantic boundness are predictive in that syntactically more tightly bound expressions coincide with closer semantic integration. The results show that this claim partly holds in Burmese, though it is not the only factor determining the choice of a specific construction in a given context.</p>
Zero-derived nouns and deverbal nominalization: Database for German
<p>This is a collection of deverbal zero-derived nouns with various information on their date of attestation, etymology, frequency, possible interpretations and their ability to realize verbal argument structure. Most of this information is extracted from lexical resources, in particular dictonaries. Examples with argument structure are extracted from natural text corpora. There are four collections in total for English, Italian, Spanish and German. The English and the Italian collections follow a parallel structure, while the Spanish and the German one are more targeted on the research purposes of the students who created them.</p>
Zero-derived nouns and deverbal nominalization: Databases for English and Italian
<p>This is a collection of deverbal zero-derived nouns with various information on their date of attestation, etymology, frequency, possible interpretations and their ability to realize verbal argument structure. Most of this information is extracted from lexical resources, in particular dictonaries. Examples with argument structure are extracted from natural text corpora. There are four (not exhaustive) collections in total for English, Italian, Spanish and German. The English and the Italian collections follow a parallel structure, while the Spanish and the German one are more targeted on the research purposes of the students who created them.</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.