Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22
datasets available to search
ShareScore release 0.9.0
Dataset results
22 results for “corpus studies”
A Corpus of Biblical Names in the Greek New Testament to Study the Additions, Omissions, and Variations across Different Manuscripts
<p>The analysis of textual variants of verses in the Ancient Greek New Testament across different manuscripts has mainly been done by close reading with manual effort. With the increasing number of transcriptions of the different manuscripts, quantitative analyses (so-called distant reading) can be used to search for patterns of omission, addition, or other variations, to formulate novel hypotheses to be investigated by close reading. In this work, we present a corpus of biblical names including spelling variation and inflections and their mentions in the transcriptions of the Ancient Greek New Testament.</p>
FiloBass: A Dataset and Corpus Based Study of Jazz Basslines
<p>Dataset to accompany the paper "FiloBass: A Dataset and Corpus Based Study of Jazz Basslines" which was published at ISMIR 2023.</p>
The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms
<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks. As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world’s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish to make, and their obtained results obtained through cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and the researchers should send the predictions of their models on these recordings. Finally, the performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were modeled using ML and expanded on the previously unlabeled data. These labels can be found in the "status_SSL" column of the "metadata_compiled.csv" file.</p>
A Corpus of Publicly Available Simulink Models for Model-based Empirical Studies
<p>Abstract: Recent years have seen many empirical studies of model-based cyber-physical systems and commercial CPS development tool chains such as Matlab/Simulink. To benefit such research, this paper presents the by-far largest corpus of freely available Simulink models to date, containing over 1,000 models.</p> <p>Surprising findings based on this corpus include that (a) tool support for metric collection is not adequate and (b) users do not reuse model components as they would in object-oriented programs.</p> <p>The paper both confirms and contradicts earlier findings that are based on significantly fewer models, suggesting the utility of the corpus for future research. While others have not yet leveraged this model corpus, we hope that our freely available corpus and infrastructure will benefit future model-based empirical research and tool development efforts, by reducing the model-collection overhead and thus easing evaluation.</p> <p>Please see our paper "A curated corpus of Simulink models for model-based empirical studies" In Proc. 4th International Workshop on Software Engineering for Smart Cyber-Physical Systems (SEsCPS)</p> <p>See https://github.com/verivital/slsf_randgen/wiki for more details including information related to downloading the Simulink models.</p> <p>http://ranger.uta.edu/~csallner/csallner_bib.html#Chowdhury18Curated</p>
childPoeDE: A corpus of German Children's Poems for Computational and Experimental Studies - Metadata
<p>The childPoeDE corpus is a collection of 1082 German poems for children created within the CHYLSA project. The poems were taken from anthologies published between 1991 and 2019. This publication includes the poem-level metadata for each poem with information about the author, the poem's length, data on case, punctuation, layout, rhyme, type-token ratio (TTR and MATTR) and lexical density. It also includes token-level metadata, namely word length and position, POS tags in different levels of granularity as well as data on onomatopoeia and sonority. Furthermore, this publication provides a word frequency table and a Python script which was used to extract some of the metadata from the texts (poemtool.py). The childPoeDE corpus does not contain all poems from the anthologies. A list of the poems that have been omitted for different reasons (length, language, typography, ...) can be accessed as well.</p> <p>Read more about the childPoeDE corpus in our data paper published in the Journal of Open Humanities Data: <a href="https://doi.org/10.5334/johd.102">The ChildPoeDE Corpus: 1082 German Children’s Poems for Computational and Experimental Studies on Poetry Reception</a>.</p> <p>DFG Schwerpunktprogramm SPP 2207 “Computational Literary Studies“<br> Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/402743989">https://gepris.dfg.de/gepris/projekt/402743989</a></li> <li><a href="https://dfg-spp-cls.github.io/">https://dfg-spp-cls.github.io<em>/</em></a></li> </ol> <p>Subproject: „CHYLSA (Children’s and Youth Literature Sentiment Analysis)“</p> <p>Online:</p> <ol> <li><a href="https://gepris.dfg.de/gepris/projekt/424250469">https://gepris.dfg.de/gepris/projekt/424250469</a></li> <li><a href="https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/">https://dfg-spp-cls.github.io/projects_en/2020/01/24/TP-CHYLSA/</a></li> </ol>
Using Corpus Studies to Find the Origins of the Madrigal: Music and Feature Values
<p>This distribution includes files associated with the experiments described in the "Using Corpus Studies to Find the Origins of the Madrigal" paper presented at the 2021 Future Directions of Music Cognition conference (http://org.osu.edu/mascats/). Our group encoded all the music from the original sources ourselves, using Sibelius, and exported the Sibelius data as PDF and MIDI files. The details of the corpus are described in the "Florence 164 Metadata.xlsx" file and in the conference paper itself.</p> <p>All files included in this archive are distributed under a "CC BY-SA 4.0" license" license (https://creativecommons.org/licenses/by-sa/4.0/). </p> <p>All features were extracted using jSymbolic 2.2 (http://jmir.sourceforge.net) directly from the MIDI encodings included here. Details on all the individual features extracted with the software are available in the jSymbolic manual (http://jmir.sourceforge.net/manuals/jSymbolic_manual/home.html). The features are presented as follows in the "F164_Extracted_Features" folder:</p> <p>- F164_Feature_Definitions.xml: Descriptions of all extracted features, encoded in ACE XML 1.0, as output directly by jSymbolic. This file does not itself include any feature values.</p> <p>- F164_Feature_Values.xml: Extracted feature values, encoded in ACE XML 1.0, as output directly by jSymbolic. The features are described in the FeatureDefinitions.xml file. Class associations are implied by the folder containing each MIDI file.</p> <p>- F164_Feature_Values_BasicCSV: Extracted feature values encoded in a CSV file, as output directly by jSymbolic. Class associations are implied by the folder containing each MIDI file.</p> <p>- F164_Feature_Values_HumanReadable.xlsx: Extracted features formatted into a human-readable Microsoft Excel file. Group averages and standard deviations have been added to the bottom.</p> <p>- F164_Feature_Values_WekaReady.csv: Extracted feature values encoded in a CSV file in a format readable by Weka (https://www.cs.waikato.ac.nz/ml/weka/). Class values have been added in a column on the right, and file paths have been removed, as required by Weka.<br> </p>
Additional Data to "A Challenge for Contrastive L1/L2 Corpus Studies"
<p>Additional plots and data to the paper "A Challenge for Contrastive L1/L2 Corpus Studies: Large Inter- and Intra-Individual Variation Across Morphological, but Not Global Syntactic Categories in Task-Based Corpus Data of a Homogeneous L1 German Group", published in Frontiers of Psychology on Nov 25th, 2021.<br> <br> This repository contains the Falko data and additional plots of the individual distributions of morphological noun classes, frequent dependency types, and parts of speech in the Falko data. The Kobalt data discussed in the paper can be found in Shadrova (2021).</p>
A corpus-based study of the acquisition of the English progressive by L1 Chinese learners: From prototypical activities to marked statives
<p>This article investigates how EFL learners’ progressive markings are influenced by the lexical aspect of verbs, modality (spoken vs. written), and proficiency levels, focusing on the controversial issue of stative verbs in progressives in L2 acquisition. Spoken (SECCL) and written (WECCL) corpus data from two proficiency levels of Chinese EFL learners and comparison data from native English speakers (COCA) were analyzed. The results suggest that in both learner and native data the progressive -<em>ing</em> is strongly associated with activity verbs, stative verbs being least likely to be inflected with the progressive<em>,</em> as predicted by the Aspect Hypothesis (Andersen and Shirai 1994, 1996). However, inconsistently with the Aspect Hypothesis, this association strengthens with higher proficiency levels. Learners’ use of stative verbs in the progressive and the overextended use of stative progressives was also found to be related to spoken vs. written mode of production and proficiency levels, with learners retreating from overextension as their proficiency increases. A usage-based account of the findings is proposed.</p>
PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)
<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>
Data and results for "A corpus-based study to triangulating experimental evidence regarding verb-noun association for action verbs"
<p>This repository provides spreadsheets containing the results of corpus-based and experimental studies for my undergraduate thesis titled "A corpus-based study to triangulating experimental evidence regarding verb-noun association for action verbs" (supervised by Gede Primahadi Wijaya Rajeg, PhD [main] and Ketut Santi Indriani, M.Hum. [associate]) in the Bachelor of English Literature (BoEL) program, Faculty of Humanities, Udayana University. The thesis explores convergences/divergences between different methods and data types for a set of verb-noun collocations for several action verbs and their synonyms. The description of the dataset is as follows:</p> <ol> <li>"data-raw": A raw dataset containing the results of an experiment conducted using Gorilla Experiment Builder. This consists of responses regarding verb-noun collocation co-occurrences from 17 participants into one. Link to Gorilla Experiment: (https://app.gorilla.sc/openmaterials/622948).</li> <li>"Corpus Analysis Results": A compiled data containing search results of frequencies found in the Corpus of Contemporary American English (COCA). The frequencies were compiled into tables for the five main verbs showing the number of co-occurrences of specific verb-noun collocations.</li> <li>"Experiment Results (1)": A compiled data containing the experiment results calculated as a total, showing the number of co-occurrences of specific verb-noun collocations across five verbs from the participant responses.</li> </ol> <p>The thesis is part of the pedagogical outcome of the <a title="CompLexico" href="https://www.cirhss.org/complexico/" target="_blank" rel="noopener"><em>CompLexico</em></a> research group at <a title="CIRHSS" href="https://www.cirhss.org/" target="_blank" rel="noopener"><em>CIRHSS</em></a>, and the Psycholinguistics course I took with I Made Sena Darmasetiyawan, PhD at BoEL, both in the Faculty of Humanities, Udayana University.</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 4. The frequency of different source segments of "but" in the Cantonese sub-corpus
<p>Likewise, Figure 4 lists all the Cantonese source segments of but and their frequency. The results show that most of the use of but was contributed by daan (hai) — its closest equivalence in Cantonese. However, there were 52 instances of the use of but which corresponded to no source segments at all in the Cantonese sub-corpus, indicating, again, the possibility of explicitation. The rest of the 16 instances of but were attributed by the use of the Cantonese markers ji (而 (frequency=8; close in meaning to however and but, indicating concession or contrast), followed by bat gwo (frequency=7) and koek (卻 (frequency=1; close in meaning to however and but, indicating concession or contrast).</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 3. The frequency of different source segments of "however" in the Cantonese sub-corpus
<p>The use of however and but in the English sub-corpus was also investigated through looking into their source segments in the Cantonese sub-corpus. Figure 3 lists all the Cantonese source segments of however and their frequency. The figures show that the use of however mostly resulted from the employment of daan (hai) in the Cantonese source text (frequency=52). There were 24 instances of however which corresponded to no source segment at all, indicating a possible trend of explicitation — a feature of interpreted or translated language. Only 21 cases of however were caused by the use of bat gwo — its closest equivalence in Cantonese. The rest 4 cases resulted from the use of the Cantonese marker ho si (可是) (frequency=3; close in meaning to both however and but in English, indicating concession or contrast) and zeon gun (儘管 (frequency=1; close in meaning to although or in spite of in English, indicating concession or topic change).</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 2. The frequency of different renditions of "daan (hai)" in the English sub-corpus
<p>The renditions of daan (hai) show a similar pattern (Figure 2). Among the 287 cases of daan (hai), the majority were interpreted into but (frequency=74) — its closest equivalence in English that signals denial and contrast. A total of 60 daan (hai) received no interpretation at all, which, again, indicates a possible mitigation strategy employed by the interpreter(s). Similar to the case of bat gwo, however — the other frequently used English marker apart from but — was the next most often employed rendition of daan (hai) in the English sub-corpus. In addition to but and however, daan (hai) was also rendered into the following English markers: although, despite, while, nevertheless, yet, having said that / that said, nonetheless, though, notwithstanding, on the other, regardless, even so, and after, most of which signal concession and topic change, with an even higher degree of subtlety as compared to however.</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 1. The frequency of different renditions of "bat gwo" in the English sub-corpus
<p>The renditions of bat gwo and daan (hai) were closely examined by looking into how they were interpreted. Figure 1 lists all the renditions of bat gwo and their frequency. The figures show that bat gwo was most often interpreted into however (frequency=22) — its closest equivalence in English. There were, however, 7 cases that bat gwo was interpreted into but — its stronger and less subtle correspondence that indicates denial and contrast. Apart from rendering into these two most common English contrastive markers, there were also 5 cases that bat gwo was not interpreted at all, suggesting a possible mitigation strategy employed by the interpreter(s). Likewise, the rest of bat gwo were rendered into other markers including nevertheless, nonetheless, while, yet, having said that / that said, all of which indicate concession and topic change, yet with an even higher degree of subtlety as compared to however.</p>
Supplementary material for "Using a parallel corpus to study patterns of word order variation: Determiners and quantifiers within the noun phrase in European languages"
<p>- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files.</p> <p>Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.</p>
Offering and showing gestures in 12- to 15-month-old infants in natural contexts: A corpus-based study
<p>Supplementary material for an original article titled "Offering and showing gestures in 12- to 15-month-old infants in natural contexts: A corpus-based study", submitted to the European Journal of Developmental Psychology.</p>
Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff
<p>Dataset for publication "Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff"</p> <p>To appear in: SKASE Journal of Theoretical Linguistics, ISSN: <a href="https://portal.issn.org/resource/ISSN/1336-782X">1336-782X</a></p> <p>Authors: Olav Mueller-Reichau (University of Leipzig, Germany) and Matthias Irmer (OntoChem GmbH, Halle (Saale), Germany)</p> <p> </p>
R Notebook and Dataset for "Usage-based perspective on argument realisation: A corpus study of Indonesian BUY verbs in applicative construction with -kan" (1.0.0)
<p>This repository contains the dataset and R codes for our paper that has been published in <a href="http://www.aa.tufs.ac.jp/en/publications/nusa">NUSA</a> (<i>Linguistic studies of languages in and around Indonesia</i>) special volume (74) on "Applicatives in Austronesian Languages".</p><h4>How to cite the paper</h4><p>Rajeg, Gede Primahadi Wijaya & I Wayan Arka. 2023. Usage-based perspective on argument realisation: A corpus study of Indonesian BUY verbs in applicative construction with -<i>kan</i>. In Jocelyn Aznar, Christian Döhler & Jozina Vander Klok (eds.), <i>NUSA (special issue on "Applicatives in Austronesian Languages")</i>, vol. 74, 83–114. <a href="https://tufs.repo.nii.ac.jp/records/2000019">https://tufs.repo.nii.ac.jp/records/2000019</a>.</p><h4>Description of the repository</h4><p>The .qmd file contains the R codes used to produce the quantitative analyses in the paper, including the statistical figures. This .qmd file also interweaves some text narratives with the codes. The file is published as a webpage at: <a href="https://gederajeg.github.io/applicative-buy/">https://gederajeg.github.io/applicative-buy/</a></p><p>The raw, annotated concordance data is located in the <a href="https://github.com/gederajeg/applicative-buy/tree/main/data">data</a> directory.</p><p>The statistical figures can also be accessed individually <a href="https://github.com/gederajeg/applicative-buy/tree/main/nusa-applicative-code_files/figure-html">here</a>.</p>
Breaking the Mold: A corpus study of numeral+noun phrases in Scottish Gaelic - Corpus Data
<p>Hello! This is the corpus data compiled in the searches done for my master's thesis: Breaking the Mold: A corpus study of numeral + noun phrases in Scottish Gaelic. The data comes from the Digital Archive of Scottish Gaelic's Corpas na Gàidhlig (<a href="https://dasg.ac.uk/corpus/">https://dasg.ac.uk/corpus/</a>). </p> <p>The relevant data are discussed in the thesis, but the full corpus is found here for anyone who wishes to view it. </p> <p>For any questions, please contact me: emma.mckenzie(at)helsinki.fi </p>
Study Comparing Corpus Callosum Atrophy as a Marker of Later Development of Cognitive Impairment in Patients With Multiple Sclerosis
ClinicalTrials.gov study NCT01250665. IPD Sharing: Not stated. Countries: 1. Publications: 1.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.