Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

31

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

31 results for “Corpus Analysis”

Learn how ShareScore rates datasets ↗
zenodo48/100

Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text

<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>

opencc-by-4.0May 2020View details →
zenodo44/100

Dataset do DH2020 [The Lusophone Digital Humanities and What they (we) are doing from the South: textual corpus analysis and FAIR principles to tackle Hegemony]

<p>Planilha de dados recuperados do Google Scholar utilizado na an&aacute;lise e apresenta&ccedil;&atilde;o da pesquisa emp&iacute;rica intitulada - <strong>The Lusophone Digital Humanities and What they (we) are doing from the South: textual corpus analysis and FAIR principles to tackle Hegemony </strong>- no evento <strong>DH2020 Ottawa</strong>: <a href="https://hcommons.org/deposits/item/hc:32051/">https://hcommons.org/deposits/item/hc:32051/</a></p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria&mdash;Hausa, Igbo, Nigerian-Pidgin, and Yor&ugrave;b&aacute;&mdash;consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Corpus de revistas Proyecto Digitization and Analysis of Cultural Transfers in Colombian Literary Magazines (1892–1950)

<p>Coprpus de revistas Proyecto Digitization and Analysis of Cultural Transfers in Colombian Literary Magazines (1892&ndash;1950)</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

The COUGHVID crowdsourcing dataset: A corpus for the study of large-scale cough analysis algorithms

<p><strong>Overview</strong></p> <p>Cough audio signal classification has been successfully used to diagnose a variety of respiratory conditions, and there has been significant interest in leveraging Machine Learning (ML) to provide widespread COVID-19 screening. The COUGHVID dataset provides over 30,000 crowdsourced cough recordings representing a wide range of subject ages, genders, geographic locations, and COVID-19 statuses. Furthermore, experienced pulmonologists labeled more than 2,000 recordings to diagnose medical abnormalities present in the coughs, thereby contributing one of the largest expert-labeled cough datasets in existence that can be used for a plethora of cough audio classification tasks.&nbsp;As a result, the COUGHVID dataset contributes a wealth of cough recordings for training ML models to address the world&rsquo;s most urgent health crises.</p> <p><strong>Private Set and Testing Protocol</strong></p> <p>Researchers interested in testing their models on the private test dataset should contact us at coughvid@epfl.ch, briefly explaining the type of validation they wish&nbsp;to make, and their obtained results obtained through&nbsp;cross-validation with the public data. Then, access to the unlabeled recordings will be provided, and&nbsp;the researchers should&nbsp;send the predictions of their models on these recordings. Finally,&nbsp;the&nbsp;performance metrics of the predictions will be sent to the researchers. The private testing data is not included in any file within our Zenodo record, and it can only be accessed by contacting the COUGHVID team at the aforementioned e-mail address.</p> <p><strong>New Semi-Supervised Labeling</strong></p> <p>The third version of the COUGHVID dataset contains thousands of additional recordings obtained through October 2021. Additionally, the recordings containing coughs were re-labeled according to a semi-supervised learning algorithm that combined the user labels with those of the expert physicians, which were&nbsp;modeled using ML and expanded on the previously unlabeled data. These labels can be found in the &quot;status_SSL&quot; column of the &quot;metadata_compiled.csv&quot; file.</p>

opencc-by-4.0Sep 2020View details →
zenodo44/100

English corpus from UniLeipzig, for keyboard layout analysis

<p>An English corpus and basic frequency analysis created from the UniLeipzig collection. Intended use is for keyboard layout analysis and development. The source sentences were selected for being typeable on a standard US ANSI keyboard, and artificially paragraphed in modern block style. Status is "beta". The analysis files may contain errors. The .csv files are known to work in LibreOffice, when imported into a spreadsheet using "tab" as a separator and NO string delimiter (delete the ' or " in the dropdown).</p> <p>Source files from https://wortschatz.uni-leipzig.de/en/download/English , their requested citation is&nbsp;</p> <p>D. Goldhahn, T. Eckart &amp; U. Quasthoff: Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages.<br>In: <em>Proceedings of the 8th International Language Resources and Evaluation (LREC'12), 2012</em></p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Oral cancer speech corpus for the paper "Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens"

<p>Dataset accompanying the paper &quot;<em>Objective speech outcomes after surgical treatment for oral cancer: An acoustic analysis of a spontaneous speech corpus containing 32.850 tokens</em>&quot;</p> <p>The zip file contains five folders:</p> <p>- <strong>Database:</strong> contains csv files for each speaker which contain the processed features</p> <p>- <strong>Recordings: </strong>the original recording from the YouTube Oral Cancer speech dataset, without further preprocessing</p> <p>- <strong>Recordings_Normalised:</strong> same as recordings but after minimal audio preprocessing (min-max scaling)</p> <p>- <strong>Textgrids: </strong>contains the textgrids which are annotated on the word-level and on phoneme-level</p> <p>- <strong>TIMIT selection: </strong>contains the textgrids for the TIMIT speakers. We unfortunately cannot share the audio date as it is not open source. More information can be found <a href="https://catalog.ldc.upenn.edu/LDC93s1">here.</a></p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)

<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>

opencc-by-sa-4.0Oct 2017View details →
zenodo40/100

Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights

<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook&nbsp;</a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images

<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>

opencc-by-4.0Oct 2018View details →
zenodo40/100

Language Function Analysis 2011 Corpus (LFA-11)

<p>The Language Function Analysis 2011 Corpus (LFA-11) is a German text corpus of promotional text, reviews and blog posts on music and smartphones. The texts were manually classified with respect to their topic relevance, language function, and sentiment polarity.</p> <p>The purpose of the corpus is to provide textual data for the development and evaluation of approaches to language function analysis and sentiment analysis. Therefore, each text is classified by language function (personal, commercial, or informational) as well as by sentiment (positive, negative, neutral).</p> <p>The corpus consists of two separated collections, which contain the texts about <em>music</em> and <em>smartphones</em> respectively. The music collection consists of 2,713 promotional texts and reviews from both users and professionals. The smartphone collection contains 2,093 blog posts on smartphones from the <a href="http://spinn3r.com/">Spinn3r corpus</a>.</p>

opencc-by-4.0Nov 2011View details →
zenodo36/100

Corpus Linguistic Analysis of the BMSatire Descriptions corpus

<p>Data and documentation related to the analysis of the BMSatire Descriptions corpus, produced as part of the British Academy funded project &#39;<a href="https://curatorialvoice.github.io/">Curatorial Voice: legacy descriptions of art objects and their contemporary uses</a>&#39;.</p> <p>All data are derived from text written by M. Dorothy George and published between 1935 and 1954 as volumes 5 to 11 of the&nbsp;<em><a href="https://en.wikipedia.org/wiki/Catalogue_of_Political_and_Personal_Satires_Preserved_in_the_Department_of_Prints_and_Drawings_in_the_British_Museum">Catalogue of Political and Personal Satires Preserved in the Department of Prints and Drawings in the British Museum</a></em>. This text is published in lightly edited form by the British Museum via ResearchSpace as linked open data at&nbsp;<a href="https://public.researchspace.org/sparql">https://public.researchspace.org/sparql</a>. The data, text and images available via this service are published under a&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)</a>&nbsp;license (Research Space,&nbsp;<a href="https://public.researchspace.org/resource/Termsofuse">2016</a>; accessed&nbsp;<a href="https://www.zotero.org/jwbaker/items/itemKey/9LY9J5XY">10 September 2018</a>).</p> <p>Code is licensed under a&nbsp;<a href="https://github.com/CuratorialVoice/code/blob/master/LICENSE">GNU General Public License v3.0</a>.</p>

opencc-by-nc-sa-4.0Jun 2019View details →
zenodo36/100

Paderborn Genre Analysis Corpus 2012 (PaGA-12)

<p>The Paderborn Genre Analysis 2012 corpus (PaGA-12) contains 1,639 HTML documents of 26 genres. All documents were collected from 2009-10-18 to 2009-11-20, and each document is manually assigned to exactly one genre. For each genre, the corpus provides at least 50 documents.</p> <p>All HTML documents contain German text only, and framesets are removed. The corpus is delivered in form of a MySQL database dump; the database structure is detailed in a README file delivered with the corpus.</p>

opencc-by-4.0Dec 2011View details →
zenodo36/100

Data sets of the analysis of 1sg and 2sg subject expression with the verbs 'creer' and 'saber' in a corpus of spoken Spanish

<p>These are the two data sets used for the quantitative analysis in the paper &quot;&ldquo;Perspectival factors and <em>pro</em>-drop &ndash; A corpus study of speaker/addressee pronouns with <em>creer </em>&lsquo;think/believe&rsquo; and <em>saber </em>&lsquo;know&rsquo; in spoken Spanish&rdquo; (Peter Herbeck;<em> to appear </em>in <em>Glossa &ndash; a journal of general linguistics</em>). It contains the&nbsp;values of the annotation of null and overt speaker/addressee pronouns in the Madrid and Alcal&aacute; samples of the corpus PRESEEA (2014-). The first data set was used for the quantitative analysis of subject expression (null/overt) with the verbs <em>creer&nbsp;</em>and <em>saber&nbsp;</em>according to person (1sg vs. 2sg) and polarity (negative vs. positive verb forms). The second data set was used for the analysis of 1sg subject expression according to the complement type of <em>creer&nbsp;</em>and <em>saber</em>.</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

Corpora, Corpus Analysis, Resources and Tools

<p>During the fourth Project Presentation Session on <strong>Tuesday 24.07.2018</strong> the following 3 projects were presented:</p> <ul> <li><strong>Yu-Hua Chen</strong> (University of Nottingham Ningbo China, China): &quot;Beyond Borders, Beyond Words: Issues &amp; Challenges in Developing An Open-Access Multimodal Corpus of L2 Academic English from An EMI University in China&quot;</li> <li><strong>Theodora Konach</strong> (Jagiellonian University,Krakow, Poland): &quot;Towards an Ethical Framework for the Digitalisation of Intangible Cultural Heritage in Museums - Digital Humanities in the Comparative Intellectual Property Law&quot;</li> <li><strong>Marco Passarotti</strong> (Universit&agrave; Cattolica del Sacro Cuore, Milan, Italy): &quot;What&#39;s Going On in Milan? A Practical Introduction to Resources and Tools for Latin at the CIRCSE Research Centre&quot;</li> </ul>

opencc-by-4.0Jul 2018View details →
zenodo32/100

Data accompanying the dissertation "Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)" (part of data-nh)

<p>This dataset includes research data that resulted from the analysis of the bibliography Bib-ACM&eacute; (see https://github.com/cligs/bibacme) and the corpus Conha19 (see https://github.com/cligs/conha19) as part of the work on the dissertation &quot;Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)&quot; by Ulrike Henny-Krahmer. Both the bibliography and the corpus include novels from Argentina, Cuba, and Mexico published between 1830 and 1910 which were analyzed by subgenre on the levels of metadata and text (see the dissertation for details).</p> <p>The dataset is part of &quot;data-nh&quot; (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.</p>

openother-pdJan 2021View details →
zenodo32/100

BHAAV (भाव) - A Text Corpus for Emotion Analysis from Hindi Stories

<p>The first and largest Hindi text corpus, named BHAAV (भाव), which means emotions&nbsp;in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus consists of 20,304 sentences collected from 230 different short stories spanning across 18 genres such as प्रेरणादायक (Inspirational) and रहस्यमयी (Mystery). Each sentence has been annotated into one of the five emotion categories anger, joy, suspense, sad, and neutral)&nbsp;by three native Hindi speakers with at least ten years of formal education in Hindi.</p>

opencc-by-4.0Sep 2019View details →
zenodo28/100

Development of the Quranic WAY Metaphor in Tafsir: A Corpus Analysis Approach

<p>This dataset contains the necessary data and scripts to recreate all steps and operations described in "Development of the Quranic WAY Metaphor in Tafsir: A Corpus Analysis Approach" by Adrian Bernhard. It also provides Docker configuration files and a database dump, allowing for quick recreation of the research infrastructure on which the study is based. Simply adapt the Docker configuration to include the database dump during setup.&nbsp;</p>

openMay 2024View details →
geo24/100

Quantitative Analysis of PPARD RNA-Seq Transcriptomes of Mouse Gastric Corpus Epithenial Cells by Next Generation Sequencing (NGS)

GEO Series GSE111050. Mus musculus. 28 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJan 2019View details →
geo24/100

Transcriptional analysis reveals the molecular details of atrophic gastritis in Helicobacter pylori infected corpus mucosa

GEO Series GSE27411. Homo sapiens. 18 samples. Type: Expression profiling by array.

openGEO-OpenFeb 2014View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record