Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

31

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

31 results for “Corpus Analysis”

Learn how ShareScore rates datasets ↗
geo24/100

Transcriptome analysis of canine (Canis familiaris) corpus luteum (CL) during non-pregnant late luteal phase compared with normal luteolyzing CL collected at prepartum progesterone (P4) decrease and f

GEO Series GSE98657. Canis lupus familiaris. 15 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenNov 2017View details →
geo24/100

miRNA analysis in myofibroblats derived from normal stomach, both antrum (A) and corpus (C) separately and gastric cancers

GEO Series GSE76218. synthetic construct; Homo sapiens. 49 samples. Type: Non-coding RNA profiling by array.

openGEO-OpenFeb 2016View details →
zenodo24/100

Corpus analysis of the polysemy of the verb taper 'hit' in non-standard French using Twitter data

<p>This data package contains (i) an annotated corpus sample of 10 000 tweets with the verb taper &#39;hit&#39; in non-standard French (e.g. taper un coma, taper (dans) un kebab, taper une &eacute;quipe 3-0)&nbsp; (NON-FINAL VERSION) and (ii) the Python script used to retrieve the tweets.</p>

opencc-by-4.0Aug 2023View details →
geo20/100

Microarray analysis of the transcriptome in the primate corpus luteum during chorionic gonadotropin administration simulating early pregnancy.

GEO Series GSE25335. Macaca mulatta. 20 samples. Type: Expression profiling by array.

openGEO-OpenNov 2011View details →
geo20/100

RNA Sequencing Analysis of Wild Type Caput, Corpus, and Cauda Epididymal Transcriptomes

GEO Series GSE138517. Mus musculus. 9 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2020View details →
zenodo20/100

Revised primary and secondary analysis of the DIALLS corpus

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
geo16/100

Single-cell RNA-seq analysis of subventricular zone-derived neural progenitors mobilized to demyelinated corpus callosum

GEO Series GSE144201. Mus musculus. 1 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2021View details →
geo16/100

Single-cell RNA-seq analysis of microglial cells in demyelinated corpus callosum

GEO Series GSE144202. Mus musculus. 2 samples. Type: Expression profiling by high throughput sequencing.

openGEO-OpenJun 2021View details →
zenodo16/100

Corpus of 19th century Spanish American novels for family resemblance analysis (part of data-nh)

<p>This dataset contains different formats of a corpus of 19th century Spanish American novels which were used in a family resemblance analysis as a part of the dissertation &quot;Genre Analysis and Corpus Design: 19th Century Spanish American Novels (1830-1910)&quot; by Ulrike Henny-Krahmer. The texts are included as plain text files, linguistically annotated files (using TreeTagger), as text files with only noun lemmas, and as chunks of 1,000 noun tokens derived from the lemmatized texts. The texts were prepared in this way to be used with topic modeling. Because 22 of the novels still are under copyright, this dataset has restricted access. The other 234 novels are in the open domain. This dataset is part of &quot;data-nh&quot; (see https://github.com/cligs/data-nh), which is the whole collection of research data accompanying the above-mentioned dissertation.</p>

restrictedJan 2021View details →
zenodo16/100

Bibliometric Analysis of Exhibition and Trade Fair Research: A Corpus of 69 Articles and Cluster Insights from Journal of Convention and Event Tourism

<p>This dataset comprises a bibliometric analysis of 69 articles focused on the exhibition and trade fair sector, drawn from Scopus-indexed publications within the <em>Journal of Convention and Event Tourism</em>. The dataset includes detailed information on authorship, publication trends, and the clustering of research themes, providing valuable insights into the academic landscape of this niche within the MICE (Meetings, Incentives, Conferences, and Exhibitions) industry. This resource is essential for researchers and professionals seeking to understand the development and focus areas of scholarly work in exhibitions and trade fairs</p>

restrictedcc-by-4.0Aug 2024View details →
zenodo12/100

Multilingual fine-grained sentiment analysis corpus

<p>A sentiment annotated corpus based on Fallout New Vegas. The corpus has the following sentiments: <em>neutral, anger, disgust, fear, happy, pained, sad, surprised</em> in the following languages: <em>English, German, Italian, Spanish and French</em>.</p> <p>Please cite the following paper:&nbsp;Mika H&auml;m&auml;l&auml;inen, Khalid Alnajjar, and Thierry Poibeau. 2022. Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. In <em>FDG&rsquo;22: Proceedings of the 17th International Conference on the Foundations of Digital Games (FDG &rsquo;22)</em></p> <p>&nbsp;</p>

restrictedAug 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record