Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4,694

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

4,694 results for “data analysis”

Learn how ShareScore rates datasets ↗
zenodo36/100

Simulation Data from STORMI and SWMF Models for the May 2024 Geomagnetic Storm Analysis

<p>Simulation Data from STORMI and SWMF Models for the May 2024 Geomagnetic Storm Analysis.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

MRI images and code for data analysis

<p>Author: Samuele Ceolin</p> <p>Luxembourg Institute of Science and Technology, Environmental Research and Innovation Department</p> <p>Correspondence: samuele.ceolin94@gmail.com</p> <p>Copyright &copy; 2019-2024 Luxembourg Institute of Science and Technology. All Rights Reserved.</p> <p>The code in this repository is distributed under the terms of the <a href="https://www.gnu.org/licenses/gpl-3.0.html" target="_blank" rel="nofollow noreferrer noopener">GNU General Public License 3.0</a> or any later version, unless stated otherwise. All other material is licensed under a <a href="https://creativecommons.org/licenses/by/4.0/" target="_blank" rel="nofollow noreferrer noopener">Creative Commons Attribution 4.0 International license</a>, unless stated otherwise.</p> <p><strong>General purpose</strong></p> <p>Repository containing the MRI data and images (MRI_data.zip) and containing the code for the extrapolation of root properties and data analysis (MRI_code.zip).</p> <p>See the file README.md contained in "MRI_code.zip" for more information.</p>

opencc-by-4.0Feb 2024View details →
zenodo36/100

Transcriptomic Analysis Data for MSTN Mutations and Mechanisms of Muscle Hypertrophy in a New Guinea Pig Breed

<p>This dataset contains raw RNA-seq data from six guinea pig muscle samples, split into two groups:</p> <ul> <li><strong>Native guinea pigs (B1 to B3):</strong> Control group with no selective breeding.</li> <li><strong>Kuri breed guinea pigs (B4 to B6):</strong> Synthetic hybrid group selectively bred for increased muscle mass.<br>Each sample has paired-end FASTQ files (e.g., B1_1.fq.gz and B1_2.fq.gz).</li> </ul>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Data and Software for "Probabilistic Trade-offs Analysis for Sustainable and Equitable Management of Climate-Induced Water Risks"

<p><span>Research data supporting the study "Probabilistic trade-offs analysis for sustainable and equitable management of climate-induced water risks"</span></p> <p><span>This repository provides data of the Stochastic Dual Dynamic Programming (SDDP) model, and the output results of the simulations of the various policies and climate scenarios considered in this study, as well as the code used for postprocessing and visualizing the results.</span></p> <p><strong><span>Contents</span></strong></p> <ol> <li><strong><span>Data: Model Inputs</span></strong><span><br>This folder contains the physical river network, reservoir and water demand, and economic data derived from the observed database.<br>The key files are:</span></li> <ul> <li><span>Input_HydrologicalData</span></li> <li><span>Input_SystemData</span></li> </ul> <li><strong><span>Results: Model Output Analysis</span></strong><span><br>This folder includes outputs from the Stochastic Dual Dynamic Programming (SDDP) model under various policies and climate scenarios. The results showcase optimized sectoral water use, including irrigated areas, hydropower generation, and allocations for agriculture, energy, and urban demands across spatial locations (upstream and downstream).<br>Key files include:</span></li> <ul> <li><strong><span>SDDP Model Outputs</span></strong><span> (MATLAB format): </span></li> <ul> <li><span>EnergyPriority_Baseline.mat</span></li> <li><span>EnergyPriority_2070.mat</span></li> <li><span>EnergyPriority_2100.mat</span></li> <li><span>AgriculturePriority_Baseline.mat</span></li> <li><span>AgriculturePriority_2070.mat</span></li> <li><span>AgriculturePriority_2100.mat</span></li> </ul> <li><strong><span>Extracted Model Results</span></strong><span> (Excel format): </span></li> <ul> <li><span>Organized for each policy and climate scenario to facilitate analysis.</span></li> </ul> </ul> <li><strong><span>Software: Data Analysis and Visualization</span></strong><span><br>Python scripts designed for outputs data analysis and visualization are included to reproduce the primary figures from the study.<br>Scripts provided:</span></li> <ul> <li><span>CDF_outflow.py</span><span>: Analyzes cumulative distribution functions for river discharge.</span></li> <li><span>CDF_sectors.py</span><span>: Examines sectoral water use distributions.</span></li> <li><span>PCP_SI.py</span><span>: Generates Parallel Coordinate Plots for trade-offs analysis.</span></li> </ul> <li><strong><span>Instructions: README File</span></strong><span><br>A comprehensive README file explains:</span></li> <ul> <li><span>Details of model input data.</span></li> <li><span>Instructions to interpret the SDDP model outputs.</span></li> </ul> </ol> <p><strong><span>Instructions:</span></strong><span><br></span><span>The Python scripts process Excel files from the model output results folder to generate and visualize the figures for the paper. Each step is clearly documented within the scripts.</span></p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Carmichael_etal_Data_Analysis_All_Stress_Responses

<p>Data for the submission "Reconciling variability in multiple stressor effects using environmental performance curves". This includes growth rate data for 12 bacterial taxa exposed to gradients of temperature, pH and salinity.&nbsp;</p> <p>The R script for the data analysis pipeline is included, as well as a README file describing the data layout.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Associated code and data for "A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia (doi: 10.1007/978-1-0716-4290-0_18)".

<p>This deposit contains the data, code, and analysis to recreate the results in the manuscript - Tuulikki Suomela, Liang Zhang, Julio Vera, Heiko Bruns, Xin Lai. A Practical Guideline for MicroRNA Sequencing Data Analysis in Chronic Lymphocytic Leukemia. Methods Mol. Biol., 2883, 403&ndash;426. <a href="https://www.researchgate.net/publication/387267721_A_Practical_Guideline_for_MicroRNA_Sequencing_Data_Analysis_in_Chronic_Lymphocytic_Leukemia">https://doi.org/10.1007/978-1-0716-4290-0_18</a>.</p> <p>The pipeline allows users to perform end-to-end analysis of bulk miRNA sequencing data, including quality control of FastQ files, mapping of read counts to miRNA genes using miRBase or Reference genome, quantification of miRNA read counts, differential gene expression analysis using DEseq2, gene set enrichment analysis using curated cancer hallmark gene sets, and identification of miRNA targets.</p> <p>If you have used the code for your research, please cite the original publication. Thank you very much.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Processed datasets and codes for differential expression analysis on polulation-level RNA-seq data

<p>This version includes codes and data necessary to reproduce all results in our response to the correspondences ("Response to 'Neglecting normalization impact in semi‑synthetic RNA‑seq data simulation generates artificial false positives' and 'Winsorization greatly reduces false positives by popular differential expression methods when analyzing human population samples'") (<a href="https://doi.org/10.1186/s13059-024-03232-8">https://doi.org/10.1186/s13059-024-03232-8</a>).</p> <p>It also includes a README file to guide the reproduction of the results in our original publication and resources for the goodness of fit test in the original publication, "Exaggerated False Positives by Popular Differential Expression Methods When Analyzing Human Population Samples" (<a href="https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4">https://genomebiology.biomedcentral.com/articles/10.1186/s13059-022-02648-4</a>).</p>

opencc-by-4.0Apr 2004View details →
zenodo36/100

Lightning Declines Over Shipping Lanes Follow Regulation of Fuel Sulfur: Data Analysis

Open the record for dataset details and reuse information.

opencc-by-4.0May 2024View details →
zenodo36/100

Nearctic Salpinginae COI phylogenetic analysis and map data

<p>This dataset includes three files associated with the "Revision of the Nearctic Salpinginae (Coleoptera: Salpingidae) manuscript by Johnston and Pollock.</p> <p>&nbsp;</p> <p>The first file is a Nexus file for Mesquite which includes the final alignment of COI DNA barcodes and the results of their analysis under maximum likelihood.</p> <p>The second two files are an R script for the creation of maps from the associated csv file of point data for specimen records of Nearctic salpingines.</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Supporting data: Divergent Estimates of China's Forest Carbon Sink Can Not Be Well Coordinated based on a Collective Analysis

<p>Differences in the carbon sink results obtained from the inventory method, DGVMs, ATM, and EC were compensated using relevant data under a harmonized definition of the NetFCsink. All the carbon flux results used have been stored in an xlsx file, categorized according to the corresponding method.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Digital Twin or Digital Model: An Analysis of Definitions along the Product Lifecycle - Research data

<p>This research data contains the statements of the authors Grieves, Stark and Tao with regard to selected characteristics of Digital Twins. According to these statements different case studies along the product life cycle are classified as Digital Twin or Digital Model.</p> <p>Version 2 added a change in characteristic 2.</p>

opencc-by-nc-nd-4.0Oct 2024View details →
zenodo36/100

Data from Coral Dark Gene analysis

<p>Results files from manuscript "Cosmopolitan gene families with known functions are hotspots for the evolution of novel genes in stony corals".</p> <p>&nbsp;</p> <p><code>Orthogroups_analysis.tar.gz</code> HTML files displaying all expression, protein structure, and phylogeny results for each of the selected dark orthogroups.</p> <p><code>AlphaFold2.tar.gz</code> AlphaFold2 PDB files with the inferred structures of each of the selected dark orthogroup's representative sequences.</p> <p><code>OGs.tar.gz</code> Files listing proteins in each orthogroup, the annotations assigned to each protein, and the size and classification (restricted/shared, taxonomic rank) assigned to each orthogroup.</p> <p><code>Datasets.tar.gz</code> Files from each of the 129 coral and outgroup genome and transcriptome datasets used in this study. For transcriptomes this generally includes protein and CDS (if produced by authors) files. For genomes this generally includes protein, CDS (if produced by authors), assembly (if applicable), and gene gff3 (if applicable) files.</p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Analysis data for location- and scale-invariant power transformations

<p>This repository contains various files and folders related to the machine learning experiments in a forthcoming manuscript on location- and scale-invariant power transformations.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Mitigating autocorrelation during spatially resolved transcriptomics data analysis

<p>Here we include the marmoset brain and mouse gut STARmap data introduced in the corresponding manuscript, "Mitigating autocorrelation during spatially resolved transcriptomics data analysis". We also include the mouse brain STARmap PLUS data that was used to demonstrate cross-species spatial integration and was previously published in Shi, He, Zhou et al. 2022.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Documents used in the PLANET4B D1.1 analysis of biodiversity discourse by news outlets - 2010 and 2022 Data

<p>Data used to analyse biodiversity discourse in news outlet as part of Deliverable D1.1. of the Planet4B Project.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Data for article "Personal approach for cancer treatment: a meta-analysis of Phase II Clinical trials"

<p>We conducted systematic review and meta-analysis to provide a comprehensive overview of outcomes in patients who underwent personalized genomics-based versus non-personalized treatment in oncology. The PubMed searches detected 803 studies based on phase II clinical trials&rsquo; results published from 2010 to 2021. We selected 50 studies, having 81 arms and 6536 patients for the analysis. We compared Response Rate (RR), medians and 1-year rates of Overall Survival (OS) and Progression-Free Survival (PFS) between genomics-based personalized and non-personalized arms. This repository contains final dataset (Dataset file) and t<span lang="EN-US">he information on 803 studies identified in the literature search (803 studies description file)</span>.&nbsp;</p> <p>The searches, study selection, data extraction and synthesis were performed in accordance to PRISMA (preferred reporting items for systematic review and meta-analysis) guidelines. The research protocol was registered in PROSPERO (International prospective register of systematic reviews, <a href="https://www.crd.york.ac.uk/PROSPERO" rel="nofollow">https://www.crd.york.ac.uk/PROSPERO</a>), record ID CRD42024504021.&nbsp;</p> <p>We performed proportional meta-analysis using the RStudio program, utilizing the R programming language and packages "meta", "metafor," and "tidyverse", the code is available at github: https://github.com/MikhailPot/PreciseOnco_meta-analysis&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Reference Manager Data Citation Analysis

<p>DESCRIPTION:</p> <p>This package contains data used to analyze citation metadata completeness and correctness for several common reference managers used in scholarly research and several common repositories in the Earth, space, and environmental sciences.</p> <p>METHODS:</p> <p>Metadata fields for import and export methods and for 8 metadata fields (authors/creators, publisher, DOI, dataset title, version, access date, publication date, and resource type) were collected from reference managers via all import methods available (app or wizard and plugin) during summer 2024 from most recent software versions of all. To encode data, citation information for each dataset as imported by Reference Manager was compared to that registered for the DOI with DataCite. Correct metadata for each of 8 fields for both import and export was encoded as 0, incorrect as 1, and missing as '' or nan. See publication and software package for more information.</p> <p>FILES:</p> <p>FOLDER 'coded-data' contains files that include information (DOIs) about the data examined in this study, preserved copies of exported data citations used in the data interpretation and processing, and the processed data itself encoded in columns.</p> <p>FOLDER 'datacite-metadata-profiles' includes the raw metadata from each dataset DOI at the time of analysis, included for reproducibility purposes. &nbsp;</p> <p>FOLDER 'bibtex-files' includes the downloaded .bib files, where available, for each dataset DOI examined.</p> <p>See README file for more information.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

PoreScript: Semi-automated Pore Size Analysis Algorithm Data Set

<p>This data set contains files related to the PoreScript semi-automatic pore size image analysis algorithm. The three MATLAB files needed for the PoreScript algorithm&nbsp;are named the following:&nbsp;</p> <p>(<a href="https://zenodo.org/api/files/47f5a723-b2de-41b6-baed-f1a6335c4b84/Jenkins_RelativeIntensityFinder_no_crop.m">Jenkins_RelativeIntensityFinder_no_crop.m</a>,&nbsp;<a href="https://zenodo.org/api/files/47f5a723-b2de-41b6-baed-f1a6335c4b84/Jenkins_UserInterface_no_crop.m">Jenkins_UserInterface_no_crop.m</a>,&nbsp;<a href="https://zenodo.org/api/files/47f5a723-b2de-41b6-baed-f1a6335c4b84/Jenkins_PoreSizeCalculator_no_crop.m">Jenkins_PoreSizeCalculator_no_crop.m</a>).</p> <p>Access the latest version of the program here:<a href="https://github.com/djenkins95/PoreScript_Update_9_26_23"> <strong>https://github.com/djenkins95/PoreScript_Update_9_26_23</strong></a></p> <p>Updated MATLAB files are more accessible to a wider range of&nbsp;SEM software. The updated version&nbsp;asks for the known length of your scale bar in&nbsp;pixels. There are many ways to measure the length of your scale bar. I recommend using the free software FIJI. Use the *Straight*&nbsp;(drawing tool to trace your scale bar), then click Analyze &gt; Measure to determine the length in pixels.&nbsp;It should be noted that the length in pixels will be the same for any image taken on the same instrument, at the same magnification, and saved as the same file type (e.g., .tiff), so you can reference the length in future data sets without needed to remeasure the scale bar.</p> <p>The Zenodo&nbsp;repository includes the unanalyzed SEM images, analyzed images, raw pore size data, analyzed pore size data, and older .m versions.</p>

opencc-by-4.0Oct 2021View details →
zenodo36/100

LNT Model is not an "Assumption": Re-Analysis of Epidemiological Data Empirically Supports LNT

<p><em>Introduction:</em>&nbsp;In &ldquo;Keeping ICRP Recommendations Fit for Purpose&rdquo;[1], LNT model is described as &ldquo;LNT is the most appropriate evidence-based assumption to use for radiological protection purposes (p.10).&rdquo; According to our critical literature survey on radiological epidemiology[2], some limitations were identified:&nbsp; (1) aggregation of individual level data, (2) model formulation, (3) model estimation, (4) model selection, (5) results interpretation. In this paper we focus (4) model selection and demonstrate LNT was the best model.</p> <p><em>Data and Method:</em>&nbsp;Using &ldquo;Life Span Study Report 14. Cancer and non-cancer disease mortality data, 1950-2003 [3]&rdquo;, solid cancer mortality was re-analyzed. In addition to the L, Q, LQ, hadn searched threshold model, kinked&ndash;at-2Gy model that assumes LQ for less than 2Gy and L for larger than 2 Gy, and Linear model with threshold as a parameter were estimated. Following [3], Poisson regression model was applied and model fit was compared with AIC and BIC.</p> <p><em>Results</em>: Among estimated models, Linear (BIC=18317.9) and grid search threshold at 20mSv (BIC=18318.1) was selected as the best models. Directly estimated threshold was -23.2 mSv and it was statistically insignificant (z=-0.087,p&gt;0.1). Model fit of kinked-at-2Gy was poorer than these models (BIC=18321.2).</p> <p><em>Conclusion:</em>&nbsp;Based on these results, we can conclude LNT model is the best model for a-bomb survivor solid cancer mortality. According to our literature survey, LNT is supported Description of LNT model in &ldquo;Keeping ICRP Recommendations Fit for Purpose&rdquo; should be modified accordingly: &ldquo;LNT is the scientifically supported model, it is reasonable LNT to use for radiological protection purposes.&rdquo;</p>

opencc-by-2.0Nov 2021View details →
dryad36/100

Data and code to replicate: Diet analysis using generalized linear models derived from foraging processes using R package mvtweedie

<p>Diet analysis integrates a wide variety of visual, chemical and biological identification of prey.  Samples are often treated as compositional data, where each prey is analyzed as a continuous percentage of the total.  However, analyzing compositional data results in analytical challenges, e.g., highly parameterized models or prior transformation of data.  Here, we present a novel approximation involving a Tweedie generalized linear model (GLM).  We first review how this approximation emerges from considering predator foraging as a thinned and marked point process (with marks representing prey species and individual prey size).  This derivation can motivate future theoretical and applied developments.  We then provide a practical tutorial for the Tweedie GLM using new package <i>mvtweedie</i> that extends capabilities of widely used packages in R (<i>mgcv</i> and <i>ggplot2</i>) by transforming output to calculate prey compositions.  We demonstrate this approach and software using two examples. Tufted puffins (<i>Fratercula cirrhata</i>) provisioning their chicks on a colony in the northern Gulf of Alaska show decadal prey switching among sand lance and prowfish (1980-2000) and then Pacific herring and capelin (2000-2020), while wolves (<i>Canis lupus ligoni</i>) in Southeast Alaska forage on mountain goats and marmots in northern uplands and marine mammals in seaward island coastlines. </p>

opencc-zeroNov 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record