Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

173

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

173 results for “Statistical analysis”

Learn how ShareScore rates datasets ↗
dryad36/100

Data from: Estimating transmission dynamics and serial interval of the first wave of COVID-19 infections under different control measures: A statistical analysis in Tunisia from February 29 to May 5, 2020

Open the record for dataset details and reuse information.

publicOct 2020View details →
dryad36/100

Data from: Statistical analysis of the presidential elections in Belarus in 2020

Open the record for dataset details and reuse information.

publicAug 2021View details →
dryad36/100

BayesW time-to-event analysis posterior outputs and summary statistics

Open the record for dataset details and reuse information.

publicMar 2021View details →
dryad36/100

Collision between biological process and statistical analysis revealed by mean-centering

Open the record for dataset details and reuse information.

publicSep 2020View details →
zenodo32/100

Prescribed Burn Statistical Analysis and R Code

<p>Paper: Diversity and composition of fungal soil communities across prescribed burn areas in temperate hardwood forests</p> <p>Authors: S.D. Russell &amp; M.C. Aime</p> <p>All of the R code for the paper and the supplemental materials required to run it. This includes the BIOM file and the metadata file. The analysis is separated by the total fungal community and a separate analysis (and R file) for the ECM analysis.</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Input data files for RSS-NET analysis of IBD GWAS summary statistics and NK cell regulatory network

<p>Details of these data files are provided in https://suwonglab.github.io/rss-net/ibd2015_nkcell.</p> <p>Contact:<code> xiangzhu[at]psu.edu </code></p>

opencc-by-4.0Mar 2020View details →
zenodo32/100

A Statistical Analysis of Error in MPI Reduction Operations Dataset

<p>This is the dataset used to generate the figures and information contained in the paper <em>A Statistical Analysis of Error in MPI Reduction Operations </em>by Samuel D. Pollard and Boyana Norris, to appear in IEEE&#39;s<em> Fourth International Workshop on Software Correctness for HPC Applications</em>, 2020.<br> <br> A description is provided in the README.md as well as the software dependencies required to re-generate these data. The log files and tab-separated-values files (tsv) allows a user to analyze the same data we used for the paper. The file datasets-pollard-correctness2020.tar.bz2 decompresses to about 5.7GB.</p> <p>The source code used to generate these data is available <a href="https://github.com/sampollard/reduce-error">on Github</a>.</p>

openNov 2020View details →
zenodo32/100

Data and statistical analysis scripts for manuscript on high-throughput phenotyping of root economics in wheat

<p>This repository contains the root data, GEMMA files, and R code for generating the statistics and figures used in a preprint describing high-throughput phenotyping of root respiration in winter wheat.</p> <p><strong>Functional phenomics and genetics of the root economics space in winter wheat using high-throughput phenotyping of respiration and architecture</strong></p> <p>Haichao&nbsp;Guo,&nbsp;Habtamu&nbsp;Ayalew,&nbsp;Anand&nbsp;Seethepalli,&nbsp;Kundan&nbsp;Dhakal,&nbsp;Marcus&nbsp;Griffiths,&nbsp;Xue-Feng&nbsp;Ma,&nbsp;Larry M.&nbsp;York</p> <p>bioRxiv&nbsp;2020.11.12.380238;&nbsp;doi:&nbsp;<a href="https://doi.org/10.1101/2020.11.12.380238">https://doi.org/10.1101/2020.11.12.380238</a></p> <p><strong>analysis.R -&nbsp;</strong>A script for processing the&nbsp;included 4 CSV<em>&nbsp;</em>files of collected data.&nbsp;It is intended to run directly in RStudio and will automatically set the working directory in that case. The GEMMA folder includes files as output from GEMMA for genetic analysis, as described in the methods section of the preprint. These files are required by the R code for making Manhattan plots and other output. It will automatically create an output folder and the output text and figure files.</p> <p>The protocol for the respiration measurements are available in a separate Zenodo repository:&nbsp;<a href="https://doi.org/10.5281/zenodo.4247873">https://doi.org/10.5281/zenodo.4247873</a></p> <p>Please cite this repository and the preprint if any data or R code is used in your work.</p>

opencc-by-4.0Nov 2020View details →
dryad32/100

Data from: Accounting for multiple comparisons in statistical analysis of the extensive bioassay data on glyphosate

<p>Glyphosate is a widely used herbicide worldwide. In 2015, the International Agency for Research on Cancer (IARC) reviewed glyphosate cancer bioassays and human studies, declared that the evidence for carcinogenicity of glyphosate is sufficient in experimental animals.  We analyzed ten glyphosate rodent bioassays, including those in which IARC found evidence of carcinogenicity, using a meta-analytic procedure that adjusts for the large number of tumors eligible for statistical testing and provides valid false-positive probabilities.  The test statistics for these global tests are functions of p-values from a standard test for dose-response trends applied to each specific type of tumor.  We evaluated three global tests, using as test statistics the smallest p-value from a standard statistical test for dose-response trend and the number of such tests for which the p-value is less than or equal to 0.05 or 0.01.  The false-positive probabilities obtained from two implementations of these three global tests are: smallest p-value: 0.26, 0.17, p-values ≤ 0.05: 0.08, 0.12, p-values ≤ 0.01: 0.06, 0.08.  In addition, we found more evidence for negative dose-response trends than positive.  Thus, we found no strong evidence that glyphosate is an animal carcinogen.  The main cause for the discrepancy between IARC's finding and ours appears to be that IARC did not account for the large number of statistical tests performed in the bioassays they reviewed and the resulting multiple comparison problem.  This work provides a more comprehensive analysis of the animal carcinogenicity data for this important herbicide than previously available.</p>

opencc-zeroMar 2020View details →
dryad32/100

Data from: Analysis of statistical correlations between properties of adaptive walks in fitness landscapes

The fitness landscape metaphor has been central in our way of thinking about adaptation. In this scenario, adaptive walks are an idealized dynamics that mimics the uphill movement of an evolving population towards a fitness peak of the landscape. Recent works in experimental evolution have demonstrated that the constraints imposed by epistasis are responsible for reducing the number of accessible mutational pathways towards fitness peaks. Here we exhaustively analyze the statistical properties of adaptive walks for two empirical fitness landscapes and for theoretical NK landscapes. Some general scenario can be drawn from our simulation study. Regardless the dynamics, we observe that the shortest paths are more regularly used. Although the accessibility of a given fitness peak is reasonably correlated to the number of monotonic pathways towards it, the two quantities are not exactly proportional. A negative correlation predictability and mean path divergence is established, and so with the decrease of the number of effective mutational pathways ensues the convergence of the attraction basin of fitness peaks. On the other hand, other features are not conserved among fitness landscapes, such as the relationship between accessibility and predictability.

opencc-zeroJan 2020View details →
zenodo32/100

Comparative analysis of statistical methods used for detecting differential expression in label-free mass spectrometry proteomics - Data Supplement

<p>This the is Data Supplement for the article &quot;Comparative analysis of statistical methods used for detecting differential expression in label-free mass spectrometry proteomics&quot; submitted to the Journal of Proteomics 2015.</p>

opencc-zeroJun 2015View details →
zenodo32/100

Final geometries and energies, statistical analysis and estimated errors of single metals and bimetallics for CO2 to methanol conversion

<p>The dataset accommodate all the extra data discussed in:<br>Pisal, P., Krejč&iacute;, O. &amp; Rinke, P. Machine learning accelerated descriptor design for catalyst discovery in CO<sub>2</sub> to methanol conversion. <em>npj Comput Mater</em> <strong>11</strong>, 213 (2025). https://doi.org/10.1038/s41524-025-01664-9&nbsp;</p> <p>The datased contains four types of data:</p> <ol> <li>All the final geometries and energies of adsorbated (*H, *O, *OCHO &amp; *OCH3) and all the 158 single metals and bimetallic alloys on all the surfaces with Miller indices in {-2, -1, ... 2} optimized with Open Catalyst Project (OCP) 20 <em>equiformer_V2</em> machine-learned force-field model. These are in the <a href="https://zenodo.org/api/records/15587232/draft/files/geometries_and_energies.zip/content" target="_blank" rel="noopener noreferrer">geometries_and_energies.zip</a> file organized by the metal/alloys name, with the final geometries and enerigies in a json file, using a json ASE format.</li> <li>All the estimated mean absolute errors (MAE) of predicted adsorption energies for all the considered metals and bimetallic alloys in&nbsp;<a href="https://zenodo.org/api/records/15587232/draft/files/Estimated_MAEs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Estimated_MAEs_metals_bimetallics.csv</a> and xlsx file. The data content is identical, files differs only by a format.</li> <li>All the adsorption energy disctibutions (AEDs) for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> <li>All the statistical information of the adsorption energies for all the 158 metals/alloys and adsorbates in <span><a href="https://zenodo.org/api/records/15587232/draft/files/Statistics_AEDs_metals_bimetallics.csv/content" target="_blank" rel="noopener noreferrer">Statistics_AEDs_metals_bimetallics.csv</a></span> and xlsx files. The data content is identical, files differs only by a format.</li> </ol>

opencc-by-4.0Aug 2024View details →
zenodo32/100

Statistical Analysis of Feature-based Molecular Networking Results from Non-Targeted Metabolomics Data

<p>This folder contains the following used for the publication:</p><ul><li>MASSIVE Repositories: MSV000082312 and MSV000085786. This contains the original data in both .raw and .mzxml formats.</li><li>MZmine 3 files: The feature table (SD_BeachSurvey_GapFilled_quant.csv), the associated mgf file, the batch file (.xml) used for MZmine 3 to obtain the feature table, the mgf file for SIRIUS annotations (SD_BeachSurvey_SIRIUS_fixed.mgf)</li><li>SIRIUS and CANOPUS summary files (.tsv files)</li><li>FBMN Result files</li></ul>

opencc-by-4.0Oct 2023View details →
zenodo32/100

FIGURE 2. Cytochrome b maximum likelihood phylogram for the genus Carollia. Support statistics from a maximum likelihood bootstrap analysis and a in On the phylogenetic position of Carollia manu Pacheco et al., 2004 (Chiroptera: Phyllostomidae: Carolliinae)

FIGURE 2. Cytochrome b maximum likelihood phylogram for the genus Carollia. Support statistics from a maximum likelihood bootstrap analysis and a Bayesian analysis are indicated at each resolved node. For the maximum likelihood analysis (ML), white indicates bootstrap frequencies ≤ 50%, grey indicates bootstrap frequencies between 50% and 75%, and black indicates bootstrap frequencies ≥ 75%. For the Bayesian analysis (BPP), white indicates posterior probabilities &lt;0.95, whereas black indicates posterior probabilities ≥ 0.95.

opennotspecifiedOct 2013View details →
dryad32/100

cDNA sequence of E2 gene family in Arabidopsis thaliana and data of statistical analysis

<p>E2 ubiquitin-conjugating enzymes act as a heart role in the ubiquitination process and are responsible for catalysis ubiquitin transfer. Although the function of ubiquitin-protein ligases (E3s) in plant response to diverse abiotic stress by targeting specific substrates has been well studied, the E2s' involvement in environmental responses and their downstream targets are not well understood. Here, we demonstrated that the E2 ubiquitin-conjugating enzyme 18 (UBC18) regulates the stability of FREE1 to modulate iron deficiency stress. UBC18 affects the ubiquitination of FREE1 and promotes its degradation, overexpression of<em> UBC18</em> in plants decreases their sensitivity to iron deficiency by reducing the level of FREE1, and high accumulation of FREE1 in<em> </em>the<em> ubc18</em> mutant resulted in sensitivity to iron deficiency. In addition, we demonstrated the lysine residues K227, K295, K315, and K540 are required for FREE1 ubiquitination and stability regulation, and mutation of these lysines of FREE1 residues resulted in sensitivity to iron starvation in plants. Taken together, our findings reveal a mechanism of UBC18 in response to iron deficiency stress by altering the abundance of FREE1, and further elucidate the role of ubiquitination sites in FREE1 stability regulation and the plant iron deficiency response.</p>

opencc-zeroJan 2024View details →
zenodo32/100

Understanding Citywalk Through Social Media: A Spatial-Statistical Analysis in Shanghai

<p>The uploaded file contains the original Citywalk social media posts, coordinates and classifications of different Citywalk POIs, sentiment scores, <span>Shannon Diversity Index</span> for different grids, public transport statistics, public transport accessibility statistics, green area statistics, and the Grid Citywalk Activity Index.</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".

<p>This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.<br><br><strong>Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Figure 3.</strong> TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (<em>n</em> = 61,838).<br><br><strong>Extended Data Figure 1.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (<em>n</em> = 21,731).<br><br><strong>Extended Data Figure 2.</strong> Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (<em>n</em> = 9,380).</p>

opencc-by-4.0Nov 2024View details →
zenodo32/100

Dataset and statistical analysis related to "Amphibian studies to investigate the endocrine disrupting properties of chemicals through the Thyroid modality: a comparison of their statistical power"

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

FIGURE. Phylogenetic analysis of Chrysosporium spp. based on ITS sequences. Statistical support values (≥50 %) are shown at nodes, and presented as ML bootstrap support/Bayesian posterior probabilities. Names in black bold are the strains isolated in this study, the coloured names are the new species. in Morphological and phylogenetic characterisations reveal nine new species of Chrysosporium (Onygenaceae, Onygenales) in China

FIGURE. Phylogenetic analysis of Chrysosporium spp. based on ITS sequences. Statistical support values (≥50 %) are shown at nodes, and presented as ML bootstrap support/Bayesian posterior probabilities. Names in black bold are the strains isolated in this study, the coloured names are the new species.

opennotspecifiedMar 2022View details →
zenodo32/100

A Uniform Retrieval Analysis of Ultra-cool Dwarfs. IV. A Statistical Census from 50 Late-T Dwarfs

<p>Posterior distributions for all model runs of A Uniform Retrieval Analysis of Ultra-cool Dwarfs. IV. A Statistical Census from 50 Late-T Dwarfs</p> <p>To view, simply download the file and extract the zipfile.</p> <p>Directory structure is as follows:</p> <p><strong>BEST_FIT_SPECTRA</strong>: Figures for best-fit model spectra for all 50 objects.</p> <p><strong>CLOUD_OPTICAL_DEPTH: </strong>Derived tau_cloud optical depth posteriors from the cloud model.</p> <p><strong>CORNER_PLOTS</strong>: Corner plots for all 50 objects.</p> <p><strong>TEMPERATURE_PROFILES</strong>: Retrieved temperature profiles for all 50 objects. Overlaid are relevant condensation curves (legend found in the text of the publication).</p>

opencc-by-4.0May 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record