Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,243
datasets available to search
ShareScore release 0.7.1
Dataset results
1,243 results for “Statistics”
Root multiple ion uptake kinetics data for maize NAM founders, statistical code, and RhizoFlux hardware plans
<p>This repository contains tabular data, R statistical code, protocols, and hardware plans associated with the following manuscript:</p> <p><strong>A multiple ion-uptake phenotyping platform reveals shared mechanisms that affect nutrient uptake by maize roots</strong></p> <p>Marcus Griffiths, Sonali Roy, Haichao Guo, Anand Seethepalli, David Huhman, Yaxin Ge, Robert E. Sharp, Felix B. Fritschi, Larry M. York</p> <p>Plant Physiology; doi: <a href="https://doi.org/10.1093/plphys/kiaa080">https://doi.org/10.1093/plphys/kiaa080</a></p> <p><strong>Equipment designs.zip</strong> - Contains the hardware plans, parts lists, and experimental protocol</p> <p><strong>ImageJ_macro.zip</strong> - Contains scripts to use within ImageJ to segment images to calculate leaf area</p> <p><strong>Supplementary_Data.zip</strong> - Contains the actual supplemental figures and tables for the manuscript as well as RNAseq data</p> <p><strong>R code & raw data.zip</strong> - Contains a single .R text file containing all the R code to generate all the figures and and supplemental figures from the include raw data files</p> <p>E-mail mgriffiths at danforthcenter.org or lmyork at noble.org with any questions.</p> <p>Version 1 was used for the preprint.</p> <p>Version 2 was used for the final submitted manuscript.</p> <p>Version 3 is the final published version.</p>
Test Collection Reliability: A Study of Bias and Robustness to Statistical Assumptions via Stochastic Simulation
<p>This archive contains the simulated collections, their diagnosis data, and the estimates of accuracy. For the full code and description, please refer to https://github.com/julian-urbano/irj2015-reliability</p>
Additional Data of O.Buß and S.Jager: "Statistical Evaluation of HTS Assays for Enzymatic Hydrolysis of ß-Keto Esters")
<p>The additional data set includes:</p> <ul> <li>Pdb structure files, which were used for calculations/pictures.</li> <li>R-scripts for calculation of statistical measures like Z-factor, SSMD and others.</li> <li>Data of different activity assays, which were used for R-script calculations.</li> <li>a SDS-PAGE (gel with all used enzymes)</li> </ul> <p>View our publication under: http://www.plosone.org/article/info%3Adoi%2F10.1371%2Fjournal.pone.0146104</p>
Classifying hot water chemistry: Application of MULTIVARIATE STATISTICS - Dataset
<p>These files are the dataset for the following paper "Classifying hot water chemistry: Application of MULTIVARIATE STATISTICS". Authors: Prihadi Sumintadireja<sup>1</sup>, Dasapta Erwin Irawan<sup>1</sup>, Yuano Rezky<sup>2</sup>, Prana Ugiana Gio<sup>3, </sup>Anggita Agustin<sup>1</sup></p>
Statistical Data from Muenster (2010-2014) as RDF
<p>This dataset is a <strong>curated statistical dataset </strong>from the Muenster City Council in RDF (Resource Description Framework) format. The City Council of Muenster has provided statistical datasets about Muenster covering the period 2010-2014 in PDF (Portable Document Format). Since PDF is not machine readable, students from the University of Muenster have tried to convert this statistical data into RDF, making therefore the data consumable by machines. The dataset covers five different topics:</p> <ul> <li>Unemployment in Muenster</li> <li>Population in Muenster</li> <li>Migration in Muenster</li> <li>Households of Muenster</li> <li>Employees subject to social insurance in Muenster</li> </ul> <p>As proofs of the usefulness of the RDF data, the students built some nice visualizations. The visualizations can be accessed from:</p> <ul> <li>https://git.io/vD547 (unemployment)</li> <li>https://git.io/vD545 (population)</li> <li>https://git.io/vD54d (migration)</li> <li>https://git.io/vD54b (households)</li> <li>https://git.io/vD5Bv (employees subject to social insurance)</li> </ul>
Raw data used for COI delineation of the Eupolybothrus species: Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar from: Eupolybothrus cavernicolus Komerički & Stoev sp. n. (Chilopoda: Lithobiomorpha: Lithobiidae): the first eukaryotic species description combining transcriptomic, DNA barcoding and micro-CT imaging data - Biodiversity Data Journal 1: e1013 (28 October 2013) https://doi.org/10.3897/BDJ.1.e1013
<p>Authors: Stoev et al. 2013 Data type: genomic The archive contains the following data: 1) fasta-Alignment as the basis for all analyses (.FASTA), 2) mega-file for the calculation of the genetic distances and the NJ tree (.MDSX), 3) NJ-tree in Newick format (.NWK), 4) graph of the TCS Software for the Statistical Parsimony method (.GRAPH) File: E_cavernicolus.rar</p>
Data for analysis in "Towards optimal cosmological parameter recovery from compressed bispectrum statistics"
<p>Measures of the three point function extracted from a suite of simulations using 4 different estimators: namely, the bispectrum, modal estimator, integrated bispectrum, line correlation function. Also supplied are power spectrum measures across the same simulations. <br> <br> The measures are done across 200 fiducial and 60 non-fiducial cosmology simulations, at 3 redshifts. Further details on what was done can be attained by reading the document, Overview.md/Overview.pdf, attached to the bundle. Even more details can be acquired by reading the paper this data was prepared for at https://arxiv.org/abs/1705.04392! </p>
Statistical Data of Vienna's population
<p>Statistical analysis of the population of Vienna by districts from 2001 to 2015.</p> <p>Input Data is taken from https://www.data.gv.at/katalog/dataset/stadt-wien_viebevlkerungseit1869wien/resource/f55512e6-81ef-4fa9-a01f-19f9c5f838c2.</p> <p>The source code for experiment which produces the data provided here can be found on https://bitbucket.org/BerwanY/dp3</p> <p>The csv-files of the districts population are containing the years as labels and the population for every year in the first row. The overview of vienna's population csv-file contains as first label 'YEAR', where the other label are the districts. So every row contains the population data of every district for one year.</p> <p>The csv-files are named as follows:</p> <ul> <li>For districts: 'district_<em>DISTRICTNUMBER</em>.csv'</li> <li>Overview: 'district_all.csv'</li> </ul> <p>, where DISTRICTNUMBER is replaced with the corresponding code of the district.</p>
Full Summary Statistics - TBDAR Genome-to-genome Study
<p><strong>TBDAR_G2G_Full_Summary_Stats.tar.gz: </strong>Full summary statistics (See README for details)</p> <p><strong>Mtb_Human_IDs.txt: </strong>Mapping between M.tb and human sequencing IDs, to faciliate joint analyses. </p> <p><strong>Supple_Data1.csv</strong>: <span lang="EN">G2G associations that meet the significance threshold of 5 × 10⁻⁸ </span></p>
Milo csQTL mapping Summary Statistics
<p>Summary statistics from Milo cell state quantitative trait loci mapping analysis on human peripheral immune cells from <a>Randolph <i>et al.</i></a> Each folder contains csQTLs per chromosome. Files are names 'chromosome_position_minorallele_majorallele_locus.tsv.gz'<i>. </i>Data are block gzip compressed and tabix indexed for rapid access based on SNP position (GRCh38 coordinates).</p>
A statistical shape model of craniosynostosis patients and 100 model instances of each pathology
<p>This dataset is part of the publication "A statistical shape model for radiation-free assessment and classification of craniosynostosis" (M. Schaufelberger et al.). It includes several 3D head models constructed of surface scans of craniosynostosis patients: The full shape model, a texture model, and submodels of four classes: sagittal suture fusion (scaphocephaly), metopic suture fusion (trigonocephaly), coronal suture fusion (brachycephaly and anterior plagiocephaly), and a control model (normocephaly and positional plagiocephaly). Each of the models is available in an .h5 file. We also include 100 mesh instances as a .ply file in a zip file. The model's statistical information can be incorporated into the [Liverpool-York child head model (Dai et al. 2019)](https://doi.org/10.1007/s11263-019-01260-7) as it uses the same vertex order and IDs (starting from index 0). If you want to synthesize new models, take a look a the demo.py file. For information about the hierarchy in the h5-file, take a look at documentation.md.</p>
Australian Statistical-Area (SA) Level Regions and Census Income Data (2011)
<p>The Australian Statistical Geography Standard (ASGS) defines a series of nested geographical areas in Australia known as Statistical Area (SA) Levels. SA3 regions are aggregations of SA2 regions, and SA2 regions are aggregations of SA1 regions. This data set contains the shapefiles of all SA1, SA2, and SA3 regions across Australia at the time of the 2011 census, originally downloaded from the Australian Bureau of Statistics (<a href="https://www.abs.gov.au/AUSSTATS/abs@.nsf/DetailsPage/1270.0.55.001July\%20201.">ABS</a>).</p><p>This data set also contains income information from the 2011 census, at the SA1 and SA2 level in New South Wales (NSW). Specifically, it contains the number of families of various types within a range of weekly income brackets.</p><p>Sainsbury-Dale et al. (2023) used a subset of this data set in a study on poverty levels in an area of (NSW) surrounding Sydney. </p><p> </p><p><strong>References</strong></p><p>Sainsbury-Dale, M., Zammit-Mangion, A., and Cressie, N. (2023) "Modelling Big, Heterogeneous, Non-Gaussian Spatial and Spatio-Temporal Data using FRK", <i>Journal of Statistical Software</i>, to appear.</p>
Dataset for A Statistical Survey of E-region Anomalous Electron Heating Using Poker Flat Incoherent Scatter Radar Observations
<p>This archive contains the complete list of anomalous electron heating (AEH) events in PFISR data between 2010 and 2023 identified by Zhang and Varney (2024), along with the code necessary to reproduce the results. The main list of AEH events is in the file AEH_event_list.csv, and the rest of this archive is supporting information for reproducibility.</p> <p>The files contained are:</p> <p>algo1.ipynb: Python notebook implementing algorithm 1.</p> <p>algo2.py: Python script implementing algorithm 2.</p> <p>algo3.ipynb: Python notebook implementing algorithm 3.</p> <p>algo4.ipynb: Python notebook implementing algorithm 4.</p> <p>cal_velo.py: Python function to calculate ion velocity.</p> <p>io_utils.py: Python functions for manipulating AMISR hdf5 files.</p> <p>Fig1.ipynb: Python notebook to recreate figure 1.</p> <p>Fig2,5.ipynb: Python notebook to recreate figures 2 and 5.</p> <p>Fig3,11.ipynb: Python notebook to recreate figures 3 and 11.</p> <p>Fig4.ipynb: Python notebook to recreate figure 4.</p> <p>Fig6.ipynb: Python notebook to recreate figure 6.</p> <p>Fig7,8,9,10.ipynb: Python notebook to recreate figures 7, 8, 9, and 10.</p> <p>PFISR_Data_Quality_Checker.ipynb: Python notebook with data preprocessing and quality checking.</p> <p>Table1.ipynb: Python notebook to extract the beamcode information needed for table 1.</p> <p>AEH_events_list.csv: Complete list of AEH events identified by algorithms 1, 3, and 4. The first column indicates the UT time of the start of the event, and 1 or 0 in the three columns denote whether the event was or was not detected by the algorithm, respectively.</p> <p>AEH_in_2010&2011.csv: Spreadsheet to facilitate direct comparisons with previous work on AEH in 2010 and 2011.</p> <p>f107.json: Smoothed F10.7 data used in this study.</p> <p>AEH_Detection_Outputs.zip: Archive of all of the raw output of the python scripts running the detection algorithms.</p> <p>AE&PAE.zip: Archive of all AE data used in this study.</p>
Data and Mathematical notebook for "Fractional-statistics-induced entanglement from Andreev-like tunneling"
<p>The uploaded files "SourceRightON_full.txt", "SourceLeftON_full.txt" and "BothSourcesON_full.txt" contain data for the work entitled "Fractional-statistics-induced entanglement from Andreev-like tunneling".</p> <p> </p> <p>The other file "New_Anyonic_data_fittings v2.nb" is the Mathematica notebook with which we perform the data analysis. When using it, please place three data files (mentioned above) in the Download folder.</p>
Full eQTL summary statistics for the PAUSE trial
<p>This dataset contains full eQTL summary statistics for the study 'Immunosuppression causes dynamic changes in expression QTLs in psoriatic skin' (<em>Nat.Comm., </em>2023). The study analyzes 375 skin samples from patients with psoriasis. For each SNP-gene pair, the dataset provides information listed below.</p> <ol> <li><strong>ID</strong>: the variant identifier, with chromosome number, position, ref and alt alleles.</li> <li><strong>Estimate</strong>: estimate</li> <li><strong>Std.Error</strong>: standard error</li> <li><strong>df</strong>: degrees of freedom</li> <li><strong>t value</strong>: t-value of the association</li> <li><strong>Pr(>|t|): </strong>p-value of the association</li> <li><strong>Gene: </strong>Ensembl gene ID</li> <li><strong>CHROM: </strong>the variant chromosome</li> </ol>
Data for fitting a statistical global burned area model for seamless integration into Dynamic Global Vegetation Models
<p>The dataset is a large R data.table object saved in RDS format. It contains global, monthly data spanning the period from 2002 to 2018, with a 0.5 degrees spatial resolution. The dataset is utilized to develop and validate statistical models for predicting global burnt areas resulting from wildfires.</p>
Рис. 10. Блок-схема фиЗико-статистического прогноЗа уроЖайности спата приморского гребешка. Fig. 10. The block diagram of physical-statistical forecast of yield of spat of the Japanese scallop. in Review of methods for the forecast of mollusk's spat productivity in sea-farms of Primorye and probable ways of their enhancement
Рис. 10. Блок-схема фиЗико-статистического прогноЗа уроЖайности спата приморского гребешка. Fig. 10. The block diagram of physical-statistical forecast of yield of spat of the Japanese scallop.
Data For Survival of the Fittest: Testing Superradiance Termination with Simulated Binary Black Hole Statistics
<p>This repository is associated with the GitHub repository: https://github.com/jacquelynzhy/Statistical_Superradiance, which includes the code to reproduce the findings of Zhu et al. (2025). Specifically, the file "Output1.dat" here represents the output of running ZEVN with the initial conditions outlined in Section 3.1 of Zhu et al. (2025), which only included BH-BH binaries. For the values generated by a ZEVN run and instructions on how to select the type of remnants you are interested in, please refer to <a href="https://ui.adsabs.harvard.edu/abs/2019MNRAS.485..889S/abstract">Spera et al. (2019)</a> and the ZEVN GitHub page at: https://gitlab.com/sevncodes/sevn.</p>
Summary statistics from "Sex-Specific Causal Relations between Steroid Hormones and Obesity—A Mendelian Randomization Study"
<p>GWAMA summary statistics of four steroid hormone levels and one steroid hormone ratio using fixed-effect model.</p> <p>When using this data, please cite: Pott J, Horn K, Zeidler R, et al.. Sex-Specific Causal Relations between Steroid Hormones and Obesity - A Mendelian Randomization Study. <em>Metabolites</em> <strong>2021</strong>, <em>11</em>, 738. https://doi.org/10.3390/metabo11110738</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>ea (effect allele)</li> <li>oa (other allele)</li> <li>eaf (effect allele frequency)</li> <li>info (minimal info score across all used studies)</li> <li>nSamples (sample size per SNP)</li> <li>nStudies (number of studies)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>I2 (SNP heterogeneity across studies)</li> <li>phenotype (phenotyp setting)</li> </ul>
Summary statistics from "Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease"
<p>GWAMA summary statistics of four steroid hormone levels using fixed-effect model and GWAS summary statistics of four other steroid hormones.</p> <p>When using this data, please cite: Pott J, Bae YJ, Horn K, et al.. Genetic Association Study of Eight Steroid Hormones and Implications for Sexual Dimorphism of Coronary Artery Disease. <em>J Clin Endocrinol Metab</em> <strong>2019</strong> Nov 1;104(11):5008-5023. doi: 10.1210/jc.2019-00757</p> <p>All txt files contain the following columns:</p> <ul> <li>markername</li> <li>chr</li> <li>bp_hg19 (base position according to hg19)</li> <li>effect_allele</li> <li>other_allele</li> <li>effect_allele_freq</li> <li>min_info (minimal info score across all used studies)</li> <li>n (sample size per SNP)</li> <li>beta (effect estimate)</li> <li>se (standard error)</li> <li>p (p-value)</li> <li>CochransQ (only in GWAMA; SNP heterogeneity across studies)</li> <li>pCochransQ (only in GWAMA; p-value of Cochrans Q value)</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.