Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,243

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,243 results for “Statistics”

Learn how ShareScore rates datasets ↗
zenodo44/100

Summary statistics for "Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer's Disease"

<p>These are the burden test results (summary statistics) for the publication:</p> <p>&quot;Exome sequencing identifies rare damaging variants in ATP8B4 and ABCA1 as risk factors for Alzheimer&rsquo;s Disease&quot;,</p> <p>Nature Genetics, 2022.</p> <p>&nbsp;</p> <p><em>Format: tab-separated-value.</em></p> <p><em>Fields:</em></p> <ul> <li><em>gene_stable_id: Ensembl gene id</em></li> <li><em>gene_name: standard gene name</em></li> <li><em>pvalue: burden test significance (likelihood ratio test, population structure correction based on&nbsp;6 PCA components)</em></li> <li><em>cmac_all: sum of minor allele dosages across all contributing samples and variants</em></li> <li><em>group: variant group (LOF, LOF+REVEL&gt;=75, LOF+REVEL&gt;=50, LOF+REVEL&gt;=25, see publication methods for further selection criteria).</em></li> <li><em>beta/se: beta/se of logistic ordinal regression (see publication methods). Positive = risk-increasing. Negative = risk-decreasing.</em></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Statistical analysis for: Mode I fracture of beech-adhesive bondline at three different temperatures

<p>This dataset collects a raw dataset and a processed dataset derived from the raw dataset. There is a document containing the analytical code for statistical analysis of the processed dataset in .Rmd format and .html format.&nbsp;<br> <br> The study examined&nbsp;some aspects of mechanical performance of solid wood composites. We were interested in certain properties of solid wood composites made using different adhesives with different grain orientations at the bondline, then treated at different temperatures prior to testing.&nbsp;</p> <p>Performance was tested by assessing fracture energy and critical fracture energy, lap shear strength, and compression strength of the composites. This document concerns only the fracture properties, which are the focus of the related paper.&nbsp;</p> <p>Notes: &nbsp;&nbsp;</p> <p>* the raw data is provided in this upload, but the processing is not addressed here. &nbsp;&nbsp;<br> * the authors of this document are a subset of the authors of the related paper.<br> * this document and the related data files were uploaded at the time of submission for review. An update providing the doi of the related paper will be provided when it is available.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Multiscale continuum figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling"

<p>Accessible versions of selected figures from&nbsp;Tratnyek et al. (2017) &quot;In silico environmental chemical science: Properties and processes from statistical and computational modelling&quot; Environ. Sci. Processes Impacts 19(3): 188-202. DOI: 10.1039/C7EM00053G.</p> <p>The Abstract Art figure shows&nbsp;a classification of variables for predictive/diagnostic models used in silico environmental chemical science, in terms of system scales and variable types. Figure 3 shows&nbsp;a continuum of system scales encompassing the whole scope of predictive/diagnostic modelling for in silico environmental chemical sciences, juxtaposing earth and biological scales.</p> <p>The published version of Figure 3 is tall, for two-column page-layouts, but a wide version of Figure 3 is provided for landscape oriented formats. The 300 dpi versions of each figure should be adequate resolution for most purposes, and therefore are recommended.&nbsp;The large versions of the figures may take significant time to download, but may be useful for high resolution applications.</p> <p>This work is from the perspectives/review paper at the beginning of a themed issue on &quot;Quantitative Structure-Activity Relationships (QSARs) and Computational Chemistry Methods in the Environmental Chemical Sciences&quot;, published in the March 2017 issue of the Royal Society of Chemistry journal Environmental Sciences: Process and Impacts. The whole collection of papers can be accessed at rsc.li/qsars.</p>

opencc-by-4.0Aug 2017View details →
zenodo44/100

nextGEMS cycle3 datasets: statistical summaries for streamed data from climate simulations

<p>This Zenodo holds the datasets used in the paper "Statistical summaries for streamed data from climate simulations" by Katherine Grayson. All the data comes from the nextGEMS cycle 3 and has been regridded for plotting purposes with resolution given in the title of each data set. The wind speed data set has been made by taking the square root of the squared and summed v10 and u10 components respectively. All data has been retrieved and regridded through the AQUA reader on the Levante supercomputer, developed as part of the Destination Earth initative. The source code to create all the figures using this data can be found in https://github.com/kat-grayson/one_pass_algorithms_paper/tree/main&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

German Student Responses to Probability Theory and Statistics Bachelor Course (WuS24): Evaluated with Rubrics

<p><strong>Description:</strong></p> <p>This dataset contains questions and answers from an introductory computer science bachelor course on statistics and probability theory at Hochschule Bonn-Rhein-Sieg. The dataset includes three questions and a total of 90 answers, each evaluated using binary rubrics (yes/no) associated with specific scores.</p> <p>&nbsp;</p> <p><strong>Dataset Components:</strong></p> <ol> <li><em>questions.csv</em>: Contains the details of the three questions. <ul> <li>Columns: <ul> <li><em>question_id</em>: Unique identifier for each question</li> <li><em>question</em>: The text of the question</li> <li><em>solution</em>: The reference answer for the question</li> <li><em>max_score</em>: The maximum score for this question</li> </ul> </li> </ul> </li> <li><em>rubrics.csv</em>: Contains the grading rubrics for each question. <ul> <li>Columns: <ul> <li><em>question_id</em>: Unique identifier for each question</li> <li><em>rubric_id</em>: Unique identifier for each rubric within a question</li> <li><em>rubric</em>: The rubric phrased as a question</li> <li><em>score</em>: The score associated with fulfilling the rubric</li> </ul> </li> </ul> </li> <li><em>answers.csv</em>: Contains 90 student answers to the questions. <ul> <li>Columns: <ul> <li><em>answer_id</em>: Unique identifier for each answer</li> <li><em>question_id</em>: Unique identifier of the question that is answered</li> <li><em>answer</em>: The text of the student's answer</li> <li><em>score</em>: The score associated with fulfilling the rubric</li> </ul> </li> </ul> </li> <li><em>answer_rubrics.csv</em>: Contains the evaluations of rubrics for each answer.<br> <ul> <li>Columns: <ul> <li><em>answer_id</em>: The identifier of the answer.</li> <li><em>question_id</em>: The identifier of the question.</li> <li><em>rubric_id</em>: The identifier of the rubric for that question.</li> <li><em>label</em>: Indicates if the rubric crierion is fulfilled for the specific answer (true / false).</li> </ul> </li> </ul> </li> </ol> <pre><strong><br>Working with the Dataset:</strong> The easiest way to work with this dataset is using the class `RubricsDataset` defined in the file `dataloader.py`. Example:<br><br></pre> <pre><code>from dataloader import RubricsDataset<br><br>dataset = RubricsDataset.from_directory("data")<br> dataset.get_question(1) # Get a dictionary containing info about the first question, including rubrics dataset.get_answers(1) # Get all the answers for the first question as a list. Each answer is a dictionary with answer, score, rubrics.</code></pre>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Participant survey for the article: More than Formulas - Integrity, Communication, Computing and Reproducibility in Statistics Education

<p>The artcile More than Formulas - Integrity, Communication, Computing and Reproducibility in Statistics Education concerns the introduction of a new course format in the Master Program in Biostatistics at the University of Zurich. This data set contains the results fo a survey among the participants in this new course.</p> <p>Sepcifically it contains the answers of 22 participants to the following questions:</p> <p>1) Did you use the following concepts or tools since you took STA472?&nbsp;<br>Good practice for...</p> <p>... spreadsheets<br>... file and folder organization<br>... version control<br>... dynamic reporting<br>... LaTeX<br>... presentation slide design&nbsp;<br>... oral presentations<br>... designing graphs<br>... designing tables<br>... structure for manuscript<br>... logic of a paragraph<br>... writing style<br>... writing R functions<br>... using unit tests<br>... setting up simulations<br>... code styling<br>... writing vectorized code<br>... writing parallelized code<br>... containerizing code</p> <p>Answers are in the scale: never since, rarely, sometimes, often, frequently, I do not know</p> <p>2) If you used the above concepts at least rarely, did the training of STA472 help you?</p> <p>Good paractice for...</p> <p>... spreadsheets<br>... file and folder organization<br>... version control<br>... dynamic reporting<br>... LaTeX<br>... presentation slide design&nbsp;<br>... oral presentations<br>... designing graphs<br>... designing tables<br>... structure for manuscript<br>... logic of a paragraph<br>... writing style<br>... writing R functions<br>... using unit tests<br>... setting up simulations<br>... code styling<br>... writing vectorized code<br>... writing parallelized code<br>... containerizing code</p> <p>Answers are in the scale: Not really &nbsp; Somewhat &nbsp;Definitively &nbsp; I do not know I do not use this concept</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Precomputed data accompanying the paper "Reduced Data-Driven Turbulence Closure for Capturing Long-Term Statistics"

<p>Together with the python code in <a href="https://github.com/rik-stra/tau_orthogonal_method_for_2D_turbulence" target="_blank" rel="noopener">https://github.com/rik-stra/tau_orthogonal_method_for_2D_turbulence</a>, this dataset lets the user interact with and reproduce the results in the paper "Reduced Data-Driven Turbulence Closure for Capturing Long-Term Statistics".</p> <p>&nbsp;All data in this dataset is computed with the provided code.</p>

openmit-licenseJul 2024View details →
zenodo44/100

Population Statistics on NCT02332590

<p>Data presented as mean or median (change), in parts with quantiles (Q1, Q3) published by</p> <table> <tbody> <tr> <td> <p>C. Gaby, G. Burmester, V. Strand, J. Msihid, M. Zilberstein, T. Kimura, B. van Hoogstraten, S. H. Boklage, J. Sadeh, M. H. G. Neil and A. Boyapati, "Sarilumab and adalimumab differential effects on bone remodelling and cardiovascular risk biomarkers, and predictions of treatment outcomes," <em>Arthritis Research &amp; Therapy, </em>vol. 22, no. 1, p. 70, 2020.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Data and Statistical analysis for: "Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research"

<p>Data and Statistical analysis for: &quot;Predator in the pool? A quantitative evaluation of non-indexed open access journals in aquaculture research&quot; published in&nbsp;<em>Frontiers in Marine Science</em></p>

openmit-licenseMar 2018View details →
zenodo44/100

Genome-wide association summary statistics for varicose veins of lower extremities

<p>The dataset contains summary statistics for the discovery and the replication stages of the large-scale genome-wide associations study for varicose veins of lower extremities. The discovery stage was based on genetic association data provided by the Neale Lab (<a href="https://vk.com/away.php?to=http%3A%2F%2Fwww.nealelab.is%2F&amp;cc_key=">http://www.nealelab.is/</a>) for 337,199 UK biobank individuals. Phenotype &ldquo;varicose veins of lower extremities&rdquo; was defined based on International Classification of Disease (ICD-10) billing code &ldquo;I83&rdquo; present in the electronic patient record. Data were adjusted for two potential confounders &ndash; body mass index and deep venous thrombosis. A replication cohort (N=71,256) was generated by means of reverse meta-analysis of two overlapping datasets: genetic association data for 408,455 UK Biobank participants provided by the Gene ATLAS database (<a href="https://vk.com/away.php?to=http%3A%2F%2Fgeneatlas.roslin.ed.ac.uk%2F&amp;cc_key=">http://geneatlas.roslin.ed.ac.uk/</a>), and the above mentioned data provided by the Neale Lab.</p> <p>Please, note, that in Shadrina et al&nbsp;(PLOS&nbsp;Genetics 2019) we only used &quot;discovery&quot; dataset, while in biorxiv preprint (https://doi.org/10.1101/368365) both discovery and replication datasets were used.&nbsp;</p> <p>The data are provided on an &quot;AS-IS&quot; basis, without warranty of any type, expressed or implied, including but not limited to any warranty as to their performance, merchantability, or fitness for any particular purpose. If investigators use these data, any and all consequences are entirely their responsibility. By downloading and using these data, you agree that you will cite the appropriate publication in any communications or publications arising directly or indirectly from these data; for utilisation of data available prior to publication, you agree to respect the requested responsibilities of resource users under 2003 Fort Lauderdale principles; you agree that you will never attempt to identify any participant.&nbsp;</p> <p><strong>When using downloaded data, please cite corresponding paper and this repository:</strong></p> <ol> <li> <p>Shadrina, A. S., Sharapov, S. Z., Shashkova, T. I. &amp; Tsepilov, Y. A. Varicose veins of lower extremities: Insights from the first large-scale genetic study. <em>PLOS Genet.</em> <strong>15,</strong> e1008110 (2019).</p> </li> <li>Alexandra S. Shadrina, Sodbo Zh. Sharapov, Tatiana I. Shashkova, &amp; Yakov A. Tsepilov. (2018). Genome-wide association summary statistics for varicose veins of lower extremities (Version 1) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1323484</li> </ol> <p><strong>Funding:</strong></p> <p>The work of ASS was supported by the Russian Science Foundation [Project No 17-75-20223].&nbsp;<br> The work of YAT was supported by the Russian Ministry of Science and Education under the 5-100 Excellence Programme.&nbsp;<br> The work of SZS was supported by the Institute of Cytology and Genetics [Project No 0324-2018-0017].</p> <p><strong>Column headers - discovery</strong></p> <ol> <li>SNP: SNP rsID</li> <li>b: effect size of effect allele</li> <li>se: standard error of effect size</li> <li>chi2: T^2 value of effect allele</li> <li>Pval: P-value of association (without GC correction)</li> <li>N:&nbsp;sample size</li> <li>Chr: chromosome</li> <li>Pos: position (GRCh37 build)</li> <li>A1: effect allele (coded as &quot;1&quot;)</li> <li>A2: reference allele (coded as &quot;0&quot;)</li> </ol> <p><strong>Column headers - replication</strong></p> <ol> <li>SNP: SNP rsID</li> <li>A1: effect allele (coded as &quot;1&quot;)</li> <li>A2: reference allele (coded as &quot;0&quot;)</li> <li>N: Total sample size</li> <li>Z: Z-value of effect allele</li> <li>P: P-value of association (without GC correction)</li> </ol>

opencc-by-4.0Jul 2018View details →
zenodo44/100

Asti City Statistics

<p>JSON data related to average statistics on children food habits and physical activities (for the city of Asti)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Milan Schools Statistics

<p>JSON data related to average statistics on children food habits and physical activities (per school)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Asti Post Code Statistics

<p>JSON data related to average statistics on children food habits and physical activities (per postcode)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Asti Schools Statistics

<p>JSON data related to average statistics on children food habits and physical activities (per school)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Milan Post Code Statistics

<p>JSON data related to average statistics on children food habits and physical activities (per postcode)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Milan City Statistics

<p>JSON data related to average statistics on children food habits and physical activities (for the city of Milan)</p>

opencc-by-sa-4.0Jul 2018View details →
zenodo44/100

Aššur and His Friends: A Statistical Analysis of Neo-Assyrian Texts

<p>This is the data used for and generated during our research for the article &quot;A&scaron;&scaron;ur and His Friends: A Statistical Analysis of Neo-Assyrian Texts&quot;, published in <em>Journal of Cuneiform Studies </em>71 (2019).</p>

opencc-by-sa-3.0Mar 2019View details →
zenodo44/100

Dataset for "Reflectance spectra of seven lunar swirls examined by statistical methods: A space weathering study"

<p>This archive corresponds to the source code, raw data, and results described in the article &quot;Reflectance spectra of seven lunar swirls examined by statistical methods: A space weathering study&quot; by Chrbolkov&aacute; et al. (2019) published in Icarus journal. See AA_README.txt for more information.</p>

opencc-by-4.0May 2019View details →
zenodo44/100

Computed Basic Statistics of Hydraulics and Discharge Measures at USGS River Monitoring Stations

<p>The shared table contains&nbsp;basic statistics (average, standard deviation, minimum, maximum, and coefficient of variation [%]) river channel hydraulics and discharge&nbsp;records of the 4472 USGS river monitoring stations. The required raw data are free to access&nbsp;by the USGS-<em>National Water Information System</em>&nbsp;(<a href="https://waterdata.usgs.gov/nwis">https://waterdata.usgs.gov/nwis</a>). Hydraulics and discharge records that&nbsp;measured&nbsp;at&nbsp;each USGS monitoring site,&nbsp;given a long time period, were assembled, assessed, and finally used for computing the basic statistics.&nbsp;&nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo44/100

Development and validation of statistical shape models of the primary functional bone segments of the foot.

<p>This dataset comprises manually segmented three-dimensional point clouds (.STL) of magnetic resonance images&nbsp;of the primary functional segments of the foot -&nbsp;first metatarsal, midfoot (second-to-fifth metatarsals, cuneiforms, cuboid, and navicular), calcaneus, and talus. These data were used to create statistical shape models of the foot bones, utilising the GIAS2 toolbox&nbsp;(https://pypi.org/project/gias2/).</p>

opencc-by-4.0Sep 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record