Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
173
datasets available to search
ShareScore release 0.9.0
Dataset results
173 results for “Statistical analysis”
Replication Package for: Mapping Firms' Locations in Technological Space: A Topological Analysis of Patent Statistics
<p>This replication package contains the data and the code to generate the paper’s main results, as well as the Online Appendix, for “Mapping Firms’ Locations in Technological Space: A Topological Analysis of Patent Statistics” by Emerson G. Escolar, Yasuaki Hiraoka, Mitsuru Igami, and Yasin Ozcan (published in <em>Research Policy</em>, volume 52, issue 8, October 2023; full text available online at https://doi.org/10.1016/j.respol.2023.104821).</p>
Data and statistical analysis scripts for manuscript on pennycress roots & response to nitrate using 3D gel system
<p>Data and statistical analysis scripts for manuscript on pennycress roots & response to nitrate using 3Dgel system</p> <blockquote> <p><strong>A temporal analysis and response to nitrate availability of 3D root system architecture in diverse pennycress (<em>Thlaspi arvense</em> L.) accessions</strong> - [<a href="https://doi.org/10.3389/fpls.2023.1145389">https://doi.org/10.3389/fpls.2023.1145389</a>]</p> </blockquote> <p>The following files contains:</p> <ul> <li><code>gel_data_preprocessing_20221024.R</code> - R statistics script for pre-processing data files from 3Dgel system GIARoots & DynamicRoots raw output</li> <li><code>gel_dataprocessing_20221229.R</code> - R statistics script for data processing of pre-processed 3D gel data</li> <li><code>TaGNS_N_Spring32.zip</code> - CSV data files and R statistics script for Spring32 grown under high, low, trace and zero N treatments.</li> <li><code>TaGNE_N_Accessions.zip</code> - CSV data files and R statistics script for 3 accessions under high and trace N treatments.</li> <li><code>TaGAA_N_Accessions.zip</code> - CSV data files and R statistics script for 24 diverse pennycress lines grown under high N conditions.</li> </ul>
Dataset, statistical analysis code, and supplementary material of juvenile ravens' responses towards acoustic cues of different social categories
Open the record for dataset details and reuse information.
Statistical analysis code for output from a model used to simulate foot-and-mouth disease dynamics in the United Kingdom
Open the record for dataset details and reuse information.
Data from: Estimating transmission dynamics and serial interval of the first wave of COVID-19 infections under different control measures: A statistical analysis in Tunisia from February 29 to May 5, 2020
<p>Background: Describing transmission dynamics of the outbreak and impact of intervention measures are critical to planning responses to future outbreaks and providing timely information to guide policy makers decision. We estimate serial interval (SI) and temporal reproduction number (R<sub>t</sub>) of SARS-CoV-2 in Tunisia.</p> <p>Methods: We collected data of investigations and contact tracing between March 1, 2020 and May 5, 2020 as well as illness onset data during the period February 29-May 5, 2020 from National Observatory of New and Emerging Diseases of Tunisia. Maximum likelihood (ML) approach is used to estimate dynamics of R<sub>t</sub>.</p> <p>Results: 491 of infector-infectee pairs were involved, with 14.46% reported pre-symptomatic transmission. SI follows Gamma distribution with mean 5.30 days [95% CI 4.66-5.95] and standard deviation 0.26 [95% CI 0.23-0.30]. Also, w<span>e estimated large changes in </span>R<sub>t</sub><span> in response to the combined lockdown interventions. The </span>R<sub>t</sub><span> moves from </span>3.18 [95% CI 2.73-3.69] <span>to 1.77 [95% CI 1.49-2.08] with </span>curfew<span> prevention measure, and under the epidemic threshold (0.89 </span>[95% CI 0.84-0.94]) by national lockdown measure<span>.</span></p> <p><span>Conclusions: </span>Overall, our findings highlight contribution of <span>interventions</span> to interrupt transmission of SARS-CoV-2 in Tunisia.</p>
Collision between biological process and statistical analysis revealed by mean-centering
<p>Animal ecologists often collect hierarchically-structured data and analyze these with linear mixed-effects models. Specific complications arise when the effect sizes of covariates vary on multiple levels (e.g., within vs among subjects). Mean-centering of covariates within subjects offers a useful approach in such situations, but is not without problems.</p> <p>A statistical model represents a hypothesis about the underlying biological process. Mean-centering within clusters assumes that the lower level responses (e.g. within subjects) depend on the deviation from the subject mean (relative) rather than on absolute values of the covariate. This may or may not be biologically realistic. We show that mismatch between the nature of the generating (i.e., biological) process and the form of the statistical analysis produce major conceptual and operational challenges for empiricists.</p> <p>We explored the consequences of mismatches by simulating data with three response-generating processes differing in the source of correlation between a covariate and the response. These data were then analyzed by three different analysis equations. We asked how robustly different analysis equations estimate key parameters of interest and under which circumstances biases arise. </p> <p>Mismatches between generating and analytical equations created several intractable problems for estimating key parameters. The most widely misestimated parameter was the among-subject variance in response. We found that no single analysis equation was robust in estimating all parameters generated by all equations. Importantly, even when response-generating and analysis equations matched mathematically, bias in some parameters arose when sampling across the range of the covariate was limited.</p> <p>Our results have general implications for how we collect and analyze data. They also remind us more generally that conclusions from statistical analysis of data are conditional on a hypothesis, sometimes implicit, for the process(es) that generated the attributes we measure. We discuss strategies for real data analysis in face of uncertainty about the underlying biological process.</p>
Analysis of performance indicators for German Library Statistics in 2019
<p>Analyse der Berechenbarkeit von Kennzahlen der ISO 11620:2014 auf Basis der Deutschen Bibliotheksstatistik des Berichtsjahres 2019.</p> <p>Analysis of the calculability of performance indicators of ISO 11620:2014 on the basis of the German Library Statistics of the reporting year 2019.</p>
Data from: Statistical analysis of the presidential elections in Belarus in 2020
<p>Elections in Belarus attract much attention around the world. The election result is declared as a victory of Mr. Lukashenko with 80% votes. It is interesting to give the simplest statistical analysis of this victory. According to Belarus law, protocols of precinct election commissions (PECs) must be posted up just after the election procedure, so that everybody could take a photograph of the protocols. Currently, 1527 of the 5767 protocols of PECs are available in the open access at <a href="https://docs.google.com/spreadsheets/d/17aK3JxBTGtzULB0-YZGOF0hJwhuViHO3/edit#gid=84585767">https://docs.google.com/spreadsheets/d/17aK3JxBTGtzULB0-YZGOF0hJwhuViHO3/edit#gid=84585767</a>. We focus an attention on two arrays of numbers taken from these photographs. Namely, the number N<sub>i</sub> of voters at some polling station and the number M<sub>i</sub> of voters for Mr. Lukashenko at the same polling station. These numbers give a possibility to calculate the average percentage of those who voted for Mr. Lukashenko, which turns out to be <span>about 60%.</span> That is, a random sample approximately of ¼ of total number of protocol gives a value that differs at about 20% from the declared total value 8<span>0%.</span> Using Monte-Carlo simulation we have calculated a probability of this event and obtain <span>less than one part in million.</span> Next we have considered N<sub>i </sub>and M<sub>i</sub> as the random variables and calculate probability distribution functions for M<sub>i</sub>/N<sub>i</sub> and M<sub>i</sub>/<N<sub>i</sub>> quantities. First function f(x) is of non Gaussian form and has a maximum at x≈0.6, and, an additional maximum at x≈0.8. Second function f(y) has only one maximum at y≈0.6. <span>One the possible explanations is that the correlation (i.e. maximum) in distribution </span>f(x) at x≈0.8 <span> arises due to artificial trimming of the percentage of those who voted for Lukashenko to </span>8<span>0% in some polling stations.</span></p>
Statistical Analysis of the Effect of Equations on Citations
<p>Statistical analysis of a data set of number of equations and number of citations of papers published in volumes 94 and 104 of the journal <em>Physical Review Letters</em>. This analysis is referred to by the paper <strong>Equation-dense papers receive fewer citations—in physics as well as biology</strong> in the <em>New Journal of Physics </em>(vol. 18, article 118003) by Andrew D Higginson and Tim W Fawcett. http://iopscience.iop.org/article/10.1088/1367-2630/18/11/118003</p>
Statistical Analysis of Overlapping Double Ion Energy Dispersion Events in the Northern Cusp (Paper Data and Code)
<p>This data and code accompanies the paper <i>Statistical Analysis of Overlapping Double Ion Energy Dispersion Events in the Northern Cusp</i>, published in Frontiers in Astronomy and Space Sciences in 2023. This upload includes a CSV list of the selected events, a human-readable table of the selected events (see below), plots of each selected event, and the code used to select the events.</p><p>The code in this repository is a fork of <a href="https://github.com/ddasilva/dmsp-dispersion-detection">https://github.com/ddasilva/dmsp-dispersion-detection</a> at the time of publication. Future updates may exist on GitHub.</p><p><br> </p>
Statistical Analysis of the submitted paper
<p><span>Statistical Analysis of the submitted paper to the Journal: Contributions to the knowledge of cumaceans (Crustacea, Peracarida) in a shallow-water hydrothermal system at Mexican Central Pacific.</span></p>
Efficient statistical method for single-cell QTL analysis
<p>This upload contains data objects associated with our paper "Efficient statistical method for single-cell QTL analysis" introducing the SAIGE-QTL method (preprint available soon!).</p>
Statistical analysis of data determining minimal selective concentration of Amoxicillin, doxycycline and enrofloxacin
<p>This repository is containing the datasets with statistical analysis of the paper: "Selection for amoxicillin, doxycycline and enrofloxacin resistant Escherichia coli at concentrations lower tha nthe ECOFF in rich media and in broiler-derived fecal fermenations."</p> <p><em>The phenotypic analysis</em></p> <ul> <li>Phenotypic amoxicillin <ul> <li>R-file: Phenotypic Amox mixed models, CSV-file: Phenotypic data Amox</li> </ul> </li> <li>Phenotypic doxycycline <ul> <li>R-file: Phenotypic Dox mixed models, CSV-file: Phenotypic data Dox</li> </ul> </li> <li>Phenotypic enrofloxacin <ul> <li>R-file: Phenotypic Enro mixed models, CSV-file: Phenotypic data Enro</li> </ul> </li> </ul> <p><em>The resistome analysis </em></p> <ul> <li>two CSV-files: Resistoom workfile and resistome_reference_file</li> <li>one R-file: Resistome_analysis_antimicrobial_classes</li> </ul> <p><em>Microbiome analysis<br></em></p> <ul> <li>Alpha- and beta-diversity analysis<br> <ul> <li>R-file Biom analysis, CSV-file: metadata2, biom-file: reads_fermentation</li> </ul> </li> <li>Microbial abundance analysis <ul> <li>R-file: Abundance plot, CSV-file: metadata2. biom-file: reads_fermentation2.biom</li> </ul> </li> </ul>
Sahana et al. Supplementary Data for Global Transboundary River Research: Databases, Case Study Analysis, and Regional Statistics for Sustainable Management
<p><span>This dataset supports our comprehensive review article on transboundary river research, exploring its implications for sustainable management worldwide. Utilizing machine learning, we analyzed 4,237 publications and conducted an in-depth desk review of 325 selected papers, examining a total of 4,713 case studies spanning 286 river basins globally. The study provides critical insights into upstream, midstream, and downstream regions, offering a complete view of challenges and opportunities in transboundary river management. Supplementary Data 1 contains the main database used in this study, sourced from Scopus, Web of Science, and Google Scholar. Additionally, Supplementary Data 2 and 3, included in the spreadsheet, offer statistics and further resources essential for understanding regional and cross-regional dynamics in river basin governance. These supplementary resources include key statistics, case study metadata, and tools, helping to facilitate a deeper exploration of basin-specific and global trends in transboundary water management. This collection of data and resources provides a valuable foundation for researchers and policymakers in advancing sustainable transboundary river management practices.</span></p>
Summary Statistics from "Genome-wide meta-analysis of phytosterols reveals five novel loci and a detrimental effect on coronary atherosclerosis"
<p>Summary statistics of meta GWAS of phytosterols.</p>
Dataset of ""Statistical atlases and automatic labelling strategies to accelerate the analysis of social insect brain evolution"
<p>Dataset of <em>Statistical atlases and automatic labelling strategies to accelerate the analysis of social insect brain evolution</em> by Sara Arganda, Ignacio Arganda-Carreras, Darcy G. Gordon, Andrew P. Hoadley, Alfonso Pérez-Escudero, Martin Giurfa and James F. A. Traniello.</p> <p>In this dataset, we are presenting:</p> <ul> <li>10 confocal brain images from <em>Pheidole spadonia </em>minors (in the original confocal TIFF format and in the open NRRD format), with manually segmented labels of 8 subregions (Optic Lobes, OL; Antennal Lobes, AL; Mushroom Body Medial Calyx, MB-MC; Mushroom Body Lateral Calyx, MB-LC; Mushroom Body Peduncle, MB-P; Central Complex, CX; Subesophageal zone, SEZ; and Rest of Central Brain, ROCB – in NRRD format) from one expert annotator.</li> <li>12 confocal brain images from <em>P. spadonia</em>, <em>P. rhea</em>, <em>P. tepicana</em> and <em>P. obtusospinosa</em> minors, with manually segmented labels of the same 8 subregions (OL; AL; MB-MC; MB-LC; MB-P; CX; SEZ; and ROCB) from one expert annotator.</li> <li>5 confocal brain images from <em>Pheidole spadonia </em>minors (“test brains”), with five sets of manually segmented labels of the same 8 subregions (OL; AL; MB-MC; MB-LC; MB-P; CX; SEZ; and ROCB) from three expert annotators (one set from annotator 1, one set from annotator 2 and three sets from annotator 3, to evaluate inter and intra person differences).</li> <li>1 group-wise template generated from the 10 confocal brain images from <em>Pheidole spadonia </em>minors, with three sets of manually segmented labels of the same 8 subregions (OL; AL; MB-MC; MB-LC; MB-P; CX; SEZ; and ROCB).</li> <li>5 group-wise templates generated from the 9 confocal brain images from <em>Pheidole spadonia </em>minors, with consensus labels of the same 8 subregions (OL; AL; MB-MC; MB-LC; MB-P; CX; SEZ; and ROCB).</li> <li>1 group-wise template generated from 12 confocal brain images from <em>P. spadonia</em>, <em>P. rhea</em>, <em>P. tepicana</em> and <em>P. obtusospinosa</em> minors, with consensus labels of the same 8 subregions (OL; AL; MB-MC; MB-LC; MB-P; CX; SEZ; and ROCB).</li> <li>7 sets of automatic labels for the 5 “test brains”: 3 sets of “Direct Labels”, 3 sets of “Consensus Labels”, 1 set of “Multispecies Template Labels”.</li> </ul> <p>Brain of minor workers were dissected from the ant head capsule in ice cold HEPES-buffered saline and were fixed and immunohistochemically stained using SYNORF1 (a monoclonal <em>Drosophila</em> synapsin I antibody obtained from the Developmental Studies Hybridoma Bank, catalog 3C11) and secondarily stained using Alexa Fluor 488 for visualization of neuropil (slightly modified from Ott, 2008). Later, brains were mounted in methyl salicylate and imaged on an Olympus Fluoview BX50 laser scanning confocal microscope with a ×20 objective at a resolution of ~0.7 × 0.7 × 5µm/voxel. All brain tissue manipulation, staining and recording was performed by Darcy G. Gordon. Brain images were obtained in TIFF format by the confocal microscope and then opened and saved as Amira Mesh (.am) stack images in Amira (version 6.0). Manual segmentation of each brain was done using Amira (version 6.0 or 2019.2). Labels were traced on eight compartments in only one brain hemisphere, except for the CX, SEZ and ROCB, which lack a clear subdivision between hemispheres. Brain grey image stacks and labels were transformed to NRRD format for template construction using the Fiji plugin SaveAsGzipNrrd<a href="#_ftn1">[1]</a>. Volume and volume similarity of labels were calculated using the Fiji toolbox MorphoLibJ<a href="#_ftn2">[2]</a>.</p> <p><strong>Acknowledgements: </strong>We thank Ming Huang (from Dr. Diana Wheeler’s laboratory) who kindly provided access to colonies from four species of the hyperdiverse ant genus <em>Pheidole</em> (<em>P. spadonia</em>, <em>P. rhea</em>, <em>P. tepicana </em>and <em>P. obtusospinosa</em>). This research was supported by National Science Foundation grants IOS 1354291 and IOS 1953393 to JT, a Marie Skłodowska-Curie Individual Fellowship BrainiAnts-660976 and Ayudas destinadas a la atracción de talento investigador a la Comunidad de Madrid en centros de I+D. This work is supported in part by the University of the <a href="https://www.sciencedirect.com/topics/engineering/basque-country">Basque Country</a> UPV/EHU grant GIU19/027.</p> <p> </p> <p><a href="#_ftnref1">[1]</a> https://github.com/iarganda/tefor</p> <p><a href="#_ftnref2">[2]</a> https://imagej.net/plugins/morpholibj</p>
Figure 5 in Statistical analysis on the cnidome of genus Hydra using Generalized Linear Models
Figure 5. Relative abundances of each type of cnidocyst for each species. (A) stenotele, (B) desmoneme, (C) atrichous isorhiza and (D) holotrichous isorhiza.
Figure 1 in Statistical analysis on the cnidome of genus Hydra using Generalized Linear Models
Figure 1. Cnidome of Hydra viridissima. (A) stenotele, (B) desmoneme, (C) atrichous isorhiza and (D) holotrichous isorhiza. Scale bar: 3 μm.
Figure 4 in Statistical analysis on the cnidome of genus Hydra using Generalized Linear Models
Figure 4. Different morphotypes of holotrichous isorhiza. (A) Hydra viridissima, (B) Hydra vulgaris pedunculata, C and (D) Hydra vulgaris. Scale bar: 2.45 μm.
Figure 8 in Statistical analysis on the cnidome of genus Hydra using Generalized Linear Models
Figure 8. GLM adjustment graphs used for comparison between species. (A) scatter plot, (B) Q-Q Plots.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.