Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
The metabolomics raw data and a supporting statistical analyses data set for publication: Metabolomic analysis revealed the absence of the principal antimicrobial compound of Pseudomonas donghuensis P482, 7-hydroxytropolone, under restricted nutrient conditions.
<p><a href="../api/records/11220997/draft/files/Metabolomic%20analyses%20raw%20files.zip/content" target="_blank" rel="noopener noreferrer">Metabolomic analyses raw files</a>, Compounds analyses, Hierarchical Condition tress and PCA Scores are uploaded.</p>
Size-dependent Copy Number Analysis Data
<p>Data<strong> </strong>used for analysis described in <em><span>Pan-cancer copy number analysis identifies optimized size thresholds and co-occurrence models for individualized risk stratification</span></em></p> <p><strong>MeningiomaSegs.zip </strong>contains the generated copy number segmentation from the sesame pipeline for 565 meningioma samples used in this project</p> <p><strong>MeningiomaClinical.xlsx </strong>contains deidentified clinical outcomes and relevant molecular data for meningioma samples</p> <p><strong>blank_CNV_table.xlsx </strong>is a blank template loaded in several R scripts in this project</p> <p><strong>hg38.14_chrom_arms_manual.txt </strong>Details the starting and ending points of chromosome arms</p> <p><strong>sampleinventory_excluded.xlsx</strong> Lists the TCGA cancer types and how many samples have both copy number and clinical data after manual exclusion of other samples</p>
Data and analysis for "Association of self-reported musculoskeletal pain with school furniture suitability and daily activities among primary school and university students"
<p>Datasets and R analysis code and report for the article "Association of self-reported musculoskeletal pain with school furniture suitability and daily activities among primary school and university students".</p>
Data+Analysis+Plotting scripts for "Constraint on the dissipative tidal deformability of neutron stars"
<p>The .zip file contains three directories. </p> <p>1. GW170817-Strain: the raw data (glitch free). Downloaded from https://gwosc.org/events/GW170817/.<br>2. Bilby-Output: the output from running our Bilby sampling scripts. These can be found at https://github.com/JLRipley314/NRTidal-D/tree/main<br>3. Plotting-Scripts: the plotting scripts we used in our paper https://arxiv.org/abs/2312.11659.</p> <p>NOTE: If you want to make sure the plotting scripts work properly, you should download bilby and related dependencies as described in https://github.com/JLRipley314/NRTidal-D/tree/main (or at https://doi.org/10.5281/zenodo.11589416)</p>
Sample data for "Live Cell Fluorescence Microscopy – An End-to-End Workflow for High-Throughput Image and Data Analysis"
<p>This repository contains:</p> <ul> <li> <p>Sample data for the "Live Cell Fluorescence Microscopy – From Sample Preparation to Numbers and Plots" methodology paper by Zahumensky & Malinsky. The paper describes the preparation of live yeast cell samples for microscopy, the subsequent semi-automatic analysis of the microscopy images using our custom-written Fiji macros, and automatic processing of the output (Results table) from the image analys using custom-written R scripts. The data provided here are real experimental data from two publications of our group: Zahumensky et al., 2022 and Vesela et al., 2023</p> </li> <li> <p>"Results tables" from the Fiji based analysis</p> </li> <li> <p>Outputs of the processing of these Results tables using our R scripts, in the form of summary tables, graphs, and statistical analyses</p> </li> </ul>
Supplementary File 10; The full data set used for the analysis presented herein:
Open the record for dataset details and reuse information.
Data of thermal analysis (TG) of Hilber et al. 2024
Open the record for dataset details and reuse information.
Inputs, results data and analysis script for the evaluation of the PDG-Arena forest growth model on beech-fir stands
<p>Supplementary files for simulations in Rouet et al. (2024): PDG-Arena: An eco-physiological model for characterizing tree-tree interactions in heterogeneous and mixed stands (doi: <a href="https://doi.org/10.1101/2024.02.09.579667" target="_blank" rel="noopener">10.1101/2024.02.09.579667</a>).</p> <p>This repository is an archive of the github repository PDG-Arena-extra (release v1.0.3), accessible at <a href="https://github.com/camille-rouet/PDG-Arena-extra/tree/v1.0.3" target="_blank" rel="noopener">https://github.com/camille-rouet/PDG-Arena-extra/tree/v1.0.3</a>.</p>
Code and Data Supplement for Using feature importance as exploratory data analysis tool on earth system models
<p>This contains:</p> <ul> <li>Code for all analyses in</li> <li>E3SM data</li> </ul> <p>For the paper Using <em>feature importance as exploratory data analysis tool on earth system models.</em></p>
Code and data for spatial and temporal magnitude clustering analysis
<p>Code used for performing spatial and temporal seismic magnitude clustering analysis. Includes documentation (README.txt) with steps on how to implement the code. The public datasets used for this study can be accessed at the following locations: </p> <ul> <li><strong>Southern California Catalog: </strong> <ul> <li>SCEDC (2013): Southern California Earthquake Center.<br> Caltech.Dataset. doi:<a href="https://dx.doi.org/10.7909/C3WD3xH1">10.7909/C3WD3xH1</a></li> </ul> </li> <li><strong>Northern California Catalog:</strong> <ul> <li>NCEDC (2014), Northern California Earthquake Data Center. UC Berkeley Seismological Laboratory. Dataset. doi:10.7932/NCEDC.</li> </ul> </li> <li><strong>Mixed-mode Laboratory Catalog:</strong> <ul> <li>Lin, Qing, et al. "Opening and mixed mode fracture processes in a quasi-brittle material via digital imaging." <em>Engineering Fracture Mechanics</em> 131 (2014): 176-193.</li> </ul> </li> <li><strong>ETAS Code:</strong> <ul> <li>Leila Mizrahi, Shyam Nandan, Stefan Wiemer 2021;<br> Embracing Data Incompleteness for Better Earthquake Forecasting. (Section 3.1)<br> <em>Journal of Geophysical Research: Solid Earth</em>; doi: <a href="https://doi.org/10.1029/2021JB022379">https://doi.org/10.1029/2021JB022379</a></li> </ul> </li> </ul>
Genomic analysis of hyperparasitic viruses associated with entomopoxviruses; supplemental data
<p>This data set includes newick files and pdb files used to generate figures in the publication 'Genomic analysis of hyperparasitic viruses associated with entomopoxviruses'.</p>
Fig. 2 in Biological Data From Post Mortem Analysis Of Otters In Hungary
Fig. 2. Distribution of otter stomach contents according to weight categories in Hungary
Fig. 1 in Biological Data From Post Mortem Analysis Of Otters In Hungary
Fig. 1. Location of otter carcasses collected in different regions of Hungary
Input data of the multi-patch geometries used in: A. Farahat, H. M. Verhelst, J. Kiendl, M. Kapl, Isogeometric analysis for multi-patch structured Kirchhoff–Love shells, Computer Methods in Applied Mechanics and Engineering 411 (2023) 116060 DOI: 10.1016/j.cma.2023.116060
Open the record for dataset details and reuse information.
Fossil Fuel Subsidy Reform - Data and Analysis
<p>This is raw data, the code for cleaning and compiling the data into a panel data, code for mapping and replication of main results for</p> <p>Droste, Chatterton, and Skovgaard (2024). A political economy theory of fossil fuel subsidy reforms in OECD countries. <em>Nature Communications 15</em>: <span>5452</span> </p>
Data Release for "First joint oscillation analysis of Super-Kamiokande atmospheric and T2K accelerator neutrino data"
<p>This archive contains the electronic version in ROOT and pdf formats of the measurements of oscillation parameters obtained with the analyses from the paper “First joint oscillation analysis of Super-Kamiokande atmospheric and T2K accelerator neutrino data”.<br><br>It is published in <a href="https://doi.org/10.1103/PhysRevLett.134.011801">Physical Review Letters</a> and is available on the <a href="https://arxiv.org/abs/2405.12488">arXiv:2405.12488 [hep-ex]</a>. </p> <p>**************************************<br>***** Results included in this release<br>**************************************<br>This release includes the results of the measurements of the different oscillation parameters obtained with the four analyses appearing in the paper. This corresponds to various 1D and 2D DeltaChi^2 and posterior probability maps as well as 2D confidence/credible regions for the parameters sin^2(theta13), sin^2(theta23), dm^2_32/|dm^2_31|, delta_cp, J_cp.</p> <p>The results are separated into different files for the four analyses. Additional details on these analyses can be found below, but two of them use a Bayesian approach (producing posterior probabilities and credible intervals/regions) and two of them follow a frequentist approach (producing DeltaChi^2 maps and confidence intervals/regions). A tag in the TGraph and histogram names also allow to differentiate the different types of intervals/regions: "cred" for credible interval from the Bayesian analysis, "conf" for confidence interval from the frequentist analysis. The tag “posterior” indicates that the object corresponds to a posterior probability distribution, while “chi2” indicates a DeltaChi^2 one.</p> <p>Results for each mass ordering hypothesis are provided, denoted "NO" for normal ordering and "IO" for inverted ordering. The Bayesian files also include results marginalised over the mass ordering, denoted by the tag "both" in the object names. The frequentist files include results profiled over the mass ordering, indicated by a tag “profMO” in the object names.<br>The Bayesian and frequentist results use different conventions for the mass splitting in the inverted ordering: the Bayesian results are in term of #Deltam^{2}_{32} for both NO and IO, whereas the frequentist results are plotted versus #Deltam^{2}_{32} for the NO, and |#Deltam^{2}_{31}| for the IO. </p> <p>A constraint on theta13 from reactor experiment measurements is used for all results in this release. It corresponds to the value in the PDG 2019 review: sin^2(2theta_13)=(8.53+-0.27) x 10^{-2}. This is commonly referred to as "the reactor constraint", and a tag “wRC” is included in the name of the different objects as a reminder that it is used for these results.</p> <p>**************************************<br>***** Brief descriptions of the four analyses<br>**************************************<br>Results are provided for the four analyses mentioned in the “Oscillation analysis” part of the paper. They were given names (Bayesian1, Bayesian2, Frequentist1, Frequentist2) based on the statistical approach they follow.</p> <p>The Bayesian analyses are based on the two T2K analyses described in Eur. Phys. J. C 83, 782 (2023), extended to include the Super-Kamiokande atmospheric data, and with modifications to use the model described in the paper to which the present release is attached to. These analyses use Markov Chain Monte Carlo methods to compute marginal likelihoods for the parameter of interests. </p> <p>For the frequentist analyses, Frequentist1 is a modified version of Bayesian1, optimized for speed to be able to address the computational challenges of producing frequentist results from an ensemble of pseudo-experiments. Frequentist2 is based on the Super-Kamiokande atmospheric analysis described in PTEP 2019, 053F01 (2019), extended to include the T2K data, and also with modifications to follow the model described in the paper. These two analyses compute profile likelihood on a grid of oscillation parameters of interest to produce measurements of these parameters.</p> <p>In terms of the differences between analyses mentioned in the paper, Bayesian2 is the analysis that does a simultaneous fit of the T2K near detector data with the events observed at SK, and the one for which the momentum scale uncertainty is not correlated between the atmospheric and T2K events observed at SK. The three other analyses use a covariance matrix to propagate the constraint on systematic uncertainties from T2K near detector data to the analysis of the events observed at SK, and treat the momentum scale uncertainty as correlated between atmospheric and T2K far detector events.</p> <p>**************************************<br>***** Example codes<br>**************************************<br>Example codes are provided for each of the four analyses, showing how to produce the pdf file from this analysis from the corresponding ROOT file. How to run these example codes is indicated in the comments at the start of each of the example files.</p> <p>**************************************<br>***** Objects inside the ROOT files<br>**************************************<br>The ROOT objects contained inside the files are named first with an identifier of which parameter(s) are being shown, followed by the reactor constraint tag, followed by the mass ordering tag.</p> <p>For the Bayesian results, there is an additional tag to indicate if the results was obtained with a prior probability uniform in deltaCP (“flatdcp”) or uniform in sin(deltaCP) (“flatsindcp”)</p> <p>A glossary is provided at the end of this readme.</p> <p>**************************************<br>*** 2D regions<br>**************************************<br>Objects of the form:<br>gr2D_varX_varY_wRC_<NO,IO,both>_<conf,cred><68,90,955,997>(_N)<br>are TGraphs corresponding to the 2D confidence ("conf") or credible ("cred") regions for the 2 variables (varX, varY). <br>N is the iterator for different TGraphs corresponding to the same region; these occur when confidence regions are discontinuous (for example when deltaCP loops over from +pi to -pi).</p> <p>68, 90, 955, 997 are the percentage credible/confidence levels.</p> <p>Most of the 2D frequentist regions were computed using the standard DeltaChi^2 values (from the Gaussian case), and therefore have only approximate coverage. However, the {sin^2(theta_23), deltaCP} confidence regions of analysis Frequentist1 were built using critical DeltaChi^2 values computed with the Feldman-Cousins method. To distinguish them from other confidence regions, a tag "FC" is included in the name of the corresponding TGraph.</p> <p>The best fit markers are also provided for the 2D results:<br>gr2D_varX_varY_wRC_<NO,IO,both>_bestfit</p> <p>The best fit markers and contour lines are generally for each MO *separately*, i.e. assuming DeltaChi^2 is 0 at the minimum or that the total posterior probability integrates to 1 in the mass ordering considered. There are some exceptions, in particular some 2D regions for (sin^2(theta_23), dcp) are also provided using a best fit over both MO to allow for comparisons with other experiments using this convention. This special set of contours has an extra tag "globalMO" in its name to distinguish it from the others.</p> <p><br>**************************************<br>*** 1D and 2D histograms<br>**************************************<br>Objects of the form<br>h1D_var_<chi2,posterior>_wRC,_<NO,IO, both, profMO><br>h2D_var1_var2_<chi2,posterior>_wRC_<NO,IO, both><br>are respectively TH1D of the DeltaChi^2 ("chi2") or posterior probability ("posterior") for oscillation parameter "var" or TH2D for the couple of parameters (var1, var2)</p> <p>The Bayesian and frequentist results use different conventions with respect to the mass ordering:<br>- DeltaChi^2 plots use a global minimum over both hierarchies<br>- Posterior probability plots integrate to unity *individually*</p> <p>**************************************<br>***** Critical values for frequentist results<br>**************************************<br>For the 1D plots, critical delta chi2 values obtained with the Feldman-Cousins method are provided for theta23 and deltaCP <br>grCritical_{variable}_chi2_wRC_{NO,IO,profMO}_conf{68, 90, 955}<br>variable: th23, dCP</p> <p>The FC-corrected confidence intervals for these 2 variables can be obtained as the region for which the corresponding 1D DeltaChi^2 histogram is below the grCritical graph of a given level.</p> <p>**************************************<br>***** Additional notes for Bayesian results<br>**************************************<br>For plots involving the mass splitting, the mass ordering is given by the sign:<br> dm32>0 is normal hierarchy (Delta m^2_{32} > 0)<br> dm32<0 is inverted hierarchy (Delta m^2_{32} < 0)</p> <p>Note that the posteriors have not been smoothed, and may contain small discontinuities due to MCMC statistical uncertainties.</p> <p>Plots with "_bestfit" appended indicate the point in the 2D parameter space (marginalized over the other parameters) with the highest posterior density, and is not necessarily the global minimum of the likelihood.</p> <p>For the 1D posterior distributions, the user can freely calculate credible intervals from the distributions. It is recommended to start at the point of the highest posterior density, and moving down in posterior density to produce highest posterior credible intervals, which is the kind of credible intervals reported in the paper. </p> <p>**************************************<br>***** Glossary of tags used in objects names<br>**************************************</p> <p>"wRC" - Uses “reactor constraint” on theta13, sin^2(2theta_13)=(8.53+-0.27) x 10^{-2}<br>"FC" - Feldman-Cousins<br>"NO" - Normal mass Ordering<br>"IO" - Inverted mass Ordering<br>"both" - Marginalised over normal and inverted mass orderings<br>"profMO" - Profiled over normal and inverted mass orderings<br>"cred" - Credible interval<br>"conf" - Confidence interval<br>"68" - 68.3% (1 sigma)<br>"90" - 90%<br>"955" - 95.5% (2 sigma)<br>"997" - 99.7% (3 sigma)<br>"chi2" - DeltaChi^2 (-2lnL) for parameter<br>"Critical" - Critical DeltaChi^2 computed using Feldman-Cousins method<br>"th13" - sin^2(theta_13)<br>"th23" - sin^2(theta_23)<br>"dCP" - delta CP<br>"dm2" - Delta m^2_{23} (NO), |Delta m^2_{13} (IO)| for confidence intervals; used for frequentist analyses results.<br>"dm32" - Delta m^{2_{23} regardless of mass ordering; in the Bayesian analyses, Delta m^2_{23} is always the variable that is plotted.<br>"jarlskog" - Jarlskog invariant<br>"flatdcp" - Using prior probability uniform in deltaCP<br>"flatsindcp" - Using prior probability uniform in sin(deltaCP)</p>
Lemonade Creek, Yellowstone National Park, USA - Microbial Community Analysis - Metabolomics Data
<p>Polar metabolomics data (targeted and untargeted) used for analysis of microbial community function over a diurnal cycle in Lemonade Creek, Yellowstone National Park, USA.</p> <p> </p> <p><code>GNPS_positive-2.xlsx</code> Comparison of GNPS data used for main metabolite analysis with targeted metabolite features. Done to support the accuracy of the GNPS results for metabolites identified outside the targeted set.</p> <p> </p> <p><code>NEG_506963_FinalEMA-HILIC_Identifications.xlsx</code> Negative ionization targeted metabolite identification quality and confidence results (prepared by the Joint Genome Institute, USA).</p> <p><code>NEG_msms_mirror_plots.tar.gz</code> Negative ionization targeted metabolite mirror plots.</p> <p><code>NEG_peak_height.tab</code> Negative ionization targeted metabolite peak height file (main results file used for abundance analysis).</p> <p><code>POS_506963_FinalEMA-HILIC_Identifications.xlsx</code> Positive ionization targeted metabolite identification quality and confidence results (prepared by the Joint Genome Institute, USA).</p> <p><code>POS_msms_mirror_plots.tar.gz</code> Positive ionization targeted metabolite mirror plots.</p> <p><code>POS_peak_height.tab</code> Positive ionization targeted metabolite peak height file (main results file used for abundance analysis).</p> <p> </p> <p><code>NEG_peak_height.csv</code> Negative ionization untargeted metabolite peak height file (main results file used for abundance analysis).</p> <p><code>POS_peak_height.csv</code> Positive ionization untargeted metabolite peak height file (main results file used for abundance analysis).</p>
FIGURE 10 in An analysis of fossil identification guides to improve data reporting in citizen science programs
FIGURE 10. An example of a †Cosmopolitodus hastalis photo enhanced by illustration.
Data from: Plastid phylogenomic analysis of Podostemaceae with an emphasis on Neotropical podostemoideae
<p>Podostemaceae are a clade of aquatic flowering plants that form important components of tropical river ecosystems. Species in the family exhibit highly derived growth forms and high vegetative phenotypic plasticity, both of which contribute to taxonomic confusion. The backbone phylogeny of the family remains poorly resolved, many species remain to be included in a molecular phylogenetic analysis, and the monophyly of many taxa remains to be tested. To address these issues, we assembled sequence data for 73 protein-coding plastid genes from 132 samples representing 68 species (~23% of described species) that span the breadth of most major taxonomic, morphological, and biogeographic groups of Podostemaceae. With these data, we conducted the first plastid phylogenomic analysis of the family with broad taxon sampling. These analyses resolved most nodes with high support, including relationships not recovered in previous analyses. No evidence of widespread, well-supported conflict among individual plastid genes and the concatenated phylogeny was observed. We present new evidence that four genera (<em>Apinagia</em>, <em>Marathrum</em>, <em>Oserya</em>, and <em>Podostemum</em>), as well as four species, are not monophyletic. In particular, we show that <em>Podostemum flagelliforme</em> should not be included in <em>Podostemum and is better recognized as Devillea flagelliformis, </em>and that <em>Marathrum capillaceum</em> is embedded within <em>Lophogyne </em>s.l.<em> </em>and should be recognized as <em>Lophogyne capillacea</em>. We also place a previously unsampled and undescribed species that likely represents a new genus. In contrast to previous studies, the neotropical genera <em>Diamantina</em>, <em>Ceratolacis</em>, <em>Cipoia,</em> and <em>Podostemum</em> are resolved as successive sister groups to a clade of all paleotropical Podostemoideae taxa sampled, suggesting a single dispersal event from the neotropics to the paleotropics in the history of the subfamily. These results provide a strong basis for improving the classification of Podostemaceae and a framework for future phylogenomic studies of the clade employing data from the nuclear genome.</p>
Improving Volcanic SO2 Cloud Modeling Through Data Fusion and Trajectory Analysis: A Case Study of 2022 Hunga Tonga Eruption
<p><strong>Dataset Overview</strong>: This dataset comprises approximately 500 clusters of aggregated observational data collected from January 16 to 20 during the ascending (ASC) and descending (DES) periods. We grouped a large number of observation points into these clusters and calculated trajectories from the center of each cluster. The choice of 500 clusters was driven by pragmatic considerations, aiming for a balance between computational feasibility and the level of detail needed for our analysis.</p> <p><strong>Data Unit Description</strong>: The "mass" values in this dataset for each cluster are calculated by multiplying the mass per unit area (<span><span>g/m2</span></span>) of individual data points by the area covered by each point, thus providing the total mass in grams (g). The "heights" are presented in units of kilometers (km), representing the observed top heights of each cluster.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.