Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Landscape composition and life-history traits influence bat movement and space use: analysis of 30 years of published telemetry data
<p>Using temperate bats, a group of particular conservation concern, we investigated how morphological traits, habitat specialization and environmental variables affect home range sizes and daily foraging movements, using a compilation of 30 years of published bat telemetry data in Northern America and Europe for the period 1988 – 2016.</p> <p>We compiled data on home range size and mean daily distance between roosts and foraging areas at both colony and individual levels from 166 studies of 3,129 radiotracked individuals of 49 bat species. We calculated multi-scale habitat composition and configuration in the surrounding landscapes of all studied roosts. Using mixed models, we examined the effects of habitat availability and spatial arrangement on bat movements, while accounting for body mass, aspect ratio, wing loading and habitat specialization.</p> <p>We found a significant effect of landscape composition on home range size and mean daily distance at both colony and individual levels. On average, home ranges were up to 42% smaller in the most habitat-diversified landscapes while mean daily distances were up to 30% shorter in the most forested landscapes. Bat home range size significantly increased with body mass, wing aspect ratio and wing loading, and decreased with habitat specialization.</p>
UK Power Station Transformer Dissolved Gas Analysis Data (2010-2015)
<p>This dataset includes dissolved gas analysis records from coolant oil in 13 UK power station transformers for various timespans between 2010-2015. They form the basis of the paper "Assessing the impact of weak and moderate geomagnetic storms on UK power station transformers" submitted to the AGU "Space Weather" Journal by Z.M. Lewis, J.A. Wild and M. Allcock in December 2021.<br> <br> Please cite Lewis et al. if using these data. The authors thank D. Barker, EDF Energy Nuclear Generation, for providing these data.</p> <p> </p> <p>The data are presented in comma separated variable format files, as described in the readme.txt file.</p>
Supplementary Data for "Sequencing the Pandemic: Rapid and High-Throughput Processing and Analysis of COVID-19 Clinical Samples for 21st Century Public Health"
<p>Supplementary material for F1000 methods manuscript. Includes raw sequencing metrics for two COVID sequencing methodologies, as well as a complete cost breakdown for each methodology.</p>
Application of Mass Multivariate Analysis on Neuroimaging Data Sets for Precision Diagnosis of Depression: Data and code
<p><strong>In the archive are the datasets and code used to create the results in article "Application of Mass Multivariate Analysis on Neuroimaging Data Sets fo Precision Diagnosis of Depression". The methods is based on the Multivariate Linear toolbox mad F. Kherif’s lab and available on GitHub. <a href="https://github.com/LREN-CHUV/MLM">https://github.com/LREN-CHUV/MLM</a>. </strong></p> <p><strong>A total of 44 patients with a current psychotic (n=19) or depressive (n=25) episode were analyzed by MLM (details can be found in the article itself) utilizing resting sate fMRI, tasks fMRI, and anatomical T1w images. </strong></p> <p><strong>The data results are provided in MATLAB format, in a matrix called MLM.mat. The matrix includes the canonical variables for each of the three imaging modalities, chi-square statistics, and p-values for each brain region. Additionally, we provided a function MLM_plot_fun that allows mapping the statistics results to a canonical three-dimensional brain as per the article. </strong></p>
ExoCAM: A 3D Climate Model for Exoplanet Atmospheres :: Model data and supplementary figures and analysis
<p>This repository contains 3D GCM model output data from the paper, "ExoCAM: A 3D Climate Model for Exoplanet Atmospheres", which is published in the Planetary Science Journal: Trapppist Habitable Atmospheres Intercomparison Special Issue. The model data includes mean climate states for the standard THAI simulations of TRAPPIST-1e, simulations using an upgraded radiative transfer, along with a large variety sensitivity experiments considering common tuning parameters of sub-grid scale cloud and convection physics. In total 43 simulations are included. Also included here are a variety of multi-panel contour plots showing basic results from all simulations as supplemental figures.</p> <p>https://iopscience.iop.org/article/10.3847/PSJ/ac3f3d</p>
Datasets for Linked Open Data Instance Level Analysis for Cultural Heritage
<p>This is the datasets used for Linked Open Data instant level quality analysis for cultural heritage (2020). 7Z and ZIP versions are available for both Excel 2006 and R 4.0.3. The compressed files include, Excel spreadsheets (.xlsx, .csv), VBA scripts (.bas), and R scripts (.r).</p> <p>Please read the full documentation in Linked_Open_Data_Instance_Level_Analysis_Procedure.pdf.</p>
[Data from:] Genetic Analysis Reveals Three Novel QTLs Underpinning a Butterfly Egg-Induced Hypersensitive Response-Like Cell Death in Brassica Rapa
<p><strong>Background</strong></p> <p>Cabbage white butterflies (<em>Pieris</em> spp.) can be severe pests of <em>Brassica</em> crops such as Chinese cabbage, Pak choi (<em>Brassica rapa</em>) or cabbages (<em>B. oleracea</em>). Eggs of <em>Pieris</em> spp. can induce a hypersensitive response-like (HR-like) cell death which reduces egg survival in the wild black mustard (<em>B. nigra</em>). Unravelling the genetic basis of this egg-killing trait in <em>Brassica</em> crops could improve crop resistance to herbivory, reducing major crop losses and pesticides use. Here we investigated the genetic architecture of a HR-like cell death induced by <em>P. brassicae</em> eggs in <em>B. rapa.</em></p> <p><strong>Results</strong></p> <p>A germplasm screening of <em>B. rapa</em> 56 accessions, representing the genetic and geographical diversity of a <em>B. rapa</em> core collection, showed phenotypic variation for cell death. An image-based phenotyping protocol was developed to accurately measure size of HR-like cell death and was then used to identify two accessions that consistently showed weak (R-o-18) or strong cell death response (L58). Screening of 160 RILs derived from these two accessions resulted in three novel QTLs for P<em>ieris</em> b<em>rassicae-</em>induced cell death on chromosomes A02 (<em>Pbc1</em>), A03 (<em>Pbc2</em>), and A06 (<em>Pbc3</em>). The three QTLs <em>Pbc1-3</em> contain cell surface receptors, intracellular receptors and other genes involved in plant immunity processes, such as ROS accumulation and cell death formation. Synteny analysis with <em>A. thaliana</em> suggested that <em>Pbc1</em> and <em>Pbc2</em> are novel QTLs associated with this trait, while <em>Pbc3</em> contains also LecRK-I.1, a gene of <em>A. thaliana</em> previously associated with cell death induced by a <em>P. brassicae</em> egg extract.</p> <p><strong>Conclusions</strong></p> <p>This study provides the first genomic regions associated with the <em>Pieris</em> egg-induced HR-like cell death in a <em>Brassica</em> crop species. It is a step closer towards unravelling the genetic basis of an egg-killing crop resistance trait, paving the way for breeders to further fine-map and validate candidate genes.</p>
Output Data of the FlexMex Analysis Tool
<p>This dataset includes the results of the FlexMex model comparison. It relies in the input dataset published under https://doi.org/10.5281/zenodo.5802178 and the application of the FlexMex Analysis Tool published under https://doi.org/10.5281/zenodo.6010392.</p> <p>For further details please refer to the included file "HowTo_FlexMexAnalysisTool.pdf".</p> <p>The FlexMex project received funding from the German Federal Ministry of Economic Affairs and Climate Action (BMWK), formerly Federal Ministry of Economic Affairs and Energy (BMWi), under grant number 03ET4077A-H.</p>
Data and Analysis for "On the Reliability of Coverage-based Fuzzer Benchmarking"
<pre><strong>Data and Analysis for "On the Reliability of Coverage-based Fuzzer Benchmarking"</strong> <strong>## Cite</strong> </pre> <pre><code>@inproceedings{benchmarking, author = {B{\"o}hme, Marcel and Szekeres, L{\'a}szl{\'o} and Metzman, Jonathan}, title = {On the Reliability of Coverage-based Fuzzer Benchmarking}, year = {2022}, booktitle = {Proceedings of the 44th International Conference on Software Engineering}, series = {ICSE '22}, pages = {1-13}, doi = {10.1145/3510003.3510230} }</code></pre> <pre> <strong>## Data Analysis</strong> The Jupyter notebook generating all tables and figures can be found in fuzzbench.manual.ipynb <strong>## Generated Images and Tables</strong> The generated data analysis artifacts are also available in this artifact. <strong>## Data</strong> All the data is available in the FuzzBench Reports and will be automatically downloaded. * 20 trials of 23 hours with 15 programs and 10 fuzzers. * Experiment name: 2021-02-17-bug-paper * Report: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/index.html * Data: https://www.fuzzbench.com/reports/2021-02-17-bug-paper/data.csv.gz * Fuzzbench Commit: [38e344fef2f1079579391a0d9dcb52319f7051f2](https://github.com/google/fuzzbench/commits/38e344fef2f1079579391a0d9dcb52319f7051f2) * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b * 30 trials of 23 hours with 11 programs and 10 fuzzers. * Experiment name: 2021-08-19-crash-s2 * Report: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/index.html and * Data: https://www.fuzzbench.com/reports/2021-08-19-crash-s2/data.csv.gz * Fuzzbench Commit: db192b60815ac87f69ee0f7f3e37aeac71949e1b The deduplicated data can be found in * 2021-02-17-bug-paper-fixed2.csv.gz * 2021-08-19-crash-s-fixed2.csv.gz * 2021-08-19-crash-s2-fixed2.csv.gz <strong>## Reproducibility</strong> </pre> <pre><code class="language-bash"># Download the precise version of FuzzBench used for the experiment git clone https://github.com/google/fuzzbench.git cd fuzzbench git checkout <Fuzzbench Commit> # Download the internal config file. curl https://storage.googleapis.com/[experiment-name]/config/experiment.yaml > /tmp/experiment-config.yaml make install-dependencies # Launch the experiment using paramters from the internal config file. PYTHONPATH=. python experiment/reproduce_experiment.py -c /tmp/experiment-config.yaml -e <new_experiment_name></code></pre> <p> </p>
NetCDF data used in analysis presented in "Assessment of the z~ time-filtered Arbitrary Lagrangian-Eulerian coordinate in a global eddy-permitting ocean model"
<p>NetCDF data used in analysis presented in "Assessment of the z~ time-filtered Arbitrary Lagrangian-Eulerian coordinate in a global eddy-permitting ocean model", submitted to Journal of Advances in Modelling the Earth System.</p> <p>The data are produced from an ensemble of six experiments based on the GO8p0 configuration of NEMO v4.0.1 on a global 1/4° grid, as described in the paper. The ensemble is intended to test the z~ vertical coordinate, and includes a control with the default "z-star" fixed coordinate, and five experiments with the z-tilde vertical coordinate, using a selection of values for the two z-tilde timescale parameters. The data includes time series of global mean ocean and ice fields; large-scale transports; and fields from diapycnal mixing analysis.</p> <p>The first part of each filename refers to the experiment from the ensemble ("zstar", "ztilde_5_30", "ztilde_10_30", "ztilde_20_30", ztilde_20_60" and "ztilde_40_60"); the following five-character string identifies the respective suite on the Met Office Rose system and the MASS archive system; and the rest of the name specifies the type of data contained in the file.</p>
Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).
<p>*******************************************************************<br> MetaDrugs workflow<br> *******************************************************************</p> <p>Data analysis pipeline for investigating drug-host-microbiome relationships in cardiometabolic disease (MetaCardis cohort).</p> <p>For questions and requests, please contact:<br> Sofia K. Forslund (sofia.forslund@mdc-berlin.de)<br> and Till Birkner (till.birkner@mdc-berlin.de)</p> <p>*******************************************************************<br> Contents:<br> -------------------------------------------------------------------<br> Data files:<br> metadata.tar.gz - archived cohort metadata files*<br> input_features.tar.gz - archived preprocessed serum and urine metabolome and gut microbiome features<br> output_complete.tar.gz - archived example analysis output files for each of the input feature file<br> output_rerun.tar.gz - archived empty directory for generating test output files as described in this document<br> <br> *Please note: Due to conflicts with Danish Data Protection laws, metadata from the Danish subset of the cohort were removed in this repository. Please reach out for a potential case-by-case access request for access to the complete set of metadata.<br> -------------------------------------------------------------------<br> Text files:<br> archived in feature_names.tar.gz:<br> atcs_names - full names for atcs drug compounds<br> contrast_names - full names for disease comparison groups<br> file_names - brief description of the files in input_features folder<br> gmm_names - full names of GMM modules<br> kegg_names - full names of KEGG modules<br> ko_names - full names of KO modules<br> metadata_names - full names of metadata features<br> mOTU_names - species names for metagenomics data<br> taxon_names - taxon names for metagenomics data<br> -------------------------------------------------------------------<br> Scripts:<br> -------------------------------------------------------------------<br> runFrame.r - main wrapper script envoking the analysis pipeline<br> -------------------------------------------------------------------<br> runFrame_rel_comb.r - script calculating drug combination effects<br> runFrame_rel.r - script calculating dosage effects<br> testCombPresenceSeparate.r - testing of significant drug combination effects beyond single drug effects<br> testDosagePresenceSeparate.pl - testing of significant drug dosage effects beyond single drug effects<br> testDosagePresenceSeparateNegative.pl - testing of unique drug dosage effects beyond single drug effects<br> -------------------------------------------------------------------<br> prettifyResults_uncollapsed.pl - wrapper scripts to create and format a single analysis output file<br> makeTables.r - wrapper script to make excel tables with analysis results<br> -------------------------------------------------------------------<br> Example output file:<br> -------------------------------------------------------------------<br> output_all_formatted_noc_uncollapsed_complete.tsv - contains all disease-drug-host-microbiome feature analysis results in one place.<br> *******************************************************************</p>
spectrapepper: A Python toolbox for advanced analysis of spectroscopic data for materials and devices.
<p>spectrapepper is a Python package that makes advanced analysis of spectroscopic data easy and accessible through straightforward, simple, and intuitive code. This library contains functions for every stage of spectroscopic methodologies, including data acquisition, pre-processing, processing, and analysis. In particular, advanced and high statistic methods are intended to facilitate, namely combinatorial analysis and machine learning, allowing also fast and automated traditional methods. The following is a short list of some main procedures that spectrapepper package enables: i) Baseline removal functions, ii) Normalization methods, iii) Noise filters, trimming tools, and despiking methods, iv) Chemometric algorithms to find peaks, fit curves, and deconvolution of spectra, v) Combinatorial analysis tools, such as Spearman, Pearson, and n-dimensional correlation coefficients, vi) Tools for Machine Learning applications, such as data merging, randomization, and decision boundaries, and vii) Sample data and examples</p>
Data for Deciphering-the-CO2-emissions-and-emission-intensity-of-cement-sector-in-China-through-decomposition-analysis
<p>The dataset contains data for Figure 7-13 in our article "<em>Deciphering the CO<sub>2</sub> emissions and emission intensity of cement sector in China through decomposition analysis</em>", and data for part of<em> China Cement Industry Dataset (CCID)</em>. </p>
Supplementary data to the Baseline Methodological User Needs Analysis
<p>Supplementary data to the Baseline Methodological User Needs Analysis</p>
Supplementary data for article "Small hydropower – small ecological footprint? A multi-annual environmental impact analysis using aquatic macroinvertebrates as bioindicators. Part 2: effects on functional diversity" by Scotti A., et al.
<p>Supplementary data for article "Small hydropower – small ecological footprint? A multi-annual environmental impact analysis using aquatic macroinvertebrates as bioindicators. Part 2: effects on functional diversity" by Scotti A., et al.:</p> <p><br> - Trait-based distances calculated for each pair of taxa;</p> <p>- CWM, CWM(LN) values, and their difference (CWMDIFF)</p> <p>Refer to the published articles for further details.</p>
Supplementary GIS data - Potential and implications of automated pre-processing of LiDAR-based digital elevation models for large-scale archaeological landscape analysis
<p>A supplementary dataset related to the paper discussing preparation of a digital elevation model derived from DMR 5G (LiDAR-based DEM of the Czech Republic) cleaned of modern artificial features. It includes data used as a clipping mask and data produced during the testing phase.</p> <p>Contents:</p> <ul> <li>..\clipping_buffers.gdb\ - Clipping buffers based on ZABAGED dataset used for masking the original data stored as ESRI geodatabase.</li> <li>..\drainages\ - Drainages with Strahler order higher than four (potential watercourses) for the original and filtered DEMs. <ul> <li>drainages_filtered - Drainges identified in the filtered DEM stored as GeoTIFF.</li> <li>drainages_original - Drainges identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\LSC\ - Locations with significant land surface curvature for the original and filtered DEMs. <ul> <li>LSC_filtered - Significant LSC identified in the filtered DEM stored as GeoTIFF. </li> <li>LSC_original - Significant LSC identified in the original DEM stored as GeoTIFF. </li> </ul> </li> <li>..\visibility\ - Viewsheds computed over the original and filtered DEMs. <ul> <li>Libice\ - Sample viewsheds computed for the early medieval hillfort of Libice. <ul> <li>Libice_visibility_filtered - Viewshed based on the filtered DEM stored as GeoTIFF. </li> <li>Libice_visibility_original - Viewshed based on the original DEM stored as GeoTIFF. </li> <li>observer_points - Observer points used for calculating the viewsheds.</li> </ul> </li> <li>regular_grid\ - Cumulative viewsheds calculated for regularly spaced points in a 10 x 10 km grid with a visibility radius of 5 km and an observer height of 2 m; a total of 574 viewsheds. <ul> <li>visibility_filtered - Cumulative viewshed for the filtered DEM stored as GeoTIFF.</li> <li>visibility_original - Cumulative viewshed for the original DEM stored as GeoTIFF. </li> <li>visibility_test_buffers - Buffers used for the viewshed calculations stored as ESRI shapefile.</li> <li>visibility_test_observers - Observer points used for the viewshed calculations stored as ESRI shapefile.</li> </ul> </li> </ul> </li> </ul> <p> </p> <p>Preprint version of the related paper:</p> <p>Novák, David and Pružinec, Filip, Potential and Implications of Automated Pre-Processing of Lidar-Based Digital Elevation Models for Large-Scale Archaeological Landscape Analysis. Available at SSRN: <a href="https://ssrn.com/abstract=4063514">https://ssrn.com/abstract=4063514</a></p>
Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity"
<p>Source data to publication "Benchmarking of Analysis Strategies for Data-Independent Acquisition Proteomics Using a Large-Scale Dataset Comprising Inter-Patient Heterogeneity".</p> <p>Data and further information at GitHub repository https://github.com/kreutz-lab/dia-benchmarking (DOI: 10.5281/zenodo.6371925)</p>
Data for analysis procedure for rectenna-based THz-spectroscopy as published in Lechelon et al., Sci. Adv. 8, eabl5855 (2022)
<p>Example of data for R-PE data processing </p>
Computational Analysis of Two-dimensional High-throughput Data from Large-scale RNAi Screens and Single-cell Transcriptomics
<p>This publication provides a singularity definition file to reproduce the computational environment along with the scripts to reproduce every figure or table in the revised manuscript using ZetaSuite Perl module and R package.</p> <p>First, generate a new folder and then download all the files into the folder.</p> <p>Then, uncompressed the files DataSets_part1.tar.gz,DataSets_part2.tar.gz,DataSets_part3.tar.gz,DataSets_part4.tar.gz, and scripts.tar.gz. within the folder.</p> <p>Next, move all the files in DataSets_part1 folder, DataSets_part2 folder,DataSets_part3 folder and DataSets_part4 folder to a new folder called DataSets.</p> <p>Finally, run the following scripts to generate the figures and tables in our manuscript.</p> <p>Regeneration of Figure2 and S2: singularity exec ZetaSuite.sif sh Figure2andS2.sh </p> <p>Regeneration of Figure3 and S3: singularity exec ZetaSuite.sif sh Figure3andS3.sh </p> <p>Regeneration of Figure4 and S4: singularity exec ZetaSuite.sif sh Figure4andS4.sh </p> <p>Regeneration of Figure5 and S5: singularity exec ZetaSuite.sif sh Figure5andS5.sh </p> <p>Regeneration of Figure6 and S6: singularity exec ZetaSuite.sif sh Figure6andS6.sh </p> <p>Regeneration of Figure7 and S7: singularity exec ZetaSuite.sif sh Figure7andS7.sh </p> <p> </p>
Live cell microscopy: From image to insight - raw data & analysis
<p>Accompanying raw and processed data as well as analysis scripts for the publication Biophysics Rev. 3, (2022); <a href="https://doi.org/10.1063/5.0082799">10.1063/5.0082799</a> "Live cell microscopy: From image to insight".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.