Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4,694
datasets available to search
ShareScore release 0.7.1
Dataset results
4,694 results for “data analysis”
Extracted data for meta-analysis: Distal versus proximal radial access in coronary angiography
<p>Data base underlying quantitative meta-analysis in the manuscript titled "Distal versus proximal radial access in coronary angiography: A meta-analysis" by Lueg, Schulze, Stöhr, & Leistner.</p> <p>Data allows meta-analysis of primary endpoint (RAO), secondary endpoints, meta-regression for moderator analysis, and publication bias analysis.</p>
Data files for analysis of scaling relations between relative and absolute humidity and rainfall extremes
<p>data belonging to: https://github.com/mister-superCC/CCscaling-Evaluation</p> <p>Contains:</p> <p>Dutch observarions in netcdf: KNMI_20201124_hourly.nc</p> <p>Model data as an R object file: DATA_SCALING_PRINCIPLES.tar.gz</p> <p>Processed model and observational data (including SFR) in: data_hourly_bootstrap.tar.gz and data_hourly_bootstrap_abs.tar.gz</p>
Data set for spectral and FT-ICR-MS analysis of dissolved organic matter in sediments of the New Britain Trench axis station
<p> </p> <p>数据集包括原始数据和简单处理所需的数据、图表和文章。</p> <p> </p> <p>表格包含样品数量、收集深度、样品名称以及相应光谱和 FT-ICR-MS 数据的名称。</p> <p> </p> <p>例如,光谱数据包括荧光相对强度、荧光指数,而 FT-ICR-MS 光谱数据包括化合物类型相对强度以及化学式、元素比、等效双键数 (DBE) 和每个样品的芳香指数 (AImod)。</p> <p> </p>
Changes in holopelagic Sargassum spp. biomass composition across an unusual year - supporting particle data and analysis source code
<p>As specified in 'Data, Materials, and Software Availability' of Tonon et al. (2024), PNAS, Vol. 121, e2312173121, https://doi.org/10.1073/pnas.2312173121, the following data and software are provided:</p> <p>Particle forward tracking (statistical data – for Fig. 2)</p> <p>Particle backward tracking (primary data – for Fig. 3)</p> <p>Analysis of primary data for gridded particle fractional coverage and mean age (Fortran source code)</p>
LAM_HTGTS data analysis
<p>These scripts are associated with the manuscript "ATM and 53BP1 regulate alternative end joining-mediated V(D)J recombination" to analysis the data available in the database Gene Expression Omnibus (GEO) under the ascension numbers GSE242952, GSE246239 and GSE263165.</p>
Data from: Estimates of late Early Cretaceous atmospheric CO2 from Mongolia based on stomatal and isotopic analysis of Pseudotorellia
<p>Our dataset includes 115 leaf cuticles related to two species of <em>Pseudotorellia</em> Florin from three stratigraphically similar samples at the Tevshiin Govi lignite mine in central Mongolia (~119.7–100.5 Ma, Aptian–Albian, Cretaceous). We apply a well-vetted paleo-CO<sub>2</sub> proxy based on leaf gas-exchange principles (the Franks model) to those leaves for paleo-CO<sub>2</sub> reconstruction, which requires leaf stomata and carbon isotope analysis. All cuticle measurements are summarized in this dataset. </p>
Data from: Tamm Review: a meta-analysis of thinning, prescribed fire, and wildfire effects on subsequent wildfire severity in conifer dominated forests of the Western US
<p>Increased understanding of how active forest management (i.e., mechanical thinning, prescribed burning, and managed wildfire) affects subsequent wildfire severity is urgently needed as people and forests face a growing wildfire crisis. In response, we reviewed scientific literature for the US West and completed a meta-analysis that answered three questions: (1) How much do treatments reduce wildfire severity within treated areas? (2) How do the effects vary with treatment type, treatment age, and forest type? (3) How does fire weather moderate the effects of treatments? We found overwhelming evidence that mechanical thinning with prescribed burning, mechanical thinning with pile burning, and prescribed burning only are effective at reducing subsequent wildfire severity, resulting in reductions in severity from 62% to 72% relative to untreated areas. In comparison, thinning only was less effective – underscoring the importance of treating surface fuels when mitigating wildfire severity is the management goal. The efficacy of these treatments did not vary among forest types assessed in this study and was high across a range of fire weather conditions. Prior wildfire had more complex impacts on subsequent wildfire severity, which varied with forest type and initial wildfire severity. Across treatment types, we found that effectiveness of treatments declined over time, with the mean reduction in wildfire severity decreasing nearly threefold when wildfire occurred greater than 10 years after initial treatment. Our meta-analysis provides up-to-date information on the extent to which active forest management reduces wildfire severity and facilitates better outcomes for people and forests during future wildfire events. </p>
Gene Enrichment Map Data from gProfiler Analysis - Selected MPK Interactions of Arabidopsis thaliana
<p>Gene enrichment analysis results for the selected predicted MPK interactions are included in the supplementary materials.</p>
Changes Monitoring in Hongjiannao Lake from 1987-2023 using Google Earth Engine and Analysis of Climatic and Anthropogenic Forces (Climatic Data)
<p>This dataset presents temporal (1987 to 2023) climatic data for the weather station near Hongjiannao Lake.</p>
Raw Bibliobmetric Data for the article "The Scientific Landscape of Phytoremediation of Tailings: A Bibliometric and Scientometric Analysis"
<p>Bibliometric data for CiteSpace related to the scientific article "The Scientific Landscape of Phytoremediation of Tailings: A Bibliometric and Scientometric Analysis".</p>
Data from: A 3D geometric morphometric analysis of the bovid distal humerus, with special reference to Rusingoryx atopocranion (Pleistocene, Eastern Africa)
<p>The family Bovidae [Mammalia: Artiodactyla] is speciose and has extant representatives on every continent, forming key components of mammal communities. For these reasons, bovids are ideal candidates for studies of ecomorphology. In particular, the morphology of the bovid humerus has been identified as highly related to functional variables such as body mass and habitat. This study investigates the functional morphology of the bovid distal humerus in isolation due to its increased likelihood of preservation in the fossil record, and the resulting opportunity for better understanding the ecomorphology of extinct bovids. A landmark scheme of 30 landmarks was used to capture the 3D distal humerus morphology in 111 extant bovid specimens. We find that the distal humerus has identifiable morphologies associated with body mass, habitat preference, and tribe affiliation, and that some characteristics are shared between high body mass bovids and those living on hard, flat terrain which is likely due to the high stress on the bone in both cases. We directly apply our findings regarding extant bovids to the extinct alcelaphine bovid, <em>Rusingoryx</em> <em>atopocranion</em> from the mid to late Pleistocene (>33-45 ka) Lake Victoria region of Kenya. This species is known for some peculiar morphologies including a domed cranium with hollow nasal crests, and having small hooves for a bovid of its size. Another interesting aspect of <em>Rusingoryx</em>'s skeletal morphology which has not been addressed is an unusual protrusion on the lateral epicondyle of the distal humerus. Despite considerable individual variation in the <em>Rusingoryx</em> specimens, we find evidence to support its historical assignment to the tribe Alcelaphini, and that it likely preferred open grassland habitats, which is consistent with independent reconstructions of the paleoenvironment. We also provide the most accurate body mass estimate for <em>Rusingoryx</em> to date, based on distal humerus centroid size. Overall, we are able to conclude that the distal humerus in extant bovids is highly informative regarding body mass, habitat preference and tribe, and that this can be applied directly to a fossil taxon with promising results.</p>
Data and analysis scripts for the submission "Seamlessly Scaling Applications with DAPHNE"
<p>To regenerate the plots:</p> <p>You will need to install `R` (4.3.2) and the `tidyverse` package (2.0.0)</p> <p>We give a Nix flake that captures this environment (Install Nix: https://nixos.org/download/ and activate the flake feature: https://nixos.wiki/wiki/Flakes#Other_Distros.2C_without_Home-Manager)</p> <p>With Nix: `nix develop --command Rscript analysis_compas.R`</p> <p>Without Nix: `Rscript analysis_compas.R`</p> <p> </p>
Analysis Data, "Strain dynamics of contaminating bacteria modulate the yield of ethanol biorefineries"
<p>This package contains datasets in `Rdata` format underlying analyses presented in the study "Strain dynamics of contaminating bacteria modulate the yield of ethanol biorefineries", first available as a preprint on February 08, 2021:</p> <p><a href="https://www.biorxiv.org/content/10.1101/2021.02.07.430133v1.article-info">https://www.biorxiv.org/content/10.1101/2021.02.07.430133v1.article-info</a></p>
Data from: Trophic guilds differ in blood glucose concentrations: A phylogenetic comparative analysis in birds
<p>Glucose is a central metabolic compound used as a source of energy across all animal taxa. There is high interspecific variation in glucose concentration between taxa, the origin and the consequence of which remain largely unknown. Nutrition may affect glucose concentrations because carbohydrate content of different food sources may determine the importance of metabolic pathways in the organism. Birds sustain high glucose concentrations that may entail the risks of oxidative damage. We collected glucose concentration and life history data from 202 bird species from 171 scientific publications; classified them into seven trophic guilds and analysed the data with a phylogenetically controlled model. We show that glucose concentration is negatively associated with body weight and is significantly associated with trophic guilds with a moderate phylogenetic signal. After controlling for allometry, glucose concentrations were highest in carnivorous birds, which rely on high rates of gluconeogenesis to maintain their glycemia and lowest in frugivorous/nectarivorous species, which intake carbohydrates directly. However, trophic guilds with different glucose concentrations did not differ in lifespan. These results link nutritional ecology to physiology and suggest that at the macroevolutionary scale, species requiring constantly elevated glucose concentrations may have additional adaptations to avoid the risks associated with high glycemia.</p>
Data from: Targeted genotyping-by-sequencing of potato and data analysis with R/polyBreedR
<p>"Mid-density" targeted genotyping-by-sequencing (GBS) combines trait-specific markers with thousands of genomic markers at an attractive price for linkage mapping and genomic selection. A 2.5K targeted GBS assay for potato was developed using the DArTag<sup>TM</sup> technology and later expanded to 4K targets. Genomic markers were selected from the potato Infinium<sup>TM</sup> SNP array to maximize genome coverage and polymorphism rates. The DArTag and SNP array platforms produced equivalent dendrograms in a test set of 298 tetraploid samples, and 83% of the common markers showed good quantitative agreement, with RMSE (root-mean-squared-error) less than 0.5. DArTag is suited for genomic selection candidates in the clonal evaluation trial, coupled with imputation to a higher-density platform for the training population. Using the software polyBreedR, an R package for the manipulation and analysis of polyploid marker data, the RMSE for imputation by linkage analysis was 0.15 in a small half-diallel population (N=85), which was significantly lower than the RMSE of 0.42 with the Random Forest method. Regarding high-value traits, the DArTag markers for resistance to potato virus Y, golden cyst nematode, and potato wart appeared to track their targets successfully, as did multi-allelic markers for maturity and tuber shape. In summary, the potato DArTag assay is a transformative and publicly available technology for potato breeding and genetics.</p>
Simulated RNA-seq data for differential splicing analysis with covariates
<p>The repository includes alignments of simulated RNA-seq data for evaluating differential splicing detection with covariates. Starting from an empirical transcript expression matrix trained on an RNA-seq data set from lung fibroblasts (GenBank A# SRR493366) and using GENCODE v.41 as reference, 11.5 million 100 bp long paired-end reads were generated per sample, from 2,000 genes with two or more expressed isoforms. RNA-seq data was simulated for one ‘condition’, with values ‘control’, ‘disease’ and ‘stage2’, with one covariate, ‘biological sex’, with values ‘M’ and ‘F’. 5 samples each were simulated for each (condition x sex) category. Changes were simulated in the expression (DE) and/or the splicing ratio (DS) of genes as follows. Changes in expression (DE) were simulated by either halving or doubling the expression level of the gene. Changes in splicing ratios (DS) were simulated by swapping the expression levels of the gene’s top two transcript isoforms. All RNA-seq data was mapped to the hg38 genome with the spliced alignment tool STAR v2.7.10a.</p> <p> <em><u>Pairwise comparison alignment set</u></em>: Differences due to ‘condition’ between two states, ‘control’ and ‘disease’, were simulated at 600 genes, including 200 DE, 200 DS and 200 DE+DS genes. Differences in ‘biological sex’ (covariate) were represented as changes in 300 genes, including 100 from each of the DS, DE and DE+DS categories. Hence, the target gene set for differential splicing ratio (DSR)<em> pairwise comparisons </em>consists of the pooled 200 DS and 200 DS+DE genes differentially spliced between the ‘control’ and ‘disease’ states, while for differential splicing abundance (DSA)<em> pairwise comparisons </em>the target gene set is the set of 600 modified genes, 200 in each of the DS, DE and DS+DE categories.</p> <p> <em><u>Multiway (3-way) comparison alignment set:</u></em> Additional changes between ‘disease’ and ‘stage2’ were made to 100 of the previously modified genes, as well as to a set of 200 additional genes not encountered previously, for each of the categories DE, DS and DE+DS. Therefore, for <em>DSR three-way comparisons</em>, the target gene set represents the 800 genes simulated as being DS or DE+DS between any of the ‘control’, ‘disease’ and ’stage2’ categories, while for the <em>multi-way DSA comparisons</em> the target is the full set of 1,200 genes (400 DE, 400 DS and 400 DE+DS) simulated to have changed between any of the 'control', 'disease’ and ‘stage2’ states.</p> <p> <em><u>Further details:</u></em> See the ‘key’ directories in each package for the gene lists.</p>
Code for data analysis - intra-community variability of leaf-out in temperate tree canopies
<p>The code and data were used for producing the results of "Phenology across scales: an intercontinental analysis of leaf-out dates in temperate deciduous tree communities", by Delpierre et al.</p>
Data underlying the article: "Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling"
<p>This repository contains the data underlying the article: “Excuse me, there is a mutant in my bioactivity soup! A comprehensive analysis of the genetic variability landscape of bioactivity databases and its effect on activity modelling” available as a preprint on ChemRxiv.</p> <p>Main authors: Marina Gorostiola González & Olivier J.M. Béquignon (Leiden University)</p> <p>Senior author: Gerard J.P. van Westen (Leiden University)</p> <p>This analysis was performed using the code available at <a href="https://github.com/CDDLeiden/chembl_variants" target="_blank" rel="noopener">https://github.com/CDDLeiden/chembl_variants</a></p>
Survey Data for Multicriteria Satisfaction Analysis of Cargo Bike Last-Mile Delivery in European Cities
<p>SSH CENTRE (Social Sciences and Humanities for Climate, Energy aNd Transport Research Excellence) is a Horizon Europe project, engaging directly with stakeholders across research, policy, and business (including citizens) to strengthen social innovation, SSH-STEM collaboration, transdisciplinary policy advice, inclusive engagement, and SSH communities across Europe, accelerating the EU’s transition to carbon neutrality. <br>SSH CENTRE is based in a range of activities related to Open Science, inclusivity and diversity – especially with regards Southern and Eastern Europe and different career stages – including: development of novel SSH-STEM collaborations to facilitate the delivery of the EU Green Deal; SSH knowledge brokerage to support regions in transition; and the effective design of strategies for citizen engagement in EU R&I activities. Outputs include action-led agendas and building stakeholder synergies through regular Policy Insight events.<br>This is captured in a high-profile virtual SSH CENTRE generating and sharing best practice for SSH policy advice, overcoming fragmentation to accelerate the EU’s journey to a sustainable future.<br>The documents uploaded here are part of WP2 whereby novel, interdisciplinary teams were provided funding to undertake activities to develop a policy recommendation related to EU Green Deal policy. Each of these policy recommendations, and the activities that inform them, will be written-up as a chapter in an edited book collection. Three books will make up this edited collection - one on climate, one on energy and one on mobility. <br>As part of writing a chapter for the SSH CENTRE book on ‘Strengthening European mobility policy - Governance recommendations from innovative interdisciplinary collaborations’, we elicit the opinions of citizens in urban logistics policymaking through a series of surveys in different European cities. The files attached to this Zenodo webpage are therefore the dataset contains raw survey data from a study utilizing Multicriteria Satisfaction Analysis (MUSA) to evaluate public perceptions of cargo bike last-mile delivery in London, Paris, Rome, Dublin, and Warsaw. The data encompasses over 2,000 responses, detailing participants' satisfaction levels with various aspects of cargo bike delivery services, including CO2 emissions, noise, traffic, safety, and shipping costs. This dataset supports comprehensive analyses of urban logistics policies aimed at sustainable mobility solutions in these specific cities.</p>
Data from: Analysis of genotyping data reveals the unique genetic diversity represented by the breeds of sheep native to the United Kingdom
<p><strong>Background: </strong>Sheep breeds native to the United Kingdom are noted for high breed variability and exhibit a striking diversity of different traits in phenotypes and genetic diversity. Some of these traits are highly sustainable, such as seasonal wool shedding in the Wiltshire Horn, are likely to become more important as pressures on sheep production increase in coming decades. Despite their clear importance to the future of sheep farming, the genetic diversity of native UK sheep breeds is poorly characterised. This increases the risk of losing the ability to select for breed-specific traits from native breeds that might be important to the UK sheep sector in the future. Here, we use 50K genotyping to perform preliminary analysis of breed relationships and genetic diversity within native UK sheep breeds, as a first step towards a comprehensive characterisation. This study generates novel data for thirteen native UK breeds, including 6 on the UK Breeds at Risk (BAR) list, and utilises existing data from the publicly available Sheep HapMap dataset to investigate population structure, heterozygosity and admixture.</p> <p><strong>Results: </strong>In this study the commercial breeds exhibited high levels of admixture, weaker population structure and had higher heterozygosity compared to the other native breeds, which generally tend to be more distinct, less admixed, and have lower genetic diversity and higher kinship coefficients. Some breeds including the Wiltshire Horn, Lincoln Longwool and Ryeland showed very little admixture at all, indicating a high level of breed integrity but potentially low genetic diversity. Population structure and admixture were strongly influenced by sample size and sample provenance – highlighting the need for equal sample sizes, sufficient numbers of individuals per breed, and sampling across multiple flocks. The genetic profiles both within and between breeds were highly complex for UK sheep, reflecting the complexity in the demographic history of these breeds.</p> <p><strong>Conclusion: </strong>Our results highlight the utility of genotyping data for investigating breed diversity and genetic structure. They also suggest that routine generation of genotyping data would be very useful in informing conservation strategies for rare and declining breeds with small populations sizes. We conclude that generating genetic resources for the sheep breeds that are native to the UK will help preserve the considerable genetic diversity represented by these breeds, and safe guard this diversity as a valuable resource for the UK sheep sector to utilise in the face of future challenges.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.