Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,079
datasets available to search
ShareScore release 0.9.0
Dataset results
1,079 results for “source data”
PMC_visualisation_and_source_data_SS_v1.0.7
<p>This is the first release of the code and data for:</p> <p>"Public health impact of current and proposed age-expanded perennial malaria chemoprevention: a modelling study"</p> <p>Swapnoleena Sen1,2, Lydia Braunack-Mayer3,4, Sherrie L Kelly1, Thiery Masserey1,2, Josephine Malinga4,5, Joerg J Moehrle2, Melissa A Penny4,5*</p> <p>1 Swiss Tropical and Public Health Institute, Allschwil, Switzerland<br>2 University of Basel, Basel, Switzerland<br>3 <span>Institute of Social and Preventive Medicine, University of Bern, Bern, Switzerland </span><br>4 Telethon Kids Institute, Nedlands, WA, Australia<br>5 Centre for Child Health Research, The University of Western Australia, Crawley, WA, Australia</p> <p>*Correspondence to: Prof Melissa A Penny (melissa.penny@uwa.edu.au)</p> <p>In this study, we integrated an individual-based model of malaria (OpenMalaria) with pharmacological models of drug action to assess the public health impact and cost-effectiveness of perennial malaria chemoprevention (PMC), and the added benefit of further age-expanded dosing schedule (referred as PMC+).</p> <p>The details of running OpenMalaria model, data generation and analysis (including R scripts used for preparing the source data files) for this study can be found in a separate "OpenMalaria_workflow_PMC_modeling" repository (DOI:10.5281/zenodo.12721515). </p> <p>Here the plotting functionalities are described. The repository is strcutured based on figures reported in the manuscript. Each figure has a folder as per its name that includes: 1) R code to plot figure 2) source data files and 3) one PNG and one PDF version of the figure. </p> <p>Please note: i) "dependencies.R" specifies all package information and dependencies in which the simulation, analysis scripts and plotting scripts are tested and stable. <br>ii) The R scripts rely on the folder structure and working directories used by the researchers. To replicate figures, you will need to adjust the file paths.</p>
Data from: Long-term trends in the occupancy of ants revealed through use of multi-sourced datasets
<p><span>We combined participatory science data and museum records to understand long-term changes in occupancy for 29 ant species in Denmark over 119 years. Bayesian occupancy modelling indicated change in occupancy for 15 species: five increased, four declined and six showed fluctuating trends. We consider how trends may have been influenced by life-history and habitat changes. Our results build on an emerging picture that biodiversity change in insects is more complex than implied by the simple insect decline narrative.</span></p>
Data-base for : 'Partitioning carbon sources between wetland and well-drained ecosystems to a tropical first-order stream - Implications to carbon cycling at the watershed scale (Nyong, Cameroon)'
<p>Dataset of carbon (pCO2, TA, DIC, DOC, POC) and ancillary parameters (water temperature, oxygen saturation, pH, specific conducitivity) in ground and surface waters of the Nyong watershed (Cameroon). The dataset covers one entire year (in 2016) and thus allows describing the varability of carbon and ancillary paramaters concentrations induced by seasons.</p>
Data for "Measurement report: Comparison of airborne in-situ measured, lidar-based, and modeled aerosol optical properties in the Central European background – identifying sources of deviations"
<p>A unique set of data is presented, derived from measurements conducted at the rural central European observatory at Melpitz, Germany. Data derived from remote sensing (lidar), airborne platforms (helicopter, balloon), and ground-based in-situ methods is included. Measured and Mie-modeled optical aerosol parameters are presented in the dry- and ambient state. Modeled optical parameters are based on Mie-theory. For ambient state hygroscopic growth simulations are utilized.</p>
Source data for testing results of the ADS software
<p>This is the source data for the testing results of the ADS software at https://zenodo.org/record/5579390</p> <p> </p> <p> </p>
Data for: Sources of CO2 produced in freshly thawed Pleistocene-age Yedoma permafrost
<p>This dataset contains δ13C and F14C compositions of CO2 samples as well as sedimentary parameters from Pleistocene Yedoma located on Kurungnakh Island in the Lena River Delta, collected during an expedition in July/August 2017.</p> <p>Sediment samples were collected from active layer soil pits, using shovels, at three sites on an active retrogressive thaw slump: Pleistocene-aged Yedoma from intact thaw mounds (TM1, TM2), intact Holocene polygonal tundra overlaying the thaw slump (HT1) and sediments from the thaw slump floor (SF3), where Pleistocene and Holocene sediments mix as a result of erosion.</p> <p>CO2 was collected in-situ from the three sites using respiration chambers, set up on vegetation-free spots on the active layer and, in fixed intervals, from a 1.5-year laboratory incubation experiment of sediment samples collected during the expedition. The analyses were performed to compare the C-isotopy of in-situ respired CO2 with that of CO2 produced during the incubation and to determine the sources of the released CO2.</p>
Source Data for the publication: "A different perspective for nonphotochemical quenching in plant antenna complexes"
<p>This dataset contains the raw data for all the results shown in the paper either in Excel or in csv format.</p> <p>The Excel file contains one Excel sheet for each figure panel:</p> <ul> <li>Figure 2a</li> <li>Figure 2b</li> <li>Figure 3a</li> <li>Figure 3d</li> <li>Figure 4a</li> <li>Figure 4b</li> <li>Figure 4c</li> <li>Figure 5b</li> <li>Supplementary Figure 2</li> <li>Supplementary Figure 3</li> <li>Supplementary Figure 4</li> <li>Supplementary Table 1</li> <li>Supplementary Figure 5</li> <li>Supplementary Figure 6</li> <li>Supplementary Figure 7</li> <li>Supplementary Figure 11e</li> </ul> <p>The zipped csv files contain the remaining figures (one per file) and are named accordingly</p>
Vesicles clustering around Wdr35-/- cilia lack electron dense decorations although electron-dense clathrin coated vesicles are still observed budding from the mutant plasma membrane (Figure 7- source data 1)
<p>Intraflagellar transport (IFT) is a highly conserved mechanism for motor-driven transport of cargo within cilia, but how this cargo is selectively transported to cilia is unclear. WDR35/IFT121 is a component of the IFT-A complex best known for its role in ciliary retrograde transport. In the absence of WDR35, small mutant cilia form but fail to enrich in diverse classes of ciliary membrane proteins. In <i>Wdr35 </i>mouse mutants, the non-core IFT-A components are degraded and core components accumulate at the ciliary base. We reveal deep sequence homology of WDR35 and other IFT-A subunits to α and ß' COPI coatomer subunits, and demonstrate an accumulation of 'coat-less' vesicles which fail to fuse with <i>Wdr35 </i>mutant cilia. We determine that recombinant non-core IFT-As can bind directly to<u> </u>lipids and provide the first <i>in-situ</i> evidence of a novel coat function for WDR35, likely with other IFT-A proteins, in delivering ciliary membrane cargo necessary for cilia elongation.</p>
Source data for "Non-Telecentric two-photon microscopy for 3D random access mesoscale 2 imaging"
<p>Source data used in a manuscript "Non-Telecentric two-photon microscopy for 3D random access mesoscale 2 imaging"</p> <p> </p> <p> </p> <p> </p>
Source data for gamete binning for an autotetraploid potato cultivar Otava
<p>Here we provide the source data for analyzing a highly heterozygous autotetraploid potato cultivar 'Otava'.</p>
Machine learning identifies girls with central precocious puberty based on multi-source data
<p><strong>Objective: </strong>The study aimed to develop simplified diagnostic models for identifying girls with central precocious puberty (CPP), without the expensive and cumbersome gonadotropin-releasing hormone (GnRH) stimulation test, which is the gold standard for CPP diagnosis.</p> <p><strong>Materials and Methods:</strong> Female patients who had secondary sexual characteristics before 8 years old and had taken a GnRH analog (GnRHa) stimulation test at a medical center in Guangzhou, China were enrolled. Data from clinical visiting, laboratory tests and medical image examinations were collected. We first extracted features from unstructured data such as clinical reports and medical images. Then, models based on each single-source data or multi-source data were developed with Extreme Gradient Boosting (XGBoost) classifier to classify patients as CPP or non-CPP.</p> <p><strong>Results: </strong>The best performance achieved an AUC of 0.88 and Youden index of 0.64 in the model based on multi-source data. The performance of single-source models based on data from basal laboratory tests and the feature importance of each variable showed that the basal hormone test had the highest diagnostic value for a CPP diagnosis.</p> <p><strong>Conclusion: </strong>We developed three simplified models that use easily accessed clinical data before the GnRH stimulation test to identify girls who are at high risk of CPP. These models are tailored to the needs of patients in different clinical settings. Machine learning technologies and multi-source data fusion can help to make a better diagnosis than traditional methods.</p>
Data and Source codes for: Real-time Radial Tagging for Quantification of Left Ventricular Torsion
<p> </p> <p>Magnetic Resonance Imaging measurement raw data, simulation, and reconstruction codes used in our paper about ‘Real-time Radial Tagging for Quantification of Left Ventricular Torsion' (DOI:10.1002/mrm.29169).</p> <p> </p> <p> </p> <p> </p>
CytofIn_Source_Data_Files
<p>Source Data Files for the article "CytofIn enables integrated analysis of public mass cytometry datasets using generalized anchors" published at Nature Communications.</p>
Source molecular simulation data for calculating energy and friction profiles and permeability coefficients through model lipid membranes
<p>Energy files from GROMACS molecular dynamics simulations with enhanced free energy sampling contain time-dependent evolution of the free energy profiles and friction profiles (and other energies and simulation properties) that were used for calculating permeability coefficients in the publication https://www.biorxiv.org/content/10.1101/2021.07.16.452599v1</p> <p>Simulation system contains a lipid POPC or DPPC bilayer with a varying amount of cholesterol (specified as mol% in the file name). Hydrophobic level of the permeating particle is specified as "level-I", "level-II" etc. When unspecified in the file name, the particle is hydrophobic level "III". Lipids D-C14-PC denote PC lipids with both tails monounsaturated of length 14 carbon atoms. DOPC is equivalent to D-C18-PC. (Detailed description in the publication)</p> <p>Adaptive Weighted Histogram (AWH) method was used to sample the free energy profile of translocating small molecule through the lipid bilayer.</p> <p>GROMACS tool `gmx awh` reads the files and provides the described profiles.</p> <p>Files were generated by GROMACS `mdrun` simulation engine version 2019.3.</p> <p> </p> <p>Coarse-grained MARTINI 3.0 model was used for modeling the biomolecular interactions.</p> <p>Scripts to perform the simulations and the files with initial configurations and simulation settings are stored in a public GitHub repository depozited on Zenodo.org: <a href="https://doi.org/10.5281/zenodo.5082249">https://doi.org/10.5281/zenodo.5082249</a>.</p> <p> </p> <p>Abraham, M. J. et al. GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 1–2, 19–25 (2015).</p> <p>Lindahl, V., Lidmar, J. & Hess, B. Accelerated weight histogram method for exploring free energy landscapes. J. Chem. Phys. 141, 044110 (2014).</p> <p>Souza, P. C. T. et al. Martini 3: a general purpose force field for coarse-grained molecular dynamics. Nat. Methods 18, 382–388 (2021).</p> <p>Melcr, J. Git repository with analysis scripts for MD simulations of permeability through lipid membranes. (2021) doi:<a href="https://doi.org/10.5281/zenodo.5082249">https://doi.org/10.5281/zenodo.5082249</a>.</p>
Data from: Effects of input data sources on species distribution model predictions across species with different distributional ranges
<p>Species distribution models (SDMs) are a popular tool in theoretical and quantitative ecology, and constitute the most widely used modelling framework in global change science and biodiversity conservation. As main data sources, SDMs require georeferenced biodiversity observations as a response or dependent variable (e.g. species occurrence, species richness, etc) and geographic layers of environmental information as predictors or independent variables (e.g. climate, land cover, vegetation indices derived from remote sensing, etc). However, although SDMs have become one of the most important quantitative tools for addressing regular and timely biodiversity assessments worldwide, these techniques are still subject to different sources of uncertainty that have been unequally assessed. Thus, despite uncertainty related to niche-based or distribution-based models has been addressed at different stages in the modelling process, an analysis of the effect of uncertainty coming from alternative data sources on the predictive ability of SDMs is still limited.</p> <p>Citizen-collected species occurrence data (e.g. eBird) are often used for fitting SDMs when data from standardized and expert-supported surveys (e.g. Atlases) are unavailable. On the other hand, macroclimate variables are much more commonly used as predictors in SDMs than other sources of information coming from remote sensing data. We assessed the effects of using different data sources (in both response and predictor variables) on SDM performance across a wide range of bird species with contrasting distributional ranges in the Iberian Peninsula (Portugal and Spain). To do that, a SDM ensemble-forecasting approach was implemented by using bird data from two different data sources: the semi-structured eBird project and standardized Atlases. We fitted SDMs with three predictor types: macroclimate, remotely sensed ecosystem functional attributes (EFAs) from vegetation indices, and their combination. Species were grouped in four range size classes. We also used different evaluation metrics to better assess the uncertainty of model predictions. We then applied generalized linear mixed-effects models to test the effect on model performance of input data source across distributional range sizes while accounting for different accuracy metrics. Pairwise comparisons between range projections were used to assess their spatial similarity.</p> <p>Our models demonstrated the usefulness and complementarity of different input data sources when modelling species distribution across different distributional ranges. Citizen science and remote sensing data contribute to update the knowledge of the distribution of the most threatened bird species by increasing the model accuracy. These findings highlight the need to integrate different data sources to improve the model predictions at regional scale. Our framework also underlines that model uncertainty should be examined more exhaustively at early stages of the modelling process.</p> <p>To perfom and replicate this study, this dataset provides all needed files (as tables) to fit SDMs: i) the Iberian bird species occurrences at 10km UTM square as a response or dependent variable; ii) the geographic layers of environmental information at 10km UTM square for the Iberian Peninsula as predictors or independent variables, such as climate data, ecosystem functioning attributes (EFAs) and the combined climate and EFA data. The dataset is provided by four <em>*.csv</em> files named as:</p> <p><em>1) The_Iberian_bird_species_occurrences_dataset_10km.csv</em></p> <p><em>2) CHELSA_bioclimate_variables_IP10km.csv</em></p> <p><em>3) MODIS_EVI-based_EFAs_IP10km.csv</em></p> <p><em>4) Combined_bioclimate_EFA_dataset_IP10km.csv</em></p> <p>For a more detailed description of the main dataset and each of these subdatasets, please refer to the attached README file.</p> <p><strong>Keywords:</strong> bird atlas, eBird data, ecosystem functional attributes (EFAs), Iberian Peninsula, IUCN categories, Model accuracy, MODIS EVI, narrow-ranged species, remote sensing, species distribution models (SDMs), widespread species</p>
Single cell imaging of ERK and Akt activation dynamics and heterogeneity induced by G protein-coupled receptors - Scripts & Source data
<p>Source data and scripts to reproduce the figures that are part of the publication "Single cell imaging of ERK and Akt activation dynamics and heterogeneity induced by G protein-coupled receptors".</p> <p>Journal of Cell Science (2022) 135, jcs259685, DOI: 10.1242/jcs.259685</p> <p> </p> <p>An earlier version of this work is published as a preprint: "Heterogeneity and dynamics of ERK and Akt activation by G protein-coupled receptors depend on the activated heterotrimeric G proteins", DOI: <a href="https://doi.org/10.1101/2021.07.27.453948">10.1101/2021.07.27.453948</a></p>
Acute lymphoblastic leukemia displays a distinct highly methylated genome (source data)
<p>DNA methylation is tightly regulated during development and is stably maintained in normal cells. In contrast, the methylome of cancer cells is commonly characterized by a global loss of DNA methylation co-occurring with CpG island hypermethylation. In acute lymphoblastic leukemia (ALL), the commonest childhood cancer, perturbations of CpG methylation have been reported to be associated with genetic disease subtype and outcome, but data examining large cohorts at genome-wide scale are lacking. Here, we performed whole genome bisulfite sequencing (WGBS) of leukemic cells of multiple subtypes of ALL, leukemia cell lines and normal hematopoietic cells, and show that in contrast to most cancers, ALL samples exhibit CpG island hypermethylation but minimal global loss of methylation. This was most pronounced in T-ALL and accompanied by an exceptionally broad range of hypermethylation of CpG islands across patients that is influenced by TET2 and DNMT3B. These findings demonstrate a distinct methylome of ALL characterized by an unusually highly methylated genome, and provide important insights into the mechanisms underlying deregulation of methylation in cancer.</p>
EZ publication: source code, profiling, analysis and simulation data
<p>Data of the PIConGPU simulations as used in the publication: EZ: An Efficient, Charge Conserving Current Deposition Algorithm for Electromagnetic Particle-In-Cell Simulations</p> <p>Data overview:</p> <ul> <li>picongpu_source.zip: <ul> <li>source code forked from the PIConGPU mainline version 0.7.0-dev</li> <li>used input set `share/picongpu/examples/PaperThermal`</li> </ul> </li> <li>runs_charge_conservation.zip: <ul> <li>output including hdf5 dumps to validate charge conservation property for the PaperThermal setup (warm plasma)</li> </ul> </li> <li>runs_performance.zip: <ul> <li>simulation timings output for Spock CPU, Spock GPU and Summit GPU runs</li> </ul> </li> <li>runs_profiling.zip: <ul> <li>profile data for Spock GPU and Summit GPU runs</li> </ul> </li> <li>runs_singleParticleTest.zip: <ul> <li>output including hdf5 dumps to validate charge conservation property for the single particle test</li> </ul> </li> <li>analysis_scripts.zip: <ul> <li>jupyter notebooks for setup and analysis of PaperThermal setup</li> <li>python script to plot charge conservation from hdf5 simulation output over time</li> <li>bash script for statistical analysis of performance runs</li> </ul> </li> </ul>
A synaptomic analysis reveals dopamine hub synapses in the mouse striatum : Source Data
<p>Source data for our study on dopamine hub synapses</p> <p>Dopamine transmission is involved in reward processing and motor control, and its impairment plays a central role in numerous neurological disorders. Despite its strong pathophysiological relevance, the molecular and structural organization of the dopaminergic synapse remains to be established. Here, we used targeted labelling and fluorescence activated sorting to purify striatal dopaminergic synaptosomes. We provide the proteome of dopaminergic synapses with 57 proteins specifically enriched. Beyond canonical markers of dopamine neurotransmission such as dopamine biosynthetic enzymes and cognate receptors, we validated 6 proteins not previously described as enriched (Cpne7, Apba1/Mint1, Cadps2, Cadm2/SynCAM 2, Stx4, Mgll). Moreover, our data reveal the adhesion of dopaminergic synapses to glutamatergic, GABAergic or cholinergic synapses in structures we named “dopamine hub synapses”. At glutamatergic synapses, pre- and postsynaptic markers are significantly increased upon association with dopamine synapses. Dopamine hub synapses may thus support local dopaminergic signalling, complementing volume transmission thought to be the major mechanism by which monoamines modulate network activity.</p>
Code and source data for the paper: Global warming leads to larger bats with a faster life history pace in the long-lived Bechstein's bat (Myotis bechsteinii)
<p>Contains two R scripts necessary to perfom the analysis for the paper "Global warming leads to a faster life history pace in the long-lived Bechstein’s bat (Myotis bechsteinii)"</p> <ul> <li>1st Script (" Script_analysis paper_bodysize_AFR_fecundity_LRS_GAMs_revised": Descriptive statistics, calculation of all GAMs and code for figure 1, 2 and 3</li> <li>2nd Script (" Script_size specific generation times"): Calculation of reproductive and mortality rates, calculation of generation time and population growth rates (lambda) as well as code for figure 4 and 5</li> </ul> <p>And also .csv files with the data points of all figures.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.