Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
82
datasets available to search
ShareScore release 0.9.0
Dataset results
82 results for “Predictive Simulations”
Assessing the adequacy of morphological models using Posterior Predictive Simulations
Open the record for dataset details and reuse information.
High-resolution climate simulations using the Model for Prediction Across Scales - Atmosphere (MPAS-A; version 5.1)
Open the record for dataset details and reuse information.
Data from: Use of simulation-based statistical models to complement bioclimatic models in predicting continental scale invasion risks
Invasive species represent one of the greatest risks to global biodiversity and economic productivity of agroecosystems. The development of certain novel crops—e.g., herbaceous perennial biomass crops—may create a risk of novel invasions by these crops. Therefore, potential benefits and risks need to be weighed in making decisions about their introduction and subsequent management. Ideally, such a weighing will be based on good estimates of invasion risks in realistic scenarios pertaining to actual landscapes of concern regarding invasion. Most previous large-scale analyses of invasion risk have used species distribution models and their established methods. Unfortunately, these approaches are unable to incorporate local scale biotic and spatial factors that influence invasion risk. Here we present a case study for how such factors can be efficiently incorporated in large-scale analyses of invasion risk, by extending simulation models with statistical modeling tools. By these means, we predict invasion risk at the scale of the entire United States for a major biomass crop, Miscanthus × giganteus. We then combine invasion risk predictions for this method with those from bioclimatic methods, producing a map of aggregated invasion risk that can offer more nuanced predictions of invasion risk than either approach alone. Lastly, we evaluate potential risks for invasive crops that differ in invasiveness traits, to examine how geographic patterns of invasion risk vary among invaders as a result of their particular constellation of traits.
About prediction of vehicle energy consumption for eco-routing: simulation results
<p>Supplementary materials, experiment results, processing scripts</p>
Molecular simulations to investigate the impact of N6-methylation in RNA recognition: Improving accuracy and precision of binding free energy prediction
<p>Dataset relative to Molecular dynamics simulation performed for the work "Molecular simulations to investigate the impact of N6-methylation in RNA recognition: Improving accuracy and precision of binding free energy prediction".<br><br>The dataset contains data of 42 alchemical simulations and is subdivided in 4 zip files.<br><br>Folders are named following the scheme: system_configuration_forcefield.<br>Zip file C1 contains .mdp files used for all the simulations.<br><br>Folders corresponding to simulations performed with the fit5_AC ff contains:<br>- topology files (topol.top, topol_RNA_chain_A.itp, topol_RNA_chain_B.itp)<br>- index files needed to reconstruct the demuxed trajectories (replica_index.xvg , replica_index.xvg)<br>- 16 folders, one for each replica (lam0 ... lam15), containing:<br> - final configuration (confout.gro)<br> - log file (md.log)</p> <p> - input file for md run (md.tpr)<br> - energies for the concatenated trajectories recomputed for the realtive replica hamiltonian (ener_trj_conc.edr)<br><br>Folders corresponding to simlations performed with fit_A parametrization only contains .edr files corresponding to energies for the concatenated trajectory computed for 14 set of DeQs drawn from a gaussian distribution, with the relative topologies.<br><br>Supplementary materials relative to simlations performed with fit_A parametrizationcan be found in: https://zenodo.org/records/6498021</p>
Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction
<p>This Zenodo repository provides comprehensive resources for the paper titled "Spatio-temporal learning from molecular dynamics simulations for protein-ligand binding affinity prediction" published on <a href="https://academic.oup.com/bioinformatics/article/41/8/btaf429/8238154">Bioinformatics</a>. We created a dataset of 63,000 molecular dynamics simulations by performing 10 simulations of 10 ns on 6,300 complexes. Neural networks were developed to learn from this data in order to predict the binding affinities of protein-ligand complexes. The implementation of these neural networks are available on <a href="https://github.com/ICOA-SBC/MD_DL_BA" target="_blank" rel="noopener">github</a>. Our collection includes training/benchmark datasets, trained statistical models, and results on test sets (CSV & PDF files).</p> <p> </p> <p><strong>Training/benchmark datasets:</strong></p> <p>Training, validation and test sets are provided to train and evaluate the following neural networks:</p> <ul> <li>Pafnucy, Proli and Densenucy without MD data augmentation (dataset file names contain "initial")</li> <li>Pafnucy, Proli and Densenucy with MD data augmentation (dataset file names contain "MDDA")</li> <li>Pafnucy with/without MD data augmentation and Proli and Densenucy with MD data augmentation were also evaluated on the fep test set (test set file name contain "fep")</li> <li>Timenucy and Videonucy using spatiotemporal learning methods (dataset file names contain "4D")</li> <li>Pafnucy without MD data augmentation and on a reduced training set (dataset file names contain "reduced")</li> </ul> <p>For each training methodology (MD data augmentation and spatiotemporal learning), we provide the data for the whole complex, only the ligand or only the protein. Additionally for spatiotemporal learning, we provide the data with only the ligand using the tracking mode.</p> <p> </p> <p><strong>Statistical models:</strong></p> <p>We provide the models trained with Pafnucy, Proli, Densenucy, Timenucy and Videonucy. Each models were trained in 10 replicates. </p> <p>For Pafnucy, Proli, Densenucy, we provide the models trained with random and systematic rotations, as well as with or without MD data augmentation.</p> <p>For Proli, Densenucy, Timenucy and Videonucy, we provide the models trained on the whole complex, only the ligand or only the protein.</p> <p>For Pafnucy we also provide the models trained on the reduced set (5932 complexes).</p> <p> </p> <p><strong>Results on test sets (CSV & PDF files):</strong></p> <p>We provide the predictions on the PDBbind v.2016 core set.</p> <ul> <li>For spatiotemporal learning methods (Timenucy and Videonucy), there are predictions for only 83 complexes, as we did not perform simulations on the whole test set.</li> <li>For models trained with MD DA, predictions were carried on the crystallographic structures as well as on the frames extracted from the simulations performed on the test set (augmented test).</li> </ul> <p>Results on the FEP dataset are also provided for Pafnucy, Proli and Densenucy.</p> <p> </p> <p>The Raw MD data (~4.5 To) are stored, and can be visualized/downloaded, on the <a href="https://mdposit.mddbr.eu/#/browse?search=MDBind">MDDB</a>.</p> <p>This work was performed using HPC resources from GENCI-IDRIS (Grant 2021-A0100712496 & 2022-AD011013521) and CRIANN (Grant 2021002).</p>
The AF2 predicted structures and MD simulation results for the paper "In-situ structural insights into activity regulation of mammalian pyruvate dehydrogenase complex"
<p>Thank you for your interest in our work. You can discover content that interests you within the respective compressed packages, accompanied by “readme” files.</p>
Hydrograph and Recession Flows Predictions Simulations using Deep Learning: Watershed Uniqueness and Objective Functions
<p>This resource contains the results presented in the paper titled 'Hydrograph and Recession Flows Simulations using Deep Learning: Watershed Uniqueness and Objective Functions' by Abhinav Gupta and Sean McKenna.<span><span><br></span></span></p> <p>'</p>
Evaluation of Spectral Light Simulation Tools For Prediction of ipRGC-influenced Light Responses in Field Conditions
<p><span>Spectral light simulation tools can be used to optimize lighting design in built environments for occupants’ ipRGC-influenced light (IIL) responses. The accuracy of these tools has been previously tested against measurements in controlled, neutral-colored spaces. This study aims to evaluate the accuracy of ALFA and Lark spectral light simulation tools in spectrally and geometrically complex environments. The methods consist of comparison of spectral irradiance, three IIL response metrics, and photopic illuminance in two office rooms in a university building. Measurements and simulations are performed in daylight, electric light, and combination of daylight and electric light. The dataset presented here documents this research.</span></p>
Data set for Predicting hospital occupancy for covid-19 patients: a simulation approach based on archetypes of empirical services' trajectories
<p>Data set for the paper: Predicting hospital occupancy for covid-19 patients: a simulation approach based on archetypes of empirical services’ trajectories</p> <p>Based on: Marin-Garcia, J. A., Ruiz, A., Julien, M., & Garcia-Sabater, J. P. (2021). A data generator for covid-19 patients’ care requirements inside hospitals. WPOM-Working Papers on Operations Management, 12(1), 76-115. https://doi.org/10.4995/wpom.15332</p> <p> </p>
Simulated metagenomic datasets used for determining the threshold score cutoff for taxa prediction using StrainIQ software
<p>StrainIQ (Strain Identification and Quantification) is a novel tool that implements a new <em>n</em>-gram based algorithm for predicting and quantifying strain-level taxa from whole genome metagenomic sequencing data. To avoid the false positive prediction of by StrainIQ algorithm, we determined cutoff score for each site specific n-gram model called DSEM (DNA Signature Element Model) using positive and negative metagenomic datasets. The Figure2-datasets.zip has positive (from GI tract genomes) and negative (from non-GI tract genomes) simulated metagenomic datasets that were used for determining optimal cut-off based on GI tract DSEM.</p> <p>Please find the new link for this data at https://zenodo.org/record/8132164</p>
Artificial Intelligence (AI) Empowered Smile Simulations: Do They Accurately Predict Actual Post Treatment Outcomes
ClinicalTrials.gov study NCT06123585. IPD Sharing: Not stated. Countries: 1. Publications: 1.
Mindfulness and Executive Functions for Prediction of Non-technical Skills of Students in Pediatric Medical Simulations
ClinicalTrials.gov study NCT03761355. IPD Sharing: UNDECIDED. Countries: 1. Publications: 7.
A Research Study to Understand and Predict the Failure of Shoulder Fracture Fixations Using Computer Simulations
ClinicalTrials.gov study NCT04056351. IPD Sharing: NO. Countries: 2. Publications: 2.
Data from: Use of simulation-based statistical models to complement bioclimatic models in predicting continental scale invasion risks
Open the record for dataset details and reuse information.
Predicting Hydrophobicity by Learning Spatiotemporal Features of Interfacial Water Structure: Combining Molecular Dynamics Simulations with Convolutional Neural Networks
<p>Files for reproducing results from Kelkar et al. (JPCB 2020) - Predicting Hydrophobicity by Learning Spatiotemporal Features of Interfacial Water Structure: Combining Molecular Dynamics Simulations with Convolutional Neural Networks</p> <p> </p> <p>This folder contains simulations starter files and also plug-and-play datasets to test ML algorithms on molecular dynamics (MD) simulation data.</p> <p> </p> <p>All analysis scripts can also be found on GitLab on this link: https://gitlab.com/atharva-kelkar/kelkar_et_al_jpcb_2020</p>
Data from: Robust regression and posterior predictive simulation increase power to detect early bursts of trait evolution
A central prediction of much theory on adaptive radiations is that traits should evolve rapidly during the early stages of a clade's history and subsequently slowdown in rate as niches become saturated – a so-called "Early Burst". Although a common pattern in the fossil record, evidence for early bursts of trait evolution in phylogenetic comparative data has been equivocal at best. We show here that this may not necessarily be due to the absence of this pattern in nature. Rather, commonly used methods to infer its presence perform poorly when when the strength of the burst - the rate at which phenotypic evolution declines - is small, and when some morphological convergence is present within the clade. We present two modifications to existing comparative methods that allow greater power to detect early bursts in simulated datasets. First, we develop posterior predictive simulation approaches and show that they outperform maximum likelihood approaches at identifying early bursts at moderate strength. Second, we use a robust regression procedure that allows for the identification and down-weighting of convergent taxa, leading to moderate increases in method performance. We demonstrate the utility and power of these approach by investigating the evolution of body size in cetaceans. Model fitting using maximum likelihood is equivocal with regards the mode of cetacean body size evolution. However, posterior predictive simulation combined with a robust node height test return low support for Brownian motion or rate shift models, but not the early burst model. While the jury is still out on whether early bursts are actually common in nature, our approach will hopefully facilitate more robust testing of this hypothesis. We advocate the adoption of similar posterior predictive approaches to improve the fit and to assess the adequacy of macroevolutionary models in general.
Data from: Simulated hatching failure predicts female plasticity in extra-pair behavior over successive broods
While many studies have investigated the occurrence of extra-pair paternity (EPP) and its adaptive significance in wild population of birds, we still know surprising little about the plasticity in mating behavior of females at the individual level and how it affects the patterns of paternity. To address this question, we focused on the direct fertility benefit hypothesis for the function of EPP and studied if female birds react in extra-pair mating behavior after reproductive failures using a wild population of the Japanese great tit, Parus minor, a socially monogamous passerine with a moderate frequency of EPP and a high multiple brooding rate. We simulated hatching failure by replacing with artificial eggs during the egg laying period to investigate if females subsequently altered their mating behavior and became more promiscuous to improve reproductive success in their following clutches. The proportion of extra-pair offspring per clutches of both experimental and control pairs increased in the second clutches (replacement and repeat), but compared to the control pairs, the increase in the experimental pairs was significantly greater. The present study suggests that individual females appear to be making decisions based on specific cues and flexibly altering mating behavior in adaptive ways. Also, our results are compatible with one of the long-debated hypotheses for the evolutionary maintenance of EPP which predicts females gain direct fitness benefit through increased reproductive success from mating multiply.
Data from: Phylodynamic model adequacy using posterior predictive simulations
Rapidly evolving pathogens, such as viruses and bacteria, accumulate genetic change at a similar timescale over which their epidemiological processes occur, such that it is possible to make inferences about their infectious spread using phylogenetic time-trees. For this purpose it is necessary to choose a phylodynamic model. However, the resulting inferences are contingent on whether the model adequately describes key features of the data. Model adequacy methods allow formal rejection of a model if it cannot generate the main features of the data. We present TreeModelAdequacy (TMA), a package for the popular BEAST2 software, that allows assessing the adequacy of phylodynamic models. We illustrate its utility by analysing phylogenetic trees from two viral outbreaks of Ebola and H1N1 influenza. The main features of the Ebola data were adequately described by the coalescent exponential-growth model, whereas the H1N1 influenza data was best described by the birth-death SIR model.
Dynamic profiling and binding affinity prediction of NBTI antibacte-rials against DNA gyrase enzyme by multidimensional machine learning and molecular dynamics simulations
<p>The chemical libraries used in this study comprised of 199 and 133 structurally diverse novel bacterial topoisomerase inhibitors (<em>alias</em> NBTIs), with experimentally determined <em>in vitro</em> antibacterial potencies against <em>Staphylococcus aureus</em> DNA gyrase (IC<sub>50</sub>=0.007-50 µM) and <em>Escherichia coli</em> DNA gyrase (IC<sub>50</sub>=0.020-100 µM), respectively (named as NBTI<em><sub>SA</sub></em> and NBTI<em><sub>EC</sub></em>), were compiled from the literature as *.sdf file format. The chemical structures comprising both NBTI libraries were initially sketched by using ChemDraw Professional 20.1.1 suite and subsequently energetically minimized utilizing Discovery Studio’s integrated Merck Molecular Force Field (MMFF) module. Moreover 4D ligands ensembles of both libraries ready to be used for multidimensional QSAR modeling are available, as well.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.