Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Ranger models for predicting isoform abundance from UTR sequence features
<p>Each RDS file contains a ranger object, trained on transcripts after removing those associated with the held-out genes in one of the five cross-validation folds. The day and replicate number in the file name corresponds to the neuronal differentiation sample on which the model was trained. The file gene_folds.txt indicates the fold from which each gene was excluded during model training. The file transcript_gene_associations.txt contains transcript-gene associations. The file predictors.RDS contains the matrix of predictor variables.</p>
FIGURE 14 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 14. Detailed comparison of model versions.
FIGURE 13 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 13. Refined model results.
FIGURE 9 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 9. Revised weighted suitability analysis results.
FIGURE 10 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 10. Number of cells assigned to each fossil potential value for the revised model.
FIGURE 7 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 7. Photos of the 10 test sites. Numbers correspond to those in Table 5.
FIGURE 6 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 6. Weighted suitability analysis results.
FIGURE 4 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 4. Simplified flowchart showing methodology.
FIGURE 3 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 3. Slope and Aspect data for the Cedar Mountain Formation.
FIGURE 2. Landsat 8 in A fossil locality predictive model using weighted suitability analysis for the Early Cretaceous Cedar Mountain Formation, Utah, USA
FIGURE 2. Landsat 8 natural color composite image centered on the Cedar Mountain Formation.
Dataset of bike-sharing Demand Prediction model based on Spatio-Temporal Graph Convolutional Networks
<p>Dataset of bike-sharing Demand Prediction model based on Spatio-Temporal Graph Convolutional Networks</p>
Lithosphere removal (delamination) following continental collision: shape and complexity of observables predicted from 3D numerical models
<p>Upload contains most essential data files from the manuscript. The raw data amount to several TB and cannot be uploaded.</p> <p><strong>1 Model data</strong></p> <p><em>1.1 Models covered</em></p> <p>Relatively frequent output is available for the reference model (<em>R</em> in paper, see Table S2), prefixed <em>dp10</em>. Limited output is uploaded for models with geometric variations (<em>OSM -> dp19, TSM->dp26</em>).</p> <p><em>1.2 Types of data</em></p> <p>Most files are visualisation files (ending *.<em>vtr</em>). These are intended for loading into the visualisation software Paraview. Files starting with <em>comp_</em> contain the lithological composition field; other visualisation files hold the fields of selected physical parameters. Note that as of the design of the study, the maximum file size allowed for loading into Paraview is ca. 2 GB. This limits the number of fields available according to model size - larger models (geometrical variations) hold therefore less fields.</p> <p>Binary files with ending *.<em>prn</em> are saved model states, from which programs can be rerun.</p> <p>Time information mapping step numbers to years is provided in <em>times...txt</em>.</p> <p> </p> <p><strong>2 Software for reproduction</strong></p> <p>The numerical code is not in the public domain, and its use is restricted. A pre-compiled binary is provided and allows reproduction of the main/reference model. We present the input files; however, these cannot be changed freely (equivalent to software distribution), and changes will lead to termination of the program.</p> <p><em>2.1 Requirements, preparation and use</em></p> <p>The code is compiled on CentOS with gcc 4.8.2. There is no library dependency, apart from the intrinsic OpenMP and glibc (>= 2.17). The following requirements <em>must</em> be satisfied in order to run it:</p> <ul> <li>Linux OS (tested on CentOS and Fedora)</li> <li>> 160 GB shared memory</li> </ul> <p>The following provisions are recommended:</p> <ul> <li>16 cores</li> <li>high bandwidth storage</li> <li>storage space <em>O</em>(TB)</li> </ul> <p>To prepare, simply unzip the provided snapshot <em>clsd_reproduce_gcc_CentOS.zip</em> in an appropriate directory (cluster). The input file <em>init.t3c</em> sets the initial conditions; the input file <em>mode.t3c </em>controls solver and file output. <em>file.t3c</em> is the pointer to the last output step; if a model snapshot is available as binary dump (*.prn file), setting the pointer to its number allows restarting.</p> <p>For normal usage, create initial condition with executable <em>in3mg</em>, and subsequently run <em>i3mg</em>. This will create snapshots and *.vtr visualisation files.</p> <p> </p> <p>High-performance computing facilities, runtime (months), proper software environment, and training in use of the code, or in the use of the visualisation software, are not provided.</p> <p> </p>
Evaluation of Predictive Capabilities of Regression Models and Artificial Neural Networks for Density and Viscosity Measurements of Different Biodiesel-Diesel-Vegetable Oil Ternary Blends
<p>In this section, it was given that Annex Figures and Annex Tables related to the article "Evaluation of Predictive Capabilities of Regression Models and Artificial Neural Networks for Density and Viscosity Measurements of Different Biodiesel-Diesel-Vegetable Oil Ternary Blends" published in "Environmental and Climate Technologies" journal. </p>
Dataset - Enhanced flux prediction by integrating relative expression and relative metabolite abundance into thermodynamically consistent metabolic models
<p><strong>Simulation data needed to reproduce the results from the manuscript “Enhanced flux prediction by integrating relative expression and relative metabolite abundance into thermodynamically consistent metabolic models”</strong><br> by V. Pandey, N. Hadadi and V. Hatzimanikatis</p> <p>"REMI manuscript - simData" folder contains all simulation data which can be used to generate results of the paper: <br> • Expression_data: This folder contains Transcriptomics data from both studies: Ishii et al (see test_expr.mat) and Holm et al.<br> • Fluxdata: Fluxomics data can be found form the studies Ishii et al and Holm et al.<br> • Metabolomics: This contains metabolomics data of aforementioned both studies.<br> • ModelsSolutions: We generated different models using with thermodynamics (TGex, TGexM, TM) and without thermodynamics models (Gex, GexM, M). Gex indicates integration with only gene expression, GexM indicates gene expression and metabolite, and M indicates only metabolites. ‘T’ is used for thermodynamic models. Models for different mutants and conditions (e.g. pgm, pgi) can be found in the corresponding folders (TGex, TGexM, TM, Gex, GexM, and M). Variables with the ‘store’ tag comprises flux solutions, correlation values and percentage error between simulation and experiment fluxes.<br> • AlternativeMCS: We generated alternative states for MCS and saved results.<br> • FVAMM: This is the result flux variability analysis can be found in this folder.<br> • Scatter_plot: Scatter plots indicates correlation between measured and model predicted fluxes.</p> <p> </p>
T2* and quantitative susceptibility mapping in an equine model of post-traumatic osteoarthritis: prediction of mechanical and structural properties
<p>Dataset for the manuscript titled "T2* and quantitative susceptibility mapping in an equine model of post-traumatic osteoarthritis: assessment of mechanical and structural properties"</p>
Data set related to the manuscript "On the development of an original mesoscopic model to predict the capacitive properties of carbon-carbon supercapacitors"
<p>Graphical files in the agr format for all the figures in the main text of the manuscript entitled "On the development of an original mesoscopic model to predict the capacitive properties of carbon-carbon supercapacitors" (<a href="https://doi.org/10.1016/j.electacta.2019.135022">10.1016/j.electacta.2019.135022</a>).</p>
The prediction data analyzed in "Seasonal Arctic sea ice prediction using a newly developed fully coupled regional model with the assimilation of satellite sea ice observations"
<p>The outputs of seasonal predictions with the new modeling system analyzed in the article including:</p> <p>Sea ice concentration (SIC)</p> <p>Sea ice thickness (SIT)</p> <p>Sea surface temperature (SST)</p> <p>Near surface air temperature (T2) </p>
Machine learning pipeline to train toxicity prediction model of FunTox-Networks
<p>Machine Learning pipeline used to provide toxicity prediction in FunTox-Networks</p> <p>01_DATA # preprocessing and filtering of raw activity data from ChEMBL<br> - Chembl_v25 # latest activity assay data set from ChEMBL (retrieved Nov 2019)<br> - filt_stats.R # Filtering and preparation of raw data<br> - Filtered # output data sets from filt_stats.R<br> - toxicity_direction.csv # table of toxicity measurements and their proportionality to toxicity</p> <p>02_MolDesc # Calculation of molecular descriptors for all compounds within the filtered ChEMBL data set<br> - datastore # files with all compounds and their calculated molecular descriptors based on SMILES<br> - scripts<br> - calc_molDesc.py # calculates for all compounds based on their smiles the molecular descriptors<br> - chemopy-1.1 # used python package for descriptor calculation as decsribed in: https://doi.org/10.1093/bioinformatics/btt105</p> <p>03_Averages # Calculation of moving averages for levels and organisms as required for calculation of Z-scores<br> - datastore # output files with statistics calculated by make_Z.R<br> - scripts<br> -make_Z.R # script to calculate statistics to calculate Z-scores as used by the regression models<br> <br> 04_ZScores # Calculation of Z-scores and preparation of table to fit regression models<br> - datastore # Z-normalized activity data and molecular descriptors in the form as used for fitting regression models<br> - scripts<br> -calc_Ztable.py # based on activity data, molecular descriptors and Z-statistics, the learning data is calculated</p> <p>05_Regression # Performing regression. Preparation of data by removing of outliers based on a linear regression model. Learning of random forest regression models. Validation of learning process by cross validation and tuning of hyperparameters.</p> <p>- datastore # storage of all random forest regression models and average level of Z output value per level and organism (zexp_*.tsv)<br> - scripts<br> - data_preperation.R # set up of regression data set, removal of outliers and optional removal of fields and descriptors<br> - Rforest_CV.R # analysis of machine learning by cross validation, importance of regression variables and tuning of hyperparameters (number of trees, split of variables)<br> - Rforest.R # based on analysis of Rforest_CV.R learning of final models</p> <p>rregrs_output<br> # early analysis of regression model performance with the package RRegrs as described in: https://doi.org/10.1186/s13321-015-0094-2</p>
Computational model results for "Uncertainties of Glacial Isostatic Adjustment model predictions in North America associated with 3D structure"
<p>The mean GIA signals of RSL, u-dot and g-dot with 1σ, 2σ and 3σ uncertainties in North America. </p>
The effect of uncertainty in humidity and model parameters on the prediction of contrail energy forcing
<p>Previous work has shown that while the net effect of aircraft condensation trails (contrails) on the<br>climate is warming, the exact magnitude of the energy forcing per meter of contrail remains uncertain.<br>In this paper, we explore the skill of a Lagrangian contrail model (CoCiP) in identifying flight<br>segments with high contrail energy forcing. We find that skill is greater than climatological<br>predictions alone, even accounting for uncertainty in weather fields and model parameters.</p> <p>We estimate the uncertainty in weather by using the ensemble ERA5 weather reanalysis from the European<br>Centre for Medium-Range Weather Forecasts (ECMWF) as Monte Carlo inputs to CoCiP. We unbias and correct<br>under-dispersion on the ERA5 humidity data by forcing a match to the distribution of in situ humidity<br>measurements taken at cruising altitude. We set aside CoCiP energy forcing estimates calculated using<br>one of the ensemble members as a proxy for ground truth, and report the skill of CoCiP in identifying<br>segments with large positive proxy energy forcing. We further estimate the uncertainty in the model<br>parameters in CoCiP by performing Monte Carlo simulations with CoCiP model parameters drawn from<br>uncertainty distributions consistent with the literature.</p> <p>When CoCiP outputs are averaged over seasons to form climatological predictions, the skill in<br>predicting the proxy is 44%, while the skill of per-flight CoCiP outputs is 84%. If these results carry<br>over to the true (unknown) contrail EF, they indicate that per-flight energy forcing predictions can<br>reduce the number of potential contrail avoidance route adjustments by 2x, hence reducing both the cost<br>and fuel impact of contrail avoidance.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.