Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Figure 5 in Palaeoclimatic distribution models predict Pleistocene refuges for the Neotropical harvestman Geraeocormobius sylvarum (Arachnida: Opiliones: Gonyleptidae)
Figure 5. Limiting factor analysis for Geraeocormobius sylvarum, applied to current climate (above left), −6k (above right), −21k CCSM (below left) and −21k MIROC (below right). Most relevant variables recognized as limiting: bc18, precipitation of warmest quarter; bc7, temperature annual range; bc9, mean temperature of driest quarter; bc8, mean temperature of wettest quarter; bc14, precipitation of driest month and bc3, isothermality. Small points: presence records.
Figure 4 in Palaeoclimatic distribution models predict Pleistocene refuges for the Neotropical harvestman Geraeocormobius sylvarum (Arachnida: Opiliones: Gonyleptidae)
Figure 4. Marginal response curves for the most relevant variables used to model the distribution of Geraeocormobius sylvarum. Curves depict average values (black curve) and ± SD (grey curves) of the 50-replicates run; vertical dotted lines represent the lowest and highest values in the record set, i.e. they delimit the empiric range of the variable. Above: bc18, precipitation of warmest quarter (critical values: 390–560 mm) and bc7, temperature annual range (critical values: 21°C–24.3°C); bottom: bc14, precipitation of driest month (critical value:>42 mm) and bc9, mean temperature of driest quarter (critical values: 14.3°C–18.9°C, peak at 17.3°C). The y-axis displays the logistic probability of presence; horizontal line indicates the 0.50 value.
Figure 3 in Palaeoclimatic distribution models predict Pleistocene refuges for the Neotropical harvestman Geraeocormobius sylvarum (Arachnida: Opiliones: Gonyleptidae)
Figure 3. Distribution models of Geraeocormobius sylvarum for current climate (above left), −6k (above right), −21k CCSM (below left) and −21k MIROC (below right). Maps display the binary prediction (grey), with the stable area highlighted in colour. Small points: presence records. AUC for the 50-replicates run: 0.9839–0.9865; average 0.9854.
Figure 1 in Palaeoclimatic distribution models predict Pleistocene refuges for the Neotropical harvestman Geraeocormobius sylvarum (Arachnida: Opiliones: Gonyleptidae)
Figure 1. Adult male of Geraeocormobius sylvarum, from Tucumán Province, Argentina. Photo: Gonzalo D. Rubio.
Figure 2 in Palaeoclimatic distribution models predict Pleistocene refuges for the Neotropical harvestman Geraeocormobius sylvarum (Arachnida: Opiliones: Gonyleptidae)
Figure 2. Records of Geraeocormobius sylvarum, displayed along with the relevant ecoregions in subtropical South America. APAF, Alto Paraná Atlantic forests; AMF, Araucaria moist forests; HCh, Humid Chaco; DCh, Dry Chaco; Y, Southern Andean Yungas; MS, Southern Cone Mesopotamian savanna; FS, Paraná flooded savanna; SM, Serra do Mar coastal forests (nomenclature after Olson et al. 2001).
Predicting readers' prototypical eye-movement behavior using MASC, a model of Attention in the Superior Colliculus: Stimulus materials, model code, data, and statistical analyses.
<p>The goal of the present research was to determine the role of rudimentary visuo-motor pathways, from the retina and the primary visual cortex to the superior colliculus (SC), in the guidance of human eye movement during reading. To this end, we used MASC, our model of Attention in the Superior Colliculus (Adeli et al., Journal of Neuroscience 2017), a model that relies on well-established saccade-programming principles in the SC. MASC predicts sequences of fixations over an input image by spatially integrating incoming signals in the space of the SC.</p> <p>Here, MASC computed the distribution of luminance contrast over sentences' images (visual-saliency map), after blurring it proportional to retinal eccentricity (retina transformation). It then projected the visual-saliency map into SC space, using a logarithmic afferent-mapping function (magnification factor). Input signals were averaged over retinotopically organized populations of neurons (point images) of constant size, first in the visual map and then in a spatially-registered motor map. The most active population was identified through a winner-take-all process. After jitter applied to the winning population, the next fixation location was determined using inverse efferent mapping. This sequence of events was then repeated to predict following fixation locations, but inserting after each saccade an inhibitory spatial tag (Inhibition of Saccade Return; ISR -referred to as IOR in the uploaded files). All MASC's parameters, but one, were biologically determined, using electrophysiological data in macaque; the ISR window was the one fit parameter.</p> <p>MASC was tested by comparing its predicted sequences of fixations over sentences from the French-Sentence Corpus (FSC) to the eye-movement behavior of 40 French-native speakers reading the same sentences for comprehension (Albrengues et al., Plos One 2019). Then, MASC was dissected to determine the crucial processing steps enabling prediction of human behavior (10 comparison models -see the general README file). Finally, to address crucial issues in the reading literature, i.e., the role of inter-word spacing and character print size in eye-movement guidance, MASC was additionally tested in four additional display conditions: the same sentences from the FSC, but with blank spaces between words being either filled or removed, or with the screen width angle being multiplied by 2 or 4, such that characters were larger in angular size (0.5° and 1°) than in the original experiment (0.25°). MASC's predicted effects of inter-word spacing and print size were compared to previously published data.</p> <p>All material relevant to the project is reported here, including the FSC materials (bitmap and information text files), the Matlab code for our MASC model, raw simulation data for MASC and all our comparison models, as well as MASC's simulations in the different display conditions, the scripts we developed in R to transform raw simulation data into data matrices for statistical analyses of (word-based) eye-movement behavior, the resulting data matrices for all models as well as the data matrix for FSC readers, the R-scripts for statistical comparison of oculomotor behavior between data sets and conditions, literature-review tables of previously published data (for comparison with MASC's predictions), and the R-scripts generating the figures summarizing our results.</p> <p>Further information can be found in the general README file as well as in the README files attached to each folder. The authors' respective contributions to the project, the licence attached to the included materials and their condition of use are listed in the general README file.</p> <p>A manuscript reporting and discussing these modeling data is in preparation (Vitu, F., Adeli, H. & Zelinsky, G. J.); A reference will be provided here when the manuscript appears in a journal.</p> <p>Other references to be cited:</p> <p>- For the model code: Adeli, H., Vitu, F., & Zelinsky, G. J. (2017). A model of the superior colliculus predicts fixation locations during scene viewing and visual search. Journal of Neuroscience, 37(6), 1453-1467. http://www.jneurosci.org/content/37/6/1453</p> <p>- For FSC materials and data: Albrengues, C., Lavigne, F., Aguilar, C., Castet, E., & Vitu, F. (2019). Linguistic processes do not beat visuo-motor constraints, but they modulate where the eyes move regardless of word boundaries: Evidence against top-down word-based eye-movement control during reading. PLoS ONE 14(7): e0219666. https://doi.org/10.1371/journal.pone.0219666<br> </p>
Lineage-level distribution models lead to more realistic climate change predictions for a threatened crayfish
<p><b>Aim: </b>As<b> </b>climate change presents a major threat to biodiversity in the next decades, it is critical to assess its impact on species habitat suitability to inform biodiversity conservation. Species distribution models (SDMs) are a widely used tool to assess climate change impacts on species' geographical distributions. As the term suggests, the species-level is the most commonly used taxonomic unit in SDMs. However, recently it has been demonstrated that SDMs considering taxonomic resolution below (or above) the species-level can make more reliable predictions of biodiversity change when different populations exhibit local adaptation. Here, we tested this idea using the Japanese crayfish (<i>Cambaroides japonicus</i>), a threatened species encompassing two geographically structured and phylogenetically distinct genetic lineages.</p> <p><span><b>Location: </b>Northern Japan.</span></p> <p><b>Methods: </b>We first estimated niche differentiation between the two lineages of <i>C. japonicus</i> using <i>n</i>-dimensional hypervolumes, then made climate change predictions of habitat suitability using SDMs constructed at two phylogenetic levels: species and intraspecific lineage.</p> <p><b>Results: </b>Our results showed only intermediate niche overlap, demonstrating measurable niche differences between the two lineages. The species-level SDM made future predictions that predicted much broader and severe impacts of climate change. However, the lineage-level SDMs led to reduced climate change impacts overall, and also suggested that the eastern lineage may be more resilient to climate change than the western one.</p> <p><strong>Main conclusions</strong>: The two lineages of <em>C. japonicus</em> occupy different niche spaces. Compared with lineage-level models, species-level models can overestimate climate change impacts. These results not only have important implications for designing future conservation strategies for this threatened species, but also highlight the need for incorporating genetic information into SDMs to obtain realistic predictions of biodiversity change.</p>
Modeling pulsed evolution and time-independent variation improves the confidence level of ancestral and hidden state predictions
<p><span><span><span><span><span><span><span><span><span><span>Ancestral state reconstruction is not only a fundamental tool for studying trait evolution, but also very useful for predicting the unknown trait values (hidden states) of extant species. A well-known problem in ancestral and hidden state predictions is that the uncertainty associated with predictions can be so large that predictions themselves are of little use. Therefore, for meaningful interpretation of predicted traits and hypothesis testing, it is prudent to accurately assess the uncertainty of the predictions. Commonly used constant-rate Brownian motion (BM) model fails to capture the complexity of tempo and mode of trait evolution in nature, making predictions under the BM model vulnerable to lack-of-fit errors from model misspecification. Using empirical data (mammalian body size and bacterial genome size), we show that the distribution of residual Z-scores under the BM model is neither homoscedastic nor normal as expected. Consequently, the 95% confidence intervals (CIs) of predicted traits are so unreliable that the actual coverage probability ranges from 33% (strongly permissive) to 100% (strongly conservative). Alternative methods such as BayesTraits and StableTraits that allow variable rates in evolution improve the predictions but are computationally expensive. Here we develop RasperGade, a method of ancestral and hidden state prediction that uses the Levy process to explicitly model gradual evolution, pulsed evolution and time-independent variation. Using the same empirical data, we show that RasperGade outperforms both BayesTraits and StableTraits and is orders-of-magnitude faster. Our results suggest that, when predicting the ancestral and hidden states of continuous traits, the tempo and mode of evolution should always be assessed and the quality of confidence estimates should always be examined.</span></span></span></span></span></span></span></span></span></span></p>
Results on the Pre-Pazy Wing Model for the 3rd Aeroelastic Prediction Workshop
<p>Add SHARPy deformed flutter data for angles of attack 6 and 7 degrees</p>
Predictive model of transcriptional elongation control identifies trans regulatory factors from chromatin signatures
<p>Supplementary data for "Predictive model of transcriptional elongation control identifies trans regulatory factors from chromatin signatures" by Toray S. Akcan, Matthias Heinig.</p>
Dataset: A model of tension-induced fiber growth predicts white matter organization during brain folding
<p>This dataset contains models, MRI data, and code associated with "A model of tension-induced fiber growth predicts white matter organization during brain folding" in <em>Nature Communications </em>(accepted 27 Oct 2021).</p> <p>Included files:</p> <ol> <li>NatCommMRIdata.zip - MRI data and associated MATLAB code for analysis of each subject</li> <li>folding_with_fibers.mph - model file used to generate results in Main Text. Modifiable parameters include stress-dependent fiber elongation rate, initial fiber volume fractions, initial geometry, and cortical growth rate. Compatible with COMSOL Multiphysics (version 5.3).</li> <li>folding_with_fibersECM.mph - model file corresponding to Supplementary Fig. 2. Compatible with COMSOL Multiphysics (version 5.3).</li> <li>folding_with_fibers_linear.mph - model file corresponding to Supplementary Fig. 3. Compatible with COMSOL Multiphysics (version 5.3).</li> </ol> <ol> </ol>
The dataset used in "A max-margin model for predicting residue-base contacts in protein-RNA interactions"
<p>The dataset used in "A max-margin model for predicting residue-base contacts in protein-RNA interactions"</p>
Predictive modeling reveals that higher-order cooperativity drives transcriptional repression in a synthetic developmental enhancer
<p>A challenge in quantitative biology is to predict output patterns of gene expression from knowledge of input transcription factor patterns and from the arrangement of binding sites for these transcription factors on regulatory DNA. We tested whether widespread thermodynamic models could be used to infer parameters describing simple regulatory architectures that inform parameter-free predictions of more complex enhancers in the context of transcriptional repression by Runt in the early fruit fly embryo. By modulating the number and placement of Runt binding sites within an enhancer, and quantifying the resulting transcriptional activity using live imaging, we discovered that thermodynamic models call for higher-order cooperativity between multiple molecular players. This higher-order cooperativity capture the combinatorial complexity underlying eukaryotic transcriptional regulation and cannot be determined from simpler regulatory architectures, highlighting the challenges in reaching a predictive understanding of transcriptional regulation in eukaryotes and calling for approaches that quantitatively dissect their molecular nature.</p>
A Predictive Coding Approach to Modelling Perceived Drum Pattern Complexity
<p>Commented <em>R</em> script and matrices to carry out drum pattern complexity prediction.</p>
Key sulphur species predicted by photochemical models and their employed chemical networks
<p>The volume mixing ratio output of the key sulphur species computed by photochemical models for producing Fig. 1 in Tsai et al. 2023 and the chemical networks used in each model.</p> <p> </p>
Improved National-Scale Flood Prediction for Gauged and Ungauged Basins using a Spatio-temporal Hierarchical Model
<p>Composite data with NWM 2.0 streamflow, basin PET, drainage area and stoage.</p> <p>SAR data used in this study are downloaded from</p> <p><a href="https://scihub.copernicus.eu/">https://scihub.copernicus.eu/</a></p> <p> </p>
Predicting GPR40 Agonists with A Deep Learning-Based Ensemble Model
<p>This dataset includes the calculation steps and optimization process of an ensemble model, along with various results and related intermediate files</p>
Model weights and predictions for reproducible benchmarking experiments in MedMNIST v2
<p>This data repository is associated with our <a href="https://github.com/MedMNIST/experiments">GitHub code</a></p> <ol> <li><code>weights_*.zip</code>: <ul> <li>PyTorch, AutoKeras and Google AutoML Vision are provided for MedMNIST2D.</li> <li>PyTorch and AutoKeras are provided for MedMNIST3D.</li> <li>If you are using PyTorch model weights, please note that the ResNet18_224 / ResNet50_224 models are trained with images resized to 224 x 224 by <code>PIL.Image.NEAREST</code>.</li> <li>Snapshots for <code>auto-sklearn</code> are not uploaded due to the embarrassingly large model sizes (lots of model ensemble).</li> </ul> </li> <li><code>predictions.zip</code>: We also provide all standard prediction files by PyTorch, auto-sklearn, AutoKeras and Google AutoML Vision, which works with <code>medmnist.Evaluator</code>. Each file is named as <code>{flag}_{split}_[AUC]{auc:.3f}_[ACC]{acc:.3f}@{run}.csv</code>, e.g., <code>bloodmnist_test_[AUC]0.997_[ACC]0.957@autokeras_3.csv</code>.</li> </ol>
Supplementary dataset for paper: "Approximate non-linear model predictive control with safety-augmented neural networks"
<p>Supplementary dataset for paper Henrik Hose and Johannes Koehler and Melanie N. Zeilinger and Sebastian Trimpe "Approximate non-linear model predictive control with safety-augmented neural networks".</p> <p>The code to use this dataset is publicly available at <a href="https://github.com/hshose/soeampc">https://github.com/hshose/soeampc</a></p> <p>The dataset contains training and testing data to train an NN controller for three standard benchmark systems, a stir tank reactor, a quadcopter, and a chain mass system.</p> <p>For each system, there are initial conditions as comma separated value in the `x0.txt` file, the MPC input trajectory in the `U.txt` file and the corresponding predicted state sequence in the `X.txt` file. MPC parameters are provided for each system. The dataset was computed using acados for SQP solving.</p> <p>The dataset also contains pretrained neural network approximations of the dataset.These are provided in the `pretrained_models.zip` file. The neural networks were trained with tensorflow.</p>
Datasets associated with the publication: "The three-dimensional structure of fronts in mid-latitude weather systems in numerical weather prediction models".
<p>ECMWF HRES and ERA5 datasets for the paper "The three-dimensional structure of fronts in mid-latitude weather systems in numerical weather prediction models".</p> <p>Content: </p> <p>========================================================================================</p> <p>Storm Friederike:</p> <p>ECMWF HRES forecast, 18.01.2018 00:00 UTC - 23:00 UTC, hourly, GRIB-format.</p> <p>ECMWF ERA5 reanalysis, 16.01.2018 12:00 UTC - 19.01.2018 00:00 UTC, twelve-hourly data, GRIB-format.</p> <p>========================================================================================</p> <p>Storm Vladiana:</p> <p>ECMWF HRES analysis, 23.09.2016 00:00 UTC - 23.09.2016 00:00 UTC, six-hourly data, rotated North Pole (latitude: 51˚, longitude: 160˚), NetCDF-format.</p> <p>========================================================================================</p> <p>Storm Egon:</p> <p>ECMWF ERA5 reanalysis, 12.01.2017 00:00 UTC - 13.01.2017 06:00 UTC, six-hourly data, GRIB-format.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.