Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,773

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,773 results for “predictive modeling”

Learn how ShareScore rates datasets ↗
zenodo40/100

Protein language model embeddings and predictions of the human proteome

<p>Residue and sequence embeddings of the human proteome (SwissProt for organism Human, downloaded on&nbsp;2021.06.09)&nbsp;computed using bio_embeddings (bioembeddings.com) using the ProtT5 embedder at full precision (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3).</p> <p>Additionally:</p> <p>- Sequence-level&nbsp;predictions of subcellular localization in 10 classes using LA (https://www.biorxiv.org/content/10.1101/2021.04.25.441334v1)</p> <p>- Residue-level three state secondary structure prediction (alpha, sheet or other) using models reported&nbsp;in the ProtTrans paper (https://www.biorxiv.org/content/10.1101/2020.07.12.199554v3)</p> <p>&nbsp;</p> <p>Files included:</p> <p>- human.fasta --&gt; FASTA-formatted sequences of human from SwissProt</p> <p>-&nbsp;DSSP3_human_ProtT5Sec.fasta --&gt; Secondary structure predictions in three states for each residue of each protein&nbsp;in human.fasta. &quot;H&quot; stands for Helix; &quot;E&quot; stands for Sheet; &quot;C&quot; stands for Other.</p> <p>-&nbsp;subcell_human_LA_ProtT5.csv --&gt; Subcellular location (10 states) and memrane-boundness (2 states)&nbsp;for each protein in human.fasta</p> <p>-&nbsp;embeddings_file.h5 --&gt; per-residue embeddings of sequences in human.fasta. Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a matrix of length Lx1024, with L being the length of the protein sequence. Datasets are indexed using integers. The original sequence identifier (from the FASTA header) can be accessed through the &quot;original_id&quot; attribute. See&nbsp;https://docs.bioembeddings.com/v0.2.0/notebooks/open_embedding_file.html for information on how to open the file</p> <p>-&nbsp;reduced_embeddings_file.h5 --&gt; per-sequence embeddings of sequences in human.fasta (obtained by mean-pooling the residue-embeddings along the length dimension of the protein sequence). Each dataset&nbsp;in the .h5 file represents a protein sequence and contains a vector of size 1024 (meaning, each sequence has the same dimension).</p>

openafl-3.0Jun 2021View details →
zenodo40/100

Regression Model to Predict the Higher Heating Value of Poultry Waste from Proximate Analysis.

<p>The response variable is High Heating Values (HHV), while the independent variables are Fixed Carbon (FC), Volatile Matter (VM), and Ash (A).&nbsp;</p>

opencc-by-4.0Jun 2018View details →
zenodo40/100

Data for publication of "Determining the sensitive parameters of WRF model for the prediction of tropical cyclones in the Bay of Bengal using Global Sensitivity Analysis and Machine Learning"

<p>The data are made available as part of the paper &quot;Determining the sensitive parameters of WRF model for the prediction of tropical cyclones in the Bay of Bengal using Global Sensitivity Analysis and Machine Learning&quot;, submitted to Geoscientific Model Development. This data set incorporates selected post-processed files needed to reproduce the results presented in the paper.</p> <p>The data contains six zip files, that are:</p> <ul> <li>Namelist.input files for the WRF model simulations of ten tropical cyclones</li> <li>WRF model simulation outputs using the default parameter values</li> <li>WRF model simulation outputs using the optimal parameter values (which give minimum RMSE value)</li> <li>IMDAA surface observations and IMERG precipitation data</li> <li>IMD observed tracks of ten tropical cyclones</li> <li>Ipython notebooks of sensitivity analysis and machine learning codes</li> </ul> <p>The remaining files are the ncl scripts that were used to obtain the figures. The ncl scripts used the data in the zip files.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Fig. 2 in Predictive distribution modelling for rufous-necked hornbill Aceros nipalensis (Hodgson, 1829) in the core area of the Western Forest Complex, Thailand

Fig. 2. Diagram showing the conceptual framework used to model the distribution of the rufous-necked hornbill in the Western Forest Complex (WEFCOM), Thailand.

opencc-by-4.0Feb 2014View details →
zenodo40/100

Fig. 1 in Predictive distribution modelling for rufous-necked hornbill Aceros nipalensis (Hodgson, 1829) in the core area of the Western Forest Complex, Thailand

Fig. 1. Map of research locations showing the study areas where the point counts were conducted along14 line-transects, using a hand-held GPS to accurately measure the count locations or extent and position of evergreen forest where the RNH lives, in order to predict RNH distribution.

opencc-by-4.0Feb 2014View details →
zenodo40/100

Fig. 3 in Predictive distribution modelling for rufous-necked hornbill Aceros nipalensis (Hodgson, 1829) in the core area of the Western Forest Complex, Thailand

Fig. 3. Presence-absence binary models for RNH distributions after introducing equal sensitivity-specific threshold values to the continuous MaxEnt model based on the smallest home range size of the male RNH #15, as derived from Tifong (2007), during: a, the breeding season; b, the non-breeding season; and c, showing the combined habitat classification.

opencc-by-4.0Feb 2014View details →
zenodo40/100

Competition between predicting mathematical models and laboratory results for covid 19 after vaccination in Iran

<p>All of us like to find a way to end Covid 19. Specially, economy confronts with many problems. In attached JPG, i bring two predictions for covid 19 pandemic in Iran. Mathematical model predicts that after a fall in figure we will have a peak. However, right now, most of people in Iran used of vaccines. In addition, government forced on all to get vaccine. Even, students with 12 to 18 years old get vaccine. It has been heard that soonly, kids with ages between 3-12 years old will receive vaccine. As a man or woman, all of us like that this program will response and vaccines work. However, predictions of math show reverse result. Although, maybe, we should enter the factor of vaccine in mathematical model. It is good opportunity for scientists to examine response of covid 19 to program of vaccine for all. Even if we have a peak, however its height be smaller, we can say that vaccines act. Hope for ending Covid 19 in all countries.</p>

opencc-by-4.0Oct 2021View details →
dryad40/100

Airflow modelling predicts seabird breeding habitat across islands

<p>Wind is fundamentally related to shelter and flight performance: two factors that are critical for birds at their nest sites. Despite this, airflows have never been fully integrated into models of breeding habitat selection, even for well-studied seabirds. Here we use computational fluid dynamics to provide the first assessment of whether flow characteristics (including wind speed and turbulence) predict the distribution of seabird colonies, taking common guillemots (<em>Uria aalge</em>) breeding on Skomer island as our study system. This demonstrates that occupancy is driven by the need to shelter from both wind and rain/ wave action, rather than airflow characteristics alone. Models of airflows and cliff orientation both performed well in predicting high quality habitat in our study site, identifying 80% of colonies and 93% of avoided sites, as well as 73% of the largest colonies on a neighbouring island. This suggests generality in the mechanisms driving breeding distributions, and provides an approach for identifying habitat for seabird reintroductions considering current and projected wind speeds and directions.</p>

opencc-zeroOct 2021View details →
dryad40/100

Data for: Modeling the transition of death assemblages through the mixed layer predicts a downcore increase in time averaging

<p>Understanding how time averaging changes during the burial is essential for using Holocene and Anthropocene cores to analyze ecosystem change, given the many ways in which the time averaging affects biodiversity measures. Here, we use transition-rate matrices to explore how time averaging changes downcore when shells transit through a taphonomically-complex mixed layer into permanently-buried historical layers: this is a null model, without any temporal changes in rates of sedimentation or bioturbation, to contrast with downcore patterns that might be produced by human activity. Assuming stochastic burial and exhumation movements of shells between increments within the mixed layer and stochastic disintegration within increments, almost all combinations of net sedimentation, mixing, and disintegration produce a downcore increase in time averaging (interquartile range, IQR), typically associated with a decrease in kurtosis and skewness and with a shift from right-skewed to symmetrical age distributions. A downcore increase in time averaging is a null expectation wherever bioturbation generates an internally-structured mixed layer (i.e., a surface well-mixed layer is underlain by an incompletely-mixed layer), so that shells are mixed throughout the entire mixed layer at slower rate than they are buried below it by sedimentation. This downcore trend created by mixing is further amplified by the downcore decline in disintegration rate. Using data from the southern California shelf, we find that transition-rate matrices accurately reproduce the downcore changes in IQR, skewness, and kurtosis observed in sediment cores. The right-skewed distributions typical of surface death assemblages – the focus of most actualistic research – might be fossilized under exceptional conditions of episodic anoxia or sudden burial. However, such right-skewed assemblages will not typically transfer into subsurface historical layers and thus will be geologically transient. The deep-time fossil record will be dominated instead by more time-averaged assemblages with weakly skewed age distributions that form in the lower parts of the mixed layer.</p>

opencc-zeroNov 2022View details →
zenodo40/100

Protein structure model predictions for secreted fungal proteins

<p><strong>Dataset A - Alphafold2 prediction output data for 753 secreted proteins of <em>Rhizophagus irregularis </em>DAOM197198</strong>.&nbsp;Gene IDs are taken from the annotation by Yildirir et al. 2021,&nbsp;<a href="https://doi.org/10.1111/nph.17842">doi.org/10.1111/nph.17842</a></p> <p><strong>Dataset B - Alphafold2 prediction output data for 10 fungal effectors.</strong><strong>&nbsp;</strong>These are nine effectors from&nbsp;<em>Fusarium oxysporum</em>&nbsp;f. sp.<em>&nbsp;lycopersici</em>&nbsp;and RiSLM from&nbsp;<em>Rhizophagus irregularis</em>&nbsp;as well as their amino acid sequences. Signal peptides and sequences preceding a predicted Kex2 processing site were removed.</p> <p><strong>Dataset C - Alphafold2 prediction output data for 454 matches of a MycFOLD-HMM search</strong>&nbsp;across the Mycocosm genome database (<a href="https://mycocosm.jgi.doe.gov/mycocosm/home">https://mycocosm.jgi.doe.gov/mycocosm/home</a>) and 36 Glomeromycotina fungal genomes.</p>

opencc-by-4.0Dec 2022View details →
dryad40/100

Data from: Predicting primate-parasite associations with exponential random graph models

<p>Ecological associations between hosts and parasites are influenced by host exposure and susceptibility to parasites, and by parasite traits, such as transmission mode. Advances in network analysis allow us to answer questions about the causes and consequences of traits in ecological networks in ways that could not be addressed in the past.</p> <p>We used a network-based framework (exponential random graph models, or ERGMs) to investigate the biogeographic, phylogenetic, and ecological characteristics of hosts and parasites characteristics that affect the probability of interactions among nonhuman primates and their parasites. Parasites included arthropods, bacteria, fungi, protozoa, viruses, and helminths.</p> <p>We investigated existing hypotheses, along with new predictors and an expanded host-parasite database that included 213 primate nodes, 763 parasite nodes, and 2,319 edges among them. Analyses also investigated phylogenetic relatedness, sampling effort, and spatial overlap among hosts.</p> <p>In addition to supporting some previous findings, our ERGM approach demonstrated that more threatened hosts had fewer parasites, and notably, that this effect was independent of threatened hosts also having a smaller geographic range. Despite having fewer parasites, threatened host species shared more parasites with other hosts, consistent with the loss of specialist parasites and threats arising from generalist parasites that can be maintained in other, non-threatened hosts. Viruses, protozoa, and helminths had broader host ranges than bacteria or fungi, and parasites that infect non-primates had a higher probability of infecting more primate species.</p> <p>The value of the ERGM approach for investigating the processes structuring host-parasite networks provided a more complete view of the biogeographic, phylogenetic, and ecological traits that influence parasite species richness and parasite sharing among hosts. The results supported some previous analyses and revealed new associations that warrant future research, thus revealing how hosts and parasites interact to form ecological networks.</p>

opencc-zeroJan 2023View details →
zenodo40/100

US-XGB Models for Defect Prediction

<p>These files contain two pickle models featuring thousands of machine learning models drawn from the Bug Prediction and Jureczko datasets. These models were trained on extensive data and are suitable for use in bug prediction research data. The models are easy to employ and can be integrated into existing systems with minimal effort. We believe that these models will be a valuable resource for researchers and practitioners working in the field of machine learning and software engineering.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Dataset for: "Reducing OpenMP to FPGA Round-trip Times with Predictive Modelling"

<p>This archive contains the samples generated for the conference paper &quot;Reducing OpenMP to FPGA Round-trip Times with Predictive Modelling&quot; (In Proc. 18th Intl. Workshop on OpenMP (IWOMP), Chattanooga, TN, Sept. 2022, Springer LNCS vol. 13527, pp. 94&ndash;108,&nbsp;<a href="https://doi.org/10.1007/978-3-031-15922-0_7">https://doi.org/10.1007/978-3-031-15922-0_7</a>).</p> <p><strong>Abstract:</strong>&nbsp;Recent works aimed at expanding the target offloading capabilities of OpenMP to FPGA platforms. While enabling the easy construction of heterogeneous systems, the approach has to face a major hurdle: by blurring the line between software and hardware development, it forces software developers to consider hardware limitations. This can be difficult through the abstractions that OpenMP introduces over the generated hardware. The high level synthesis tools used by OpenMP compilers to generate hardware already offer predictions on hardware usage. Their value for OpenMP offloading however is questionable. This paper is based on the data mining we conducted on thousands of kernel variations. It demonstrates and proofs under which circumstances these predictions can be trusted in the context of OpenMP to FPGA offloading and concludes by showing how to derive runtime performance predictions from them. The model we present can be used without experience in hardware development and quickly predicts runtime on our benchmarks with an average Pearson correlation of 0.897. This knowledge allows developers to make fast, informed design decisions.</p>

openother-openSep 2022View details →
zenodo40/100

Supplementary Material for Automated Kinetic Models to Predict the Flame Speeds of Halocarbons

<p>Supplementary material to accompany the paper &quot;Automated Kinetic Models to Predict the Flame Speeds of Halocarbons&quot; by&nbsp;Nora Khalil, Sevy Harris, Richard H. West, submitted to the&nbsp;13th U. S. National Combustion Meeting,&nbsp;Organized by the Central States Section of the Combustion Institute,&nbsp;March 19&ndash;22, 2023,&nbsp;College Station, Texas.</p> <p>RMG input files.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
dryad40/100

Predicting daily activity time through ecological niche modeling and microclimatic data

<p><span>1. </span><span>Climate temporality is a phenomenon that affects species' activity and distribution patterns across spatial and temporal scales. Despite the global availability of microclimatic data, their use to predict activity patterns and distributions remains scarce, particularly at fine temporal scales (e.g., &lt; month). Predicting activity patterns based on climatic data may allow us to foresee some of the consequences of climate change, particularly for ectothermic vertebrates. </span></p> <p><span>2. </span><span>The Gila monster exhibits marked daily and seasonal activity patterns linked to physiology and reproduction. Here we evaluate if ecological niche models fitted using microclimate data can predict temporal activity patterns using the Gila monster (<em>Heloderma suspectum</em>) as a study system. Further, we identified if the activity patterns are related to physiological constraints.</span></p> <p><span>3. </span><span>We used dated occurrences from museum specimens and human observations to generate and test ecological niche models using minimum-volume ellipsoids. We generated hourly microclimatic data for each occurrence site for ten years using the NicheMapR package. For ecological niche modeling, we compared the traditional seasonal approach versus a daily activity pattern strategy for model construction. We tested both using the omission rate of independent observations (citizen science data). Finally, we tested if unimodal and bimodal activity patterns for each season could be recreated through ecological niche modeling and if these patterns followed known physiological constraints.</span></p> <p><span>4. </span><span>The unimodal and bimodal activity patterns previously reported directly from tracking individuals across the year were recovered by using niche modeling and microclimate across the species' geographical range. We found that upper thermal tolerances can explain the daily activity patterns of this species. </span></p> <p><span>5. </span><span>We conclude that ecological niche models trained with microclimatic data can be used to predict activity patterns at fine temporal scales, particularly on ectotherm species of arid zones coping with rapid climate modifications. Further, the use of fine temporal scale variables can lead to a better niche delimitation, enhancing the results of any research objective that uses correlative models.</span></p>

opencc-zeroFeb 2023View details →
zenodo40/100

MAGs and gapseq models for auxotrophy predictions in the human gut microbiome

<p>This dataset contains MAGs, their DNA sequence, genome statistics, quantification per sample, and their metabolic model reconstructions from two human population cohorts from northern Germany.</p>

opencc-by-4.0Feb 2023View details →
dryad40/100

Data for: Predicting age and mass at maturity from feeding behavior and diet in M. sexta: An empirical test of a life history model

<p>Feeding for most animals involves bouts of active ingestion alternating with bouts of no ingestion. In insects, the temporal patterning of bouts varies widely with resource quality and is known to affect growth, development time, and fitness. However, the precise impacts of resource quality and feeding behavior on insect life history traits is poorly understood. To explore and better understand the connections between feeding behavior, resource quality and insect life history traits, we combined laboratory experiments with a recently proposed mechanistic model of insect growth and development for a larval herbivore, <em>Manduca sexta</em>. We ran feeding trials for 4<sup>th</sup> and 5<sup>th</sup> instar larvae across different diet types (two hostplants and artificial diet) and used these data to parameterize a joint model of age and mass at maturity that incorporates both insect feeding behavior and hormonal activity. We found that the estimated durations of both feeding and non-feeding bouts were significantly shorter on low- than on high-quality diets. We then explored how well the fitted model predicted historical out-of-sample data on age and mass of <em>M</em>.<em> sexta</em>. We found that the model accurately described qualitative outcomes for the out-of-sample data, notably that a low-quality diet results in reduced mass and later age at maturity compared to high-quality diets. Our results clearly demonstrate the importance of diet quality on multiple components of insect feeding behavior (feeding and non-feeding), and partially validate a joint model of insect life history. We discuss the implications of these findings with respect to insect herbivory and discuss ways in which our model could be improved or extended to other systems.</p>

opencc-zeroMar 2023View details →
zenodo40/100

Training patches and prediction codes of deep learning (LANA) model for Landsat 8/9 cloud/shadow mask

<p>This dataset includes (i) the image patches dataset and (ii) application/prediction (not training) codes for Landsat 8 cloud and cloud shadow masking used in a paper in review and uploaded here: &nbsp;</p> <p>Hankui Zhang, Dong Luo, David Roy, A learning attention network algorithm (LANA) for accurate Landsat-8 cloud and shadow masking,&nbsp;<em>Remote Sensing of Environment</em>&nbsp;</p> <p>The documentation is in&nbsp;<a href="https://zenodo.org/api/files/5462baa5-2bba-4b0f-92aa-c17681b6464b/l8_training_data_readme_new.pdf?versionId=95253efb-6447-46d8-ae4b-ac0c93b43532">l8_training_data_readme_new.pdf</a>.&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Soil organic carbon models need independent time-series validation for reliable prediction

<p>Supplementary Data 1 to the paper: Soil organic carbon models need independent time-series validation for reliable prediction</p> <p>By: Le No&euml;, J., Manzoni, S., Abramoff, R.Z., B&ouml;lscher, T., Bruni, E., Cardinael, R., Ciais, P., Chenu, C., Clivot, H., Derrien, D., Ferchaud, F., Garnier, P., Goll, D., Lashermes, G., Martin, M.P., Rasse, D., Rees, F., Sainte-Marie, J., Salmon, E., Schiedung, M., Schimel, J., Wieder, W.R., Abiven, S., Barr&eacute;, P., C&eacute;cillon, L., Guenet, B.</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

MESMAR v1: A new regional coupled climate model for downscaling, predictability, and data assimilation studies in the Mediterranean region. Article data

<p>Regional coupled and Earth System models are fundamental numerical tools for climate investigations, downscaling of&nbsp;predictions and projections, process-oriented understanding of regional extreme events, and many more applications. Here we&nbsp;introduce a newly developed coupled regional modeling framework for the Mediterranean region, called MESMAR&nbsp;(Mediterranean Earth System model at ISMAR) version 1, which is composed of the WRF atmospheric model, the NEMO oceanic&nbsp;15 model, and the HD hydrological discharge model, coupled via the OASIS coupler. The model is implemented at moderate&nbsp;resolution (about 1/12&deg; for the ocean and river routing, while twice coarser for the atmosphere) for long-term investigations.</p> <p>The gzipped tarball contains data files contained in the manuscript associated with the MESMARv1 description and&nbsp;submitted to Geoscientific Model Developments:</p> <p>MESMAR v1: A new regional coupled climate model for downscaling,&nbsp;predictability, and data assimilation studies in the Mediterranean region</p> <p>by&nbsp;Andrea Storto, Yassmin Hesham Essa, Vincenzo de Toma, Alessandro Anav, Gianmaria Sannino,<br> Rosalia Santoleri, Chunxue Yang</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record