Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,773
datasets available to search
ShareScore release 0.9.0
Dataset results
1,773 results for “Predictive model”
Data from: Predictive power of food web models based on body size decreases with trophic complexity
Food web models parameterized using body size show promise to predict trophic Interaction Strengths (IS) and abundance dynamics. However, this remains to be rigorously tested in food webs beyond simple trophic modules, where indirect and intraguild interactions could be important and driven by traits other than body size. We systematically varied predator body size, guild composition and richness in microcosm insect webs and compared experimental outcomes with predictions of IS from models with allometrically scaled parameters. Body size was a strong predictor of IS in simple modules (r2=0.92), but with increasing complexity the predictive power decreased, with model IS being consistently overestimated. We quantify the strength of observed trophic interaction modifications, partition this into density-mediated vs. behaviour-mediated indirect effects and show that model shortcomings in predicting IS is related to the size of behaviour-mediated effects. Our findings encourage development of dynamical food web models explicitly including and exploring indirect mechanisms.
Data from: Patterns of male fitness conform to predictions of evolutionary models of late-life
We studied lifetime male virility, a male fitness component, in five populations of Drosophila melanogaster. Virility was measured as the number of females, out of eight total, that a male could fertilize in 24 hours. Individual males were measured at weekly intervals until they died. Virility declined in an approximately linear fashion for the first three weeks of adult life. It then stayed low but relatively constant for another three weeks, exhibiting a clear plateau. These observations are consistent with the evolutionary theories of late-life.The results were not consistent with a simple heterogeneity theory of late-life. This is the first demonstration of a late-life plateau for a male fitness component. We also found that the virility of males that were within seven days of death was significantly lower than similarly aged males that were not about to die. This rapid deterioration of virility prior to death, or death spiral, is similar to a decline in fecundity that we had previously documented.
Data from: Phylodynamic model adequacy using posterior predictive simulations
Rapidly evolving pathogens, such as viruses and bacteria, accumulate genetic change at a similar timescale over which their epidemiological processes occur, such that it is possible to make inferences about their infectious spread using phylogenetic time-trees. For this purpose it is necessary to choose a phylodynamic model. However, the resulting inferences are contingent on whether the model adequately describes key features of the data. Model adequacy methods allow formal rejection of a model if it cannot generate the main features of the data. We present TreeModelAdequacy (TMA), a package for the popular BEAST2 software, that allows assessing the adequacy of phylodynamic models. We illustrate its utility by analysing phylogenetic trees from two viral outbreaks of Ebola and H1N1 influenza. The main features of the Ebola data were adequately described by the coalescent exponential-growth model, whereas the H1N1 influenza data was best described by the birth-death SIR model.
Data from: Are Bitcoin bubbles predictable? Combining a generalized Metcalfe's law and the LPPLS model
We develop a strong diagnostic for bubbles and crashes in Bitcoin, by analyzing the coincidence (and its absence) of fundamental and technical indicators. Using a generalized Metcalfe's law based on network properties, a fundamental value is quantified and shown to be heavily exceeded, on at least four occasions, by bubbles that grow and burst. In these bubbles, we detect a universal super-exponential unsustainable growth. We model this universal pattern with the Log-Periodic Power Law Singularity (LPPLS) model, which parsimoniously captures diverse positive feedback phenomena, such as herding and imitation. The LPPLS model is shown to provide an ex-ante warning of market instabilities, quantifying a high crash hazard and probabilistic bracket of the crash time consistent with the actual corrections; although, as always, the precise time and trigger (which straw breaks the camel's back) being exogenous and unpredictable. Looking forward, our analysis identifies a substantial but not unprecedented overvaluation in the price of Bitcoin, suggesting many months of volatile sideways Bitcoin prices ahead (from the time of writing, March 2018).
Data from: A prediction model of compressor with variable geometry diffuser based on elliptic equation and Partial Least Squares
In order to fulfill more and more extensive intake air flow range of diesel engine, variable geometry compressor (VGC) is introduced into turbocharged diesel engine. However, due to the variable diffuser vanes angle (DVA), the prediction for the performance of VGC becomes more difficult than normal compressor. In the present study, a prediction model comprised of elliptical equation and PLS (Partial Least Squares) model was proposed to predict the performance of VGC. The speed lines of pressure ratio map and efficiency map with elliptical equation were fitted, and the coefficients of elliptical equation was introduced into PLS model to build the polynomial relationship between the coefficients and relative speed, DVA. And further, the maximal order of polynomical was detailed investigated to reduce the number of sub-coefficients and acceptable fit accuracy simultaneously. The prediction model was validated with sample data and in order to present the superiority in compressor performance prediction, the prediction results of this model were compared with those of look-up table and BPNN. The validation and comparison results show that the prediction accuracy of the new developed model is acceptable, and this model is much more suitable than look-up table and BPNN under the same condition in the VGA performance prediction. Moreover, the new developed prediction model provides a novel and effective prediction solution for VGC and can be used to improve the accuracy of the thermodynamic model for turbocharged diesel engines in the future.
Data from: Using striated tooth marks on bone to predict body size in theropod dinosaurs: a model based on feeding observations of Varanus komodoensis, the Komodo monitor
Mesozoic tooth marks on bone surfaces directly link consumers to fossil assemblage formation. Striated tooth marks are believed to form by theropod denticle contact, and attempts have been made to identify theropod consumers by comparing these striations with denticle widths of contemporaneous taxa. The purpose of this study is to test whether ziphodont theropod consumer characteristics may be accurately identified from striated tooth marks on fossil surfaces. There are three major objectives; 1) experimentally produce striated tooth marks and explain how they form; 2) determine whether body size characteristics are reflected in denticle widths; 3) determine whether denticle characters are accurately transcribed onto bone surfaces in the form of striated tooth marks. Controlled feeding trials were conducted with the dental analogue Varanus komodoensis (the Komodo monitor). Goat (Capra hircus) carcasses were introduced to captive, isolated individuals. Striated tooth marks were then identified, and striation width, number, and degree of divergence were recorded for each. Denticle widths and tooth/body size characters were taken from photographs and published accounts of both theropod and V. komodoensis skeletal material, and regressions were compared among and between the two groups. Striated marks tend to be regularly striated with a variable degree of branching, and may co-occur with scores. Striation morphology directly reflects contact between the mesial carina and bone surfaces during the rostral reorientation when defleshing. Denticle width is primarily influenced by tooth size, and correlates well with body size displaying negative allometry in both groups regardless of taxon or position. When compared, striation widths fall within or below the range of denticle widths extrapolated for similar sized V. komodoensis individuals. Striation width is directly influenced by the orientation of the carina during feeding, and may underestimate but cannot overestimate denticle width. Although body size may theoretically be estimated solely by a striated tooth mark under ideal circumstances, many caveats should be considered. These include the influence of negative allometry across taxa and throughout ontogeny, the existence of theropods with extreme denticle widths, and the potential for striations to underestimate denticle widths. This method may be useful under specific circumstances, especially for establishing a lower limit body size for potential consumers.
Data from: Using viromes to predict novel immune proteins in non-model organisms
Immunity is mostly studied in a few model organisms, leaving the majority of immune systems on the planet unexplored. To characterize the immune systems of non-model organisms alternative approaches are required. Viruses manipulate host cell biology through the expression of proteins that modulate the immune response. We hypothesized that metagenomic sequencing of viral communities would be useful to identify both known and unknown host immune proteins. To test this hypothesis, a mock human virome was generated and compared to the human proteome using tBLASTn, resulting in 36 proteins known to be involved in immunity. This same pipeline was then applied to reef-building coral, a non-model organism that currently lacks traditional molecular tools like transgenic animals, gene-editing capabilities, and in vitro cell cultures. Viromes isolated from corals and compared with the predicted coral proteome resulted in 2503 coral proteins, including many proteins involved with pathogen sensing and apoptosis. There were also 159 coral proteins predicted to be involved with coral immunity but currently lacking any functional annotation. The pipeline described here provides a novel method to rapidly predict host immune components that can be applied to virtually any system with the potential to discover novel immune proteins.
SpatPPI: a geometric deep learning model for predicting protein-protein interactions involving intrinsically disordered regions
Open the record for dataset details and reuse information.
Merged HLS2 (L30), ERA5-Land inputs and sample predictions of land surface temperature for the IBM granite-geospatial-land-surface-temperature model
<p>This dataset contains merged Harmonized Landsat-Sentinel 2 (HLS2) (L30 only) and ERA5-Land data following the UTM CRS:WGS84. It has been assembled for predicting land surface temperature with a fine-tuned granite geospatial foundation model developed by IBM Research. In addition, we include sample predictions of land surface temperature derived from this model. Please see https://huggingface.co/ibm-granite/granite-geospatial-land-surface-temperature for more information on data preparation and model use.</p> <p><strong>HLS:</strong></p> <p>Masek, J., J. Ju, J. Roger, S. Skakun, E. Vermote, M. Claverie, J. Dungan, Z. Yin, B. Freitag, C. Justice. HLS Sentinel-2 MSI Surface Reflectance Daily Global 30m v2.0. 2021, distributed by NASA EOSDIS Land Processes DAAC, https://doi.org/10.5067/HLS/HLSS30.002 </p> <h4><strong>ERA5-Land:<br></strong></h4> <p>Copernicus Climate Change Service, Climate Data Store, (2024): ERA5-land post-processed daily-statistics from 1950 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI: <a href="https://doi.org/10.24381/cds.e9c9c792">10.24381/cds.e9c9c792</a> (Accessed on 05-11-2024)</p> <h4><strong>LST-predictions:</strong></h4> <p>These predictions of land surface temperature are derived from the IBM granite-geospatial-land-surface-temperature model and have been made available for Abidjan, Côte d’Ivoire and Johannesburg, South Africa for the period 2013-2023. </p> <h4>Attribution</h4> <p>Copernicus programme:</p> <p>Contains modified Copernicus Climate Change Service information [2024]. Neither the European Commission nor ECMWF is responsible for any use that may be made of the Copernicus information or data it contains.</p> <p><strong>Data</strong></p> <p>Muñoz Sabater, J., Comyn-Platt, E., Hersbach, H., Bell, B., Berrisford, P., Biavati, G., Horányi, A., Muñoz Sabater, J., Nicolas, J., Peubey, C., Radu, R., Rozum, I., Schepers, D., Simmons, A., Soci, C., Dee, D., Thépaut, J-N., Cagnazo, C., Cucchi, M. (2024): ERA5-land post-processed daily-statistics from 1950 to present. Copernicus Climate Change Service (C3S) Climate Data Store (CDS), DOI: <a href="https://doi.org/10.24381/cds.e9c9c792">10.24381/cds.e9c9c792</a> (Accessed on 05-11-2024)</p>
DeepGO-SE protein function prediction model data
<p>Data for training and running DeepGO-SE protein function prediction model</p>
About epiTCR-KDA: Knowledge Distillation model on Dihedral Angles for TCR-peptide prediction
Open the record for dataset details and reuse information.
Data Set for Predicting the Performance of ATL Model Transformations Based on Generated Models
<p>Predicting the execution time of model transformations can help to understand how a transformation reacts to a given input model without creating and transforming the respective model.</p> <p>In our previous data set (https://doi.org/10.5281/zenodo.8385957), we have documented our experiments in which we predict the performance of ATL transformations using predictive models obtained from training linear regression, random forest and support vector regression. As input for the prediction, our approach uses a characterization of the input model. In these experiments, we only used data from real models.</p> <p>However, a common problem is that transformation developers do not have enough models available to use such a prediction approach. Therefore, in a new variant of our experiments, we investigated whether the three considered machine learning approaches can predict the performance of transformations if we use data from generated models for training. We also investigated whether it is possible to achieve good predictions with smaller training data. The dataset provided here offers the corresponding raw data, scripts, and results.</p> <p>A detailed documentation is available in documentaion.pdf.</p>
Supplementary material 1 from: Keck F, Hürlemann S, Locher N, Stamm C, Deiner K, Altermatt F (2022) A triad of kicknet sampling, eDNA metabarcoding, and predictive modeling to assess richness of mayflies, stoneflies and caddisflies in rivers. Metabarcoding and Metagenomics 6: e79351. https://doi.org/10.3897/mbmg.6.79351
Figures S1–S5
Coupling deep learning and physically-based hydrological models for monthly streamflow predictions
<p>Revision in journal Water Resources Research, Paper # <strong><span>2023WR035618R</span></strong></p>
Dispersal patterns and potential distribution prediction of three rice planthopper species in China based on the ensemble model
<p>Emergence of three rice planthopper species in China from 1993 to 2022.</p>
Data from: Distribution models predict climate-related range alteration or extinction of eleven threatened tropical rainforest trees in the Western Ghats
<p>This dataset contains information related to species occurence data and species distribution modeling (SDM) analysisr of eleven threatened tree species. Occurrences are compiled from extensive field surveys in the Anamalai Hills along with data from the Global Biodiversity Information Facility (GBIF.org) and earlier work done within the southern Western Ghats, India.</p> <p>References:<br>Page, N. V., & Shanker, K. (2020). Climatic stability drives latitudinal trends in range size and richness of woody plants in the Western Ghats, India. PLOS ONE, 15(7), e0235733. https://doi.org/10.1371/journal.pone.0235733</p> <p>GBIF.org (2022) GBIF Occurrence Download, 2 August 2022. DOI:10.15468/dl.gnvuxj</p> <p><br>AUTHOR #1<br>1. Name: A.P. Madhavan<br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Email address: madhavan@ncf-india.org<br>4. ORCID: https://orcid.org/0009-0009-2754-8256</p> <p>AUTHOR #2<br>1. Name: Kshama Bhat<br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Email address: kshama@ncf-india.org<br>4. ORCID: ORCID: https://orcid.org/0000-0002-6190-2687</p> <p>AUTHOR #3<br>1. Name: Srinivasan Kasinathan<br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Email address: srini@ncf-india.org<br>4. ORCID: https://orcid.org/0000-0001-7323-6653</p> <p>AUTHOR #4<br>1. Name: Divya Mudappa <br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Email address: divya@ncf-india.org <br>4. ORCID: https://orcid.org/0000-0001-9708-4826</p> <p>AUTHOR #5<br>1. Name: Navendu Page<br>2. Work Address: Wildlife Institute of India, Post Box No. 18, Chandrabani, Dehradun, Uttarakhand 248001, India<br>3. Email address: navendu.page@gmail.com<br>4. ORCID: ORCID: https://orcid.org/0000-0002-9413-7571</p> <p>AUTHOR #6<br>1. Name: T. R. Shankar Raman <br>2. Work Address: Nature Conservation Foundation, 1311, 12th A Main, Vijayanagar 1st Stage, Mysuru 570017, Karnataka, India<br>3. Email address: trsr@ncf-india.org <br>4. ORCID: https://orcid.org/0000-0002-1347-3953</p> <p>Keywords: tropical rainforest, climate change, tree distributions, species distribution models, range shifts, Western Ghats</p> <p><br>Geographic Coverage:<br>1. Location/Study Area: Southern Western Ghats Montane Rain Forests, Southern Western Ghats Moist Deciduous Forests, India<br>2. GPS coordinates: SWG (73.95° – 80.33° E, 8.06° – 13.11°N) </p> <p>Temporal coverage<br>Starts: 2020-08-01<br>Ends: 2024-03-28</p> <p>Besides this README.txt file, the dataset includes three comma-delimited text files (csv); two R scripts, and 1 kml file of surveyed trails.</p> <p>CSV files with the data in columns as explained below:</p> <p>1) Focal_Tree_Dat.csv</p> <p>Comp: Number identifier<br>FT_ID: Unique tree no for each individual<br>Focal_tree: Scientific name of species<br>Date: Date of occurrence observation<br>Place: Area/locality description<br>Trail: Unique trail ID<br>Waypoint: Waypoint number <br>Time: Time in hh:mm format <br>Location: Specific description of occurrence locality <br>Latitude: Latitude in decimal degrees N <br>Longitude: Longitude in decimal degrees E <br>Elevation: Elevation in metres <br>Slope: Cateory of slope <br>ID_Notes: Notes on identification<br>Phenophase: Phenophase expression at the time of observation <br>GBH: Girth at breast height in centimetres (comma separated list of numbers in case of multi-stemmed trees) <br>Tree_ht: Tree height in metres<br>Canopy_ht: Maximimum height of the surrounding canopy in metres<br>Substrate: Soil substrate composition<br>Invasives: Name of invasive species (if present) <br>Stature: Vegetation strata position <br>Relatively: Stature of focal individual relative to other surrounding individuals <br>Deadwood: Description of deadwood on the tree <br>Damage: Description of damage on the bole <br>Shape: Description of tree canopy shape<br>Closure: Canopy closure at focal tree <br>Seedlings: Number of conspecific seedlings present in 5 m radius of focal tree <br>Saplings: Number of conspecific saplings present in 5 m radius of focal tree<br>Trees: Number of conspecific trees present in 5 m radius of focal tree<br>Remarks: Remarks </p> <p>2) Ffspecies.csv</p> <p>Source: Source of occurrence <br>ID: State/location of occurrence<br>Region: Biogeographic region of occurrence <br>decimalLatitude: Latitude in decimal degrees N<br>decimalLongitude: Longitude in decimal degrees E<br>species: Scientific name of species</p> <p>4) ft_surveys.csv</p> <p>Date: Date of survey of sample trail<br>Prot_type: Category indicating whether protected area or fragment <br>Place: Area/locality description<br>Route_description: Specific landmark description of trail<br>Trail: Unique trail ID <br>Trail_distance: Tracked distance of trail in km <br>Corrected_trail_distance: Corrected distance of trail in km<br>Track_filename_kml: File name of gps track<br>Sample_collected: Name of species if sample collected <br>Observers: Name of observers <br>Remarks: Remarks</p> <p>ANALYSES SCRIPTS<br>flexsdm_script.R<br>Script containing the analysis of all maxent distribution modeling and associated analysis</p> <p>Franklinia_density.Rmd<br>Script of density and abundance related analysis</p> <p> </p>
Dataset [ref. paper "Predictive modeling of drivers' brake reaction time through machine learning methods"]
Open the record for dataset details and reuse information.
Detailed Mapping: Standardizing Heat-Related Diagnoses for Predictive Modeling in Healthcare
Open the record for dataset details and reuse information.
DeepBacs – S. aureus SIM prediction dataset and CARE model
<p>Training and test images of live, membrane-labeled <em>S. aureus </em>cells for prediction of SIM super-resolution images from widefield images, as well as a trained CARE model.</p> <p>Additional information can be found on this <a href="https://github.com/HenriquesLab/DeepBacs/wiki">github wiki</a>.</p> <p>The example image shows a widefield fluorescence image and SIM reconstruction of Nile Red labelled, live <em>S. aureus </em>cells.</p> <p> </p> <p><strong>Training and test dataset</strong></p> <p><strong>Data type</strong>: Paired microscopy images (fluorescence) of low (widefield) and high resolution (SIM)</p> <p><strong>Microscopy data type</strong>: Fluorescence microscopy (Nile Red)</p> <p><strong>Microscope</strong>: GE HealthCare Deltavision OMX system (with temperature and humidity control, 37°C) equipped with an Olympus 60x 1.42NA Oil immersion objective and 2 PCO Edge 5.5 sCMOS cameras (one for DIC, one for fluorescence)</p> <p><strong>Cell type</strong>: <em>S. aureus</em> strain JE2 grown under agarose pads</p> <p><strong>File format</strong>: .tif (16-bit for widefield images and 32-bit for SIM reconstructions)</p> <p><strong>Image size</strong>: 1024 x 1024 px² (40 nm/px)<br> <strong>Image preprocessing</strong>: <em>S. aureus</em> widefield images were scaled with a factor of 2 to match the SIM reconstruction pixel size. </p> <p> </p> <p><strong>CARE model</strong></p> <p>The CARE 2D model was generated using the ZeroCostDL4Mic platform (Chamier et al., 2021). It was trained from scratch for 300 epochs on 9400 paired image patches (image dimensions: (1024 x 1024 px²), patch size: (80 x 80 px²), 100 patches/image) with a batch size of 8 and a laplace loss function, using the CARE 2D ZeroCostDL4Mic notebook (v 1.12). Key python packages used include tensorflow (v 0.1.12), Keras (v2.3.1), csbdeep (v 0.6.1), numpy (v 1.19.5), cuda (v 10.1.243). The training was accelerated using a Tesla P100GPU and data was augmented by a factor of 4 using rotation, flipping and random zoom.</p> <p>Model weights can be used with the ZeroCostDL4Mic CARE 2D notebook or the CSBDeep Fiji plugin.</p> <p> </p> <p><strong>Author(s)</strong>: Pedro Matos Pereira<sup>1,2</sup>, Mariana Pinho<sup>1,3</sup></p> <p><strong>Contact email</strong>: <a href="mailto:pmatos@itqb.unl.pt">pmatos@itqb.unl.pt</a> and <a href="mailto:mgpinho@itqb.unl.pt">mgpinho@itqb.unl.pt</a></p> <p> </p> <p><strong>Affiliation</strong>: </p> <p>1) Bacterial Cell Biology, Instituto de Tecnologia Química e Biológica António Xavier, Universidade Nova de Lisboa, Oeiras, Portugal</p> <p>2) ORCID: https://orcid.org/0000-0002-1426-9540</p> <p>3) ORCID: https://orcid.org/0000-0002-7132-8842</p>
Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu
<p>This is supporting data for the paper titled "Machine Learning Mid-Infrared Spectral Models for Predicting Modal Mineralogy of CI/CM Chondritic Asteroids and Bennu". Figure S1 compares model performance of nonnegative LSMA and PLS from Pan et al. (2015). Tables S1 through S5 provide XRD data, metadata, MIR spectra using the laboratory conversion, MIR spectra using the OTES conversion and quantitative XRD results for Murchison meteorite. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.