Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
170
datasets available to search
ShareScore release 0.9.0
Dataset results
170 results for “predictive mapping”
Predicted soil organic carbon stock at 30 m in t/ha for 0-100 cm depth global / update of the map of mangrove forest soil carbon
<p>This is the 2nd update of maps produced by <a href="https://doi.org/10.1088/1748-9326/aabe1c">Sanderman et al (2018)</a>. The improvements to the <a href="https://opengeohub.github.io/spatial-prediction-eml/spatiotemporal-prediction-of-soil-organic-carbon.html">3D spatial predictions</a> include:</p> <ul> <li> <p>new updated global mangrove coverage map (contact Thomas Worthington),</p> </li> <li> <p>spatiotemporal predictions to account for differences in spectral reflectance at the time of field work,</p> </li> <li> <p>additional SOC points <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41558-018-0162-5/MediaObjects/41558_2018_162_MOESM2_ESM.xlsx">published in Rovai et al. (2018)</a> used in model training (see gpkg file).</p> </li> </ul> <p>To open map in QGIS or similar, drag and drop the *.tif files. You can than add also the gpkg file contain the training points.</p> <p>Production steps (ensemble predictions using SuperLearner) are explained in detail at: </p> <ul> <li>R code: <a href="https://github.com/whrc/Mangrove-Soil-Carbon/">https://github.com/whrc/Mangrove-Soil-Carbon/</a> (see "R_code/GMW_mangroves_SOC_30m.R")</li> <li>Tutorial: <a href="https://envirometrix.github.io/PredictiveSoilMapping/soilmapping-using-mla.html#ensemble-predictions-using-superlearner-package">"Predictive Soil Mapping with R"</a></li> </ul> <p>Produced for the purpose of Mangrove Restoration Potential Map funded by The Nature Conservancy and IUCN. Contact TNC: Emily Landis <<a href="mailto:elandis@TNC.ORG">elandis@TNC.ORG</a>>. Contact IUCN / University of Cambridge: Thomas Worthington <<a href="mailto:taw52@cam.ac.uk">taw52@cam.ac.uk</a>>.</p> <ul> <li>The mangrove restoration potential map is available at: <a href="https://www.researchgate.net/deref/http%3A%2F%2Fmaps.oceanwealth.org%2Fmangrove-restoration%2F">http://maps.oceanwealth.org/mangrove-restoration/</a></li> </ul>
Predicted maps
<p>This dataset is the final product of research through the projects ANTARES (grant agreement No. 739570) and CYBELE (grant agreement No. 825355). The dataset consists of yield, protein content and selective harvesting soya maps at a resolution of 10 m. These maps were created by satellite images and soil properties data using machine learning algorithms. Maps are located in the Upper Austria region. Files with the name of "map yield" contain information about yield amount per pixel, while files with "map protein" denote parcels with predicted protein content. Also, the same protein map files contain an additional class column. Class 1 indicates pixels where soya have good quality (protein content > 41), while class 2 represents poorer quality.</p>
Stochastic Occupancy Grid Map Prediction in Dynamic Scenes: Dataset
<p>Three occupancy grid map (OGM) datasets for the paper titled "Stochastic Occupancy Grid Map Prediction in Dynamic Scenes" by Zhanteng Xie and Philip Dames</p> <p>1. OGM-Turtlebot2: collected by a simulated Turtlebot2 with a maximum speed of 0.8 m/s navigates around a lobby Gazebo environment with 34 moving pedestrians using random start points and goal points</p> <p>2. OGM-Jackal: extracted from two sub-datasets of the socially compliant navigation dataset (SCAND), which was collected by the Jackal robot with a maximum speed of 2.0 m/s at the outdoor environment of the UT Austin</p> <p>3. OGM-Spot: extracted from two sub-datasets of the socially compliant navigation dataset (SCAND), which was collected by the Spot robot with a maximum speed of 1.6 m/s at the Union Building of the UT Austin</p> <p>The relevant code is available at: <br> OGM prediction: https://github.com/TempleRAIL/SOGMP<br> OGM mapping with GPU: https://github.com/TempleRAIL/occupancy_grid_mapping_torch</p>
Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks
<p>Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.</p>
Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks
<p>Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.</p>
Predictive high-resolution mapping of sea floor rock cover for the UK and Ireland
<p>Predicted seabed rock cover using the machine learning algorithm Catboost and marine environmental predictors.</p> <p>Supporting data for T4.1 of Horizon 2020 project FutureMARES.</p>
Predicted Surficial Blue Carbon Maps for Blackbird Creek and St. Jones River Tidal Salt Marshes using 2014-2023 Landsat-8 OLI records
These data tables reflect the predictions of a gradient boosted trees model for predicting soil organic matter (SOM) in the surficial layer of tidal marsh soils. The model was trained on soil core data related to its corresponding spectral characteristics from decadal Landsat-8 Operational Land Imager data (see associated publication Warner et al. "Leveraging a decade of Landsat-8 spectral records for mapping blue carbon storage in tidal salt marshes"). This is a modeled data product designed to illustrate spatial patterns of organic matter storage as predicted by satellite data.
Mapping present and future predicted distribution patterns for a meso-grazer guild in the Baltic Sea
<p>Baltic Sea communities consisting of key and endemic species are threatened by climate change. Using Ecological niche modelling, we map predicted distribution patterns under recent and future climate change scenarios (2050) for a food-web consisting of a guild of meso-grazers (Idotea spp.), their host algae (Fucus vesiculosus and F. radicans) and their fish predator (Gasterosteus aculeatus). Brackish water species depend on two important abiotic factors: temperature and salinity. We assess which of these environmental factors determines the distribution limits of the grazers in the Baltic Sea today. For species in a semi-enclosed sea area such as the Baltic Sea, climate-induced changes may lead to dramatic food-web effects. We assess the consequences of the predicted climate-induced habitat range changes for this unique Baltic community.<br /> </p>
Data: Spatio-temporal prediction of soil moisture using soil maps, topographic indices and SMAP retrievals
<p><strong>Data used in:</strong></p> <p>Schönauer, M., Prinz, R., Väätäinen, K., Astrup, R., Pszenny, D., Lindeman, H., et al. (2022). Spatiotemporal prediction of soil moisture using soil maps, topographic indices and SMAP retrievals. <em>International Journal of Applied Earth Observation and Geoinformation</em>, 102730. doi: 10.1016/j.jag.2022.102730</p>
Prince Edward Island (Canada) Predictive Soil Mapping data set (30 m)
<p>Prince Edward Island (Canada) Predictive Soil Mapping data set. Training points include:</p> <ul> <li>soil organic matter (624 points): unit: percent, g/100g ) samples from the topsoil (0-23 cm depth),</li> <li>soil types (672 points),</li> </ul> <p>Covariate layers include:</p> <ul> <li>DEM derivatives (Channel_Network_Base_Level.tif, MRRTF.tif, Relative_Slope_Position.tif, Slope.tif, TWI.tif, Valley_Depth.tif, Vertical_Distance_To_Channel_Network.tif),</li> <li>landcover_2016_reclassify.tif (categorical values),</li> <li>DSS_soil_polygons.tif (soil polygons),</li> </ul> <p>Some covariates are type numeric, some type factor. To use the pre-processed data download only the RDS file. E.g. "PEI100m.soil.rds" contains all covariates layers resampled to 100 m resolution and all training points.</p>
ECOCLIMAP-SG-ML: an ensemble land cover map for numerical weather prediction
<p>This dataset contains ensemble land cover maps for numerical weather prediction at 60 m resolution over Europe. As they were<br>generated thanks to machine learning, the weights and the training data are also provided.</p>
Impact of Schistosomiasis, Soil-Transmitted Helminthiasis and Anaemia on preschool and school-age children's health condition: post treatment predictive mapping in Benin Republic.
<p>This dataset provides information about the epidemiology of schistosomiasis, soil transmitted helminthiasis and anemia alongside malnutrition among preschool and school age children in Ouake and Bembereke districts of donga and borgou departments in Benin republic.</p>
Soil texture dataset from the publication: "Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula'
<p>Clay, silt and sand distribution in Antarctic soils modeled and predicted through Machine Learning approaches, legacy soil data and environmental covariates. The coefficient of variation and quantile data represent the spatial uncertainty of the predictions. For more information about the methodology used, users are referred to the article: </p> <p>Siqueira, R.G., Moquedace, C.M., Francelino, M.R., Schaefer, C.E.G.R., Fernandes-Filho, E.I., 2023. Machine learning applied for Antarctic soil mapping: Spatial prediction of soil texture for Maritime Antarctica and Northern Antarctic Peninsula. Geoderma 432, 116405. https://doi.org/10.1016/j.geoderma.2023.116405</p> <p>The .zip file has the following folders:</p> <p>1) soil_texture_antarctica: soil texture information containing clay, silt and sand contents</p> <p>2) soil_texture_coefficient_variation: uncertainty from the coefficient of variation of the soil texture prediction</p> <p>3) soil_texture_prediction_interval: uncertainty from the prediction interval 90% (Q95% - Q5%) of the soil texture prediction</p> <p>4) soil_texture_quantile05: quantile 5% of the soil texture prediction</p> <p>5) soil_texture_quantile95: quantile 95% of the soil texture prediction</p>
Paper Repository and References for "Early software defect prediction: A systematic map and review"
<p>Context: Software defect prediction is a trending research topic, and a wide variety of the published papers focus on coding phase or after. A limited number of papers, however, includes the prior (early) phases of the software development lifecycle (SDLC).<br> Objective: The goal of this study is to obtain a general view of the characteristics and usefulness of Early Software Defect Prediction (ESDP) models reported in scientific literature. <br> Method: A systematic mapping and systematic literature review study has been conducted. We searched for the studies reported between 2000 and 2016. We reviewed 52 studies and analyzed the trend and demographics, maturity of state-of-research, in-depth characteristics, success and benefits of ESDP models. <br> Results: We found that categorical models that rely on requirement and design phase metrics, and few continuous models including metrics from requirements phase are very successful. We also found that most studies reported qualitative benefits of using ESDP models.<br> Conclusion: We have highlighted the most preferred prediction methods, metrics, datasets and performance evaluation methods, as well as the addressed SDLC phases. We expect the results will be useful for software teams by guiding them to use early predictors effectively in practice, and for researchers in directing their future efforts.</p>
Meaning maps and saliency models based on deep convolutional neural networks are insensitive to image meaning when predicting human fixations - data
<p>Data from the paper:<em> Meaning maps and saliency models based on deep convolutional neural networks are insensitive to image meaning when predicting human fixations.</em></p> <p>Preprint: https://www.biorxiv.org/content/10.1101/840256v1</p> <p>Marek A. Pedziwiatr<br> marek.pedziwi@gmail.com<br> September 2020</p> <p> </p>
Data from: Abiotic proxies for predictive mapping of near-shore benthic assemblages: implications for marine spatial planning
Marine spatial planning (MSP) should assist managers in guiding human activities towards sustainable practices and in minimizing user-conflicts in our oceans. A necessary first step is to quantify spatial patterns of marine assemblages in order to understand the ecosystem's structure, function, and services. However, the large spatial scale, high economic value, and density of human activities in near-shore habitats often makes quantifying this component of marine ecosystems especially daunting. To address this challenge, we developed an assessment method that employs abiotic proxies to rapidly characterize marine assemblages in near-shore benthic environments with relatively high resolution. We evaluated this assessment method along 300 km of the State of Maine's coastal shelf (< 100m depth)—a zone where high densities of buoyed lobster traps typically preclude extensive surveys by towed sampling gear (i.e., otter trawls). During the summer months of 2010-2013, we implemented a stratified-random survey using a small remotely operated vehicle that allowed us to work around lobster buoys and to quantify all benthic megafauna to species. Stratifying by substrate, depth, and coastal water masses, we found that abiotic variables explained a significant portion of variance (37- 59%) in benthic species composition, diversity, biomass and economic value. Generally, the density, diversity, and biomass of assemblages significantly increased with the substrate complexity (i.e., from sand-mud to ledge). The diversity, biomass and economic value of assemblages also decreased significantly with increasing depth. Lastly demersal fish densities, sessile invertebrate densities, species diversity, and assemblage biomass increased from east to west, while the abundance of mobile invertebrates and economic value decreased, corresponding mainly to the contrasting water-mass characteristics of the Maine Coastal Current system (i.e., summertime current direction, speed, and temperature). Integrating modeled predictions with existing GIS layers for abiotic conditions allowed us to scale up important assemblage attributes to define key foundational ecological principles of MSP and to find priority regions where some bottom-disturbing activities would have minimal impact to benthic assemblages. We conclude that abiotic proxies can be strong forcing functions for the assembly of marine communities and therefore useful tools for spatial extrapolations of marine assemblages in congested (heavily used) near-shore habitats.
Predictive mapping of tree species assemblages in an African montane rainforest
<p>Conservation of mountain ecosystems can benefit from knowledge of habitats and their distribution patterns. This benefit is particularly true for diverse ecosystems with high conservation values such as the "Afromontane" rainforests. We mapped the vegetation of one such forest: the rugged Bwindi Impenetrable Forest, Uganda—a World Heritage Site known for its many restricted-range plants and animal taxa including several iconic species. Given variation in elevation, terrain and human impacts across Bwindi, we hypothesised that these factors influence the composition and distribution of tree species. To test this, detailed surveys were carried out using stratified random sampling. We established 289 georeferenced sample sites (each with 15 trees ≥20 cm dbh) ranging from 1,320 to 2,467 m a.s.l. and measured 4,335 trees comprising 89 species that occurred in four or more sample sites. These data were analysed against twenty-one digitally mapped biophysical variables using various analytical techniques including non-metric multidimensional scaling (NMDS) and random forests. We identified six tree species assemblages with distinct compositions. Among the biophysical variables, elevation had the strongest correlation with the ordination (r<sup>2</sup>=0.5; <em>p</em><0.001). The "out-of-bag" (OOB) estimate of the error rate for the best final model was 50.7% meaning that nearly half of the variation was accounted for using a limited set of variables. We demonstrate that it is possible to predict the spatial pattern of such a forest based on sampling across a highly complex landscape. Such methods offer accurate mapping of composition that can guide conservation.</p>
Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks
<p>Benchmark data sets of CDPred as described in</p> <p><strong>Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks</strong></p> <p>Zhiye Guo<sup>1</sup>, Jian Liu<sup>1</sup>, Jeffrey Skolnick<sup>2</sup>, Jianlin Cheng<sup>1*</sup></p> <p><sup>1 </sup>Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211</p> <p><sup>2 </sup>School of Biological Sciences, Georgia Institute of Technology, Atlanta, GA 30332-2000</p> <p>*Corresponding author (chengji@missouri.edu)</p> <p>There is four test dataset in this package, each test dataset contains four different folders and one list file. The <strong>afpred_pdb</strong> includes all the corresponding monomer structures predicted by alphafold. The <strong>cdpred_output </strong>includes the prediction results of our tool CDPred for each dataset. The <strong>pre_gen_a3m </strong>includes the multiple sequence alignments file used by CDPred to generate prediction results. And the <strong>true_pdb </strong>includes the fasta file for the test dataset and its heavy atom distance map (h_dist) and carbon alpha distance map (real_dist) that extract from the native structure.</p> <p>HomoTest1: The homodimer test dataset contains 28 targets collect from CASP_CAPRI 10-13</p> <p>HomoTest2: The homodimer test dataset contains 23 targets collect from CASP_CAPRI 13-14</p> <p>HeteroTest1: The heterodimer test dataset contains 9 targets collect from CASP_CAPRI13-14</p> <p>HeteroTest2: The heterodimer test dataset contains 55 targets collect from PDB bank 09-2021 to 11-2021</p>
T2* and quantitative susceptibility mapping in an equine model of post-traumatic osteoarthritis: prediction of mechanical and structural properties
<p>Dataset for the manuscript titled "T2* and quantitative susceptibility mapping in an equine model of post-traumatic osteoarthritis: assessment of mechanical and structural properties"</p>
Mapping the relative accuracy of cross-ancestry prediction
<p>GWAS for six traits for SNPs from the UK Biobank arrays set, six traits for SNPs from the HapMap set, and three traits for SNPs from the ARIC study set.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.