Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,042
datasets available to search
ShareScore release 0.9.0
Dataset results
1,042 results for “model species”
Short- And Long-Term Micro & Nano Particles Exposure to Estuarine Model Species at Variable Salinities
Micro (< 5mm) & Nano (1-1000 nm) plastic (MNP) particles are ubiquitous in the environment and have been shown to have a variety of effects on aquatic organisms. The effects of MNP exposure can vary depending on the type of MNP, the concentration of MNP, the duration of exposure, and the salinity of the water. This study used 5 plastic types including polyester (PE), polypropylene (PP), polylactic acid (PLA), polyethylene terephthalate (PETE) and tire particles (TP) in two forms, solid plastic particles and microfibers. To assess potential impacts on exposed organisms, early life stages of the estuarine indicator species Inland Silverside (Menidia beryllina) and mysid shrimp (Americamysis bahia) were exposed to three concentrations at micro- and nano-size fractions, and separately to leachate, across a 5-25 PSU salinity gradient. This exposure study was performed in longer term (21 days for Inland Silverside and 28 days for mysid shrimp) and shorter term (4 days for Inland Silverside and 7 days for mysid shrimp). Following MNP exposures of 7d (A. bahia) and 96 h (M. beryllina), behavioral assays were performed post-exposure from each treatment using a Danio Vision Observation Chamber (Noldus, Wageningen, the Netherlands) for the dark: light cycle as described previously (Siddiqui et al., 2022; Siddiqui et al., 2023; Mundy et al., 2021; Segarra et al., 2021). These behavior studies provide important information for risk assessments and policy making that can establish knowledge for MNPs risk.
Figure 2. D in Neotypification of Drawida hattamimizu Hatai, 1930 (Annelida, Oligochaeta, Megadrili, Moniligastridae) as a model linking mtDNA (COI) sequences to an earthworm type, with a response to the 'Can of Worms' theory of cryptic species
Figure 2. D. hattamimizu unscaled habitus (from Watanabe, 2005, fig. 1 after Hatai's 1931 original).
3D MTHFR models in different eukaryotic species
<p>3D models for <em>Mus musculus, Gallus gallus, Danio rerio, Acanthaster planci, Arabidopsis thaliana, C.elegans MTHFR </em> protein </p>
Generalized model-based solutions to false positive error in species detection/non-detection data: DataS5.
<p>Data/code associated with empirical case study (Gray fox relative abundance estimation/prediction) in article "Generalized model-based solutions to false positive error in species detection/non-detection data" [doi pending].</p>
Data and code for the manuscript: "Varying richness need not imply non-random species co-occurrence: implications for specifying null models"
<p>Data and R code for the manuscript "Varying richness need not imply non-random species co-occurrence: implications for specifying null models".</p>
Bioclimatic data for species distribution modelling in the Amazon Basin
<p>In this dataset, bioclimatic data regarding the Amazon Basin, in the near of the cities of Manaus and Manacapuru are available. There are 11 environmental data variables, referring to temperature, atmospheric pressure, concentration of pollutants and aerosols, such as carbon monoxide, ozone, carbon dioxide, among others. These were collected by the G-159 Gulfstream aircraft during its two periods of operation (IOP1 and IOP2), available in the GOAmazon (Green Ocean Amazon) project's data repository. A spatial interpolation methodology (linear barycentric interpolation) was applied to each variable, in order to obtain a larger area of data. The species occurrence data were collected from the repositories of the ICMBio (Instituto Chico Mendes de Conservação da Biodiversidade) Portal da Biodiversidade and GBIF (Global Biodiversity Information Facility), referring to the same date and location of the environmental data. <br> </p>
Data from: Species distribution models of the Spotted Wing Drosophila (Drosophila suzukii, Diptera: Drosophilidae) in its native and invasive range reveal an ecological niche shift
<p>The Spotted Wing Drosophila (<em>Drosophila</em> <em>suzukii</em>) is native to Southeast Asia. Since its first detection in 2008 in Europe and North America, it has been a pest to the fruit production industry as it feeds and oviposits on ripening fruit. Here we aim to model the potential geographical distribution of <em>D. suzukii</em>. We performed an extensive literature review to map the current records. In total, 517 documented occurrences (96 native and 421 invasive) were identified spanning 52 countries. Next, we constructed three species distribution models (SDMs) based on occurrence records in: 1) the native range (SDMnative), 2) the invasive range in Europe (SDMEurope) and 3) a global model of all records (SDMglobal). The models aimed to investigate, whether this species will be able to occupy additional ecological niches beyond its native range and expand its current geographic distribution both globally and in Europe. The SDMs were generated using Maximum Entropy algorithms (Maxent) based on present occurrence records and bioclimatic variables (WorldClim). Predictions of habitat suitability vary greatly depending on the origins of occurrence records. According to all models, precipitation and low temperatures were key limiting factors for the distribution of <em>D. suzukii</em>, which suggests that this species requires a humid environment with mild winters in order to establish a permanent population in its invasive range. Several regions in the invasive range, not presently occupied by this species, were predicted highly suitable, especially in northern Europe, suggesting that <em>D. suzukii</em> is not occupying its full fundamental niche yet. Synthesis and applications. Based on these models of potential geographic distribution of the Spotted Wing Drosophila (<em>Drosophila</em> <em>suzukii</em>), we show a shift in the ecological niche in <em>D. suzukii</em> populations, emphasizing the importance of using presence and local environmental data. Further investigation regarding new occurrences is recommended to secure optimal pest management. Despite a continuing expansion, many countries still lack proper surveillance schemes, and we urge policymakers to initiate appropriate management programs.</p>
Data from: Complementary strengths of spatially-explicit and multi-species distribution models
<p><span><span><span><span><span><span><span><span><span><span><span> Species distribution models (SDMs) project the outcome of community assembly processes - dispersal, the abiotic environment, and biotic interactions - onto geographic space. Recent advances in SDMs account for these processes by simultaneously modeling the species that comprise a community in a multivariate statistical framework or by incorporating residual spatial autocorrelation in SDMs. However, the effects of combining both multivariate and spatially-explicit model structures on the ecological inferences and the predictive abilities of a model are largely unknown. We used data on eastern hemlock (<i>Tsuga canadensis</i>L.) and five additional co-occurring overstory tree species in 35,569 forest stands across Michigan, USA to evaluate how the choice of model structure, including spatial and non-spatial forms of univariate and multivariate models, affects ecological inference about the processes that shape community composition as well as model predictive ability.</span></span></span></span></span></span></span></span></span></span></span></p> <p><span><span><span><span><span><span><span><span><span><span> Incorporating residual spatial autocorrelation via spatial random effects did not improve out-of-sample prediction for the six tree species, although in-sample model fit was higher in the spatial models. Spatial models attributed less variation in occurrence probability to environmental covariates than the non-spatial models for all six tree species, and estimated higher (more positive) residual co-occurrence values for most species pairs. The non-spatial multivariate model was better suited for evaluating habitat suitability and hypotheses about the processes that shape community composition. Environmental correlations and residual correlations among species pairs were positively related, perhaps indicating that residual correlations were due to shared responses to unmeasured environmental covariates. This work highlights the importance of choosing a non-spatial model formulation to address research questions about the species-environment relationship or residual co-occurrence patterns, and a spatial model formulation when within-sample prediction accuracy is the main goal.</span></span></span></span></span></span></span></span></span></span></p>
Supporting data for "Mammalian species abundance across a gradient of tropical land-use intensity: A hierarchical multi-species modelling approach"
<p>Combined camera trap and live trap dataset underlying the analyses in a Biological Conservation paper (https://doi.org/10.1016/j.biocon.2017.05.007), provided in .csv format. This spatially- and temporally-replicated dataset is suitable for occupancy modelling.</p> <p>The first 3 columns in the dataset are:</p> <p>1) Trap location name – old-growth forest, logged forest and oil palm plantation locations have the prefixes "Old", "Log" and "Palm", respectively</p> <p>2) Sampling occasion number – camera trap and live trap occasions have the prefixes “Lvtrap” and “Ctrap”, respectively, and are defined in the paper</p> <p>3) Calendar year in which sampling took place (most locations were sampled in > 1 calendar years)</p> <p>Following these 3 columns, there are 66 columns for each of the mammal species detected during the study (species common names are used). The values for each species represent the number of independent captures, as defined in the paper. This can be reduced to detection/non-detection data (zeroes and ones), if needed, for occupancy modelling.</p>
Code and data for Bayesian joint species distribution model selection for community-level prediction
<p>Code and data for reproducing the analysis in the manuscript "Bayesian joint species distribution model selection for community-level prediction." Provided data include percent cover observations for 39 modeled vascular plant species within boreal forest understory communities and environmental model covariates. R code is provided to generate model inputs, apply alternative models, generate out-of-sample predictions, and calculate associated community and species log scores and alternative model evaluation metrics. Further, R source code is provided to implement the multinomial joint species distribution model defined in the manuscript. Details on the data, its processing, and the alternative model definitions and structure can be found in the main text of the manuscript. Provided data are currently being used in ongoing analyses and coordination with authors may be warranted to avoid duplicate publication. Potential users are encouraged to consider collaboration with authors when useful and appropriate. Misinterpretation of data may occur if used outside the context of the original analysis. All data are made available in their current state. While significant efforts have been made to ensure data accuracy, complete accuracy cannot be guaranteed. Data may be updated periodically. It is the responsibility of the data user to check for updated versions of the data.</p>
Data from: Integrated species distribution models to account for sampling biases and improve range wide occurrence predictions
<p><strong><span>Aim</span></strong></p> <p><span>Species distribution models (SDMs) that integrate presence-only and presence-absence data offer a promising avenue to improve information on species' geographic distributions. The use of such 'integrated SDMs' on a species range-wide extent has been constrained by the often-limited presence-absence data and by the heterogeneous sampling of the presence-only data. Here, we evaluate integrated SDMs for studying species ranges with a novel expert range map-based evaluation. We build a new understanding about how integrated SDMs address issues of estimation accuracy and data deficiency and thereby offer advantages over traditional SDMs.</span></p> <p><strong><span>Location</span></strong></p> <p><span>South and Central America.</span></p> <p><strong><span>Time period</span></strong></p> <p><span>1979-2017.</span></p> <p><strong><span>Major taxa studied</span></strong></p> <p><span>Hummingbirds.</span></p> <p><strong><span>Methods</span></strong></p> <p><span>We build integrated SDMs by linking two observation models – one for each data type – to the same underlying spatial process.</span> <span>We validate SDMs with two schemes: i) cross-validation with presence-absence data and ii) comparison with respect to the species' whole range as defined with IUCN range maps. We also compare models relative to the estimated response curves and compute the association between the benefit of the data integration and the number of presence records in each data set.</span></p> <p><strong><span>Results</span></strong></p> <p><span>The integrated SDM accounting for the spatially varying sampling intensity of the presence-only data was one of the top-performing models in both model validation schemes. Presence-only data alleviated overly large niche estimates, and data integration was beneficial compared to modelling solely presence-only data for species that had few presence points when predicting the species' whole range. On the community level, integrated models improved the species richness prediction.</span></p> <p><strong><span>Main conclusions</span></strong></p> <p><span>Integrated SDMs combining presence-only and presence-absence data are successfully able to borrow strengths from both data types and offer improved predictions of species' ranges. Integrated SDMs can potentially alleviate the impacts of taxonomically and geographically uneven sampling and to leverage the detailed sampling information in presence-absence data.</span></p>
Data from: A cost-effective blood DNA methylation-based age estimation method in domestic cats, Tsushima leopard cats (Prionailurus bengalensis euptilurus), and Panthera species, using targeted bisulfite sequencing and machine learning models
<p><span>Knowledge of individual age can help both in-situ and ex-situ conservation programs to design more efficient and suitable management plans for targeted wildlife species. DNA methylation is one of the epigenetic aging markers that has emerged as a promising tool that can estimate age with high accuracy using only a tiny amount of biological material, which can be collected in a minimally invasive way. Here, we sequenced five targeted genetic regions and used </span><span>8–23</span><span> selected CpG sites to build age estimation models with machine learning methods </span><span>with about only $3–7 per sample</span><span>, using blood samples of seven Felidae species—ranging from small to big, and domestic to endangered species: domestic cats (<em>Felis catus</em>, 139 samples), Tsushima leopard cats (<em>Prionailurus bengalensis euptilurus</em>, 84 samples), and five<em> Panthera </em>species (96 samples). </span><span>The models built achieved satisfactory accuracy—the mean absolute error of the best models was 1.966, 1.348, and 1.552 years in domestic cats, Tsushima leopard cats, and <em>Panthera</em> spp., respectively.</span><span> Our models in domestic cats and Tsushima leopard cats were applicable to individuals regardless of health conditions, indicating the high applicability of our models to samples collected from diverse situations, e.g., rescued individuals in the context of conservation. We also showed the possibility of developing universal age estimation models for the five<em> Panthera</em> spp. using two of the five genetic regions, suggesting an even lower cost to use our models for future applications.</span></p>
Resources for: Spatio-temporal integrated Bayesian species distribution models reveal lack of broad relationships between traits and range shifts
<p><strong>Aim</strong>: Climate change and habitat loss or degradation are some of the greatest threats that species face today, often resulting in range shifts. Species traits have been discussed as important predictors of range shifts, with the identification of general trends being of great interest for conservation efforts. However, studies reviewing relationships between traits and range shifts have questioned the existence of such generalized trends, due to mixed results and weak correlations, as well as analytical shortcomings. The aim of this study was to test this relationship empirically, using analytical approaches that account for common sources of bias when assessing range trends.<br><strong>Location</strong>: Tanzania, East Africa.<br><strong>Time period</strong>: 1980-1999 and 2000-2020.<br><strong>Major taxa studied</strong>: 57 savannah specialist birds found in Tanzania, belonging to 26 families and 11 orders.<br><strong>Methods</strong>: We applied recently developed integrated spatio-temporal species distribution models in R-INLA, combining citizen science and bird atlas data to estimate ranges of species, quantify range shifts, and test the predictive power of traditional trait groups, as well as exposure-related and sensitivity traits. We based our study on 40 years of bird observations in East African savannahs, a biome that has experienced increasing climatic and non-climatic pressures over recent decades. We correlated patterns of change with species traits.<br><strong>Results</strong>: We find indications of relationships identified by previous research, but low average explanatory power of traits from an ecological perspective, confirming the lack of meaningful general associations. However, our analysis finds compelling species-specific results.<br><strong>Main conclusions</strong>: We highlight the importance of individual assessments, while demonstrating the usefulness of our analytical approach for analyses of range shifts.</p>
Fig. 8. Heat map resulting from the Species Distribution Model using MaxEnt, where 1 in A revision of the genus Armillipora Quate (Diptera: Psychodidae) with the descriptions of two new species
Fig. 8. Heat map resulting from the Species Distribution Model using MaxEnt, where 1 is equal to the highest probability of distribution, while 0 is the lowest probability.
Improving distribution models of sparsely-documented disease vectors by incorporating information on related species via joint modeling
<p>A necessary component of understanding vector-borne disease risk is the accurate characterization of the distributions of their vectors. Species distribution models have been successfully applied to data-rich species but may produce inaccurate results for sparsely-documented vectors. In light of global change, vectors that are currently not well-documented could become increasingly important, requiring tools to predict their distributions. One way to achieve this could be to leverage data on related species to inform the distribution of a<strong> </strong>sparsely-documented vector based on the assumption that the environmental niches of related species are not independent. Relatedly, there is a natural dependence of the spatial distribution of a disease on the spatial dependence of its vector. Here, we propose to exploit these correlations by fitting a hierarchical model jointly to data on multiple vector species and their associated human diseases to improve distribution models of sparsely-documented species. To demonstrate this approach, we evaluated the ability of twelve models—which differed in their pooling of data from multiple vector species and inclusion of disease data—to improve distribution estimates of sparsely-documented vectors. We assessed our models on two simulated data sets, which allowed us to generalize our results and examine their mechanisms. We found that when the focal species is sparsely documented, incorporating data on related vector species reduces uncertainty and improves accuracy by reducing overfitting. When data on vector species are already incorporated, disease data only marginally improve model performance. However, when data on other vectors are not available, disease data can improve model accuracy and reduce overfitting and uncertainty. We then assessed the approach on empirical data on ticks and tick-borne diseases in Florida and found that incorporating data on other vector species improved model performance. This study illustrates the value of exploiting correlated data via joint modeling to improve distribution models of data-limited species.</p>
Dataset for Integrated Species Distribution Model for pikeperch larvae in the Porvoo-Sipoo archipelago
<p>This record contains the data required to run the code for fitting the Integrated Species Distribution Model described in <a href="https://arxiv.org/abs/2206.08817">arXiv:2206.08817 [stat.ME].</a></p> <h1>Files in this record</h1> <ul> <li><strong>transect_data.csv</strong> Line transect observations from Porvoo-Sipoo archipelago, Finland on June 2017.</li> <li><strong>expert_assessments.tif</strong> Rasterized, anonymous expert assessments. Categorical values denoting how likely a given location is to be a spawning location for pikeperch. 4 categories, with smaller values corresponding to higher probabilities.</li> <li><strong>covariate_raster_example.tif</strong> Rasterized example environmental covariate values. These are similarly structured as the covariate data used in the study and compatible with the analysis code. However, since we do not have the permission to release the original data set, these values are instead generated based on the projected planar coordinates such that they have roughly similar spatial gradients as the original covariates.</li> </ul> <h1>Detailed descriptions</h1> <h2>Transect data</h2> <h3>Location and replicate identifiers</h3> <ul> <li> <p><strong>id</strong> : transect identifier. Replicates of the same transect have the same identifier.</p> </li> <li> <p><strong>id2</strong> : alternate transect identifier, unique for each transect.</p> </li> <li> <p><strong>repeated</strong> : whether transect was replicated or not.</p> </li> <li> <p><strong>X_euref</strong> : easting coordinate, EUREF_FIN_TM35FIN, for the transect starting location in [meters]</p> </li> <li> <p><strong>Y_euref</strong> : northing coordinate, EUREF_FIN_TM35FIN, for the transect starting location in [meters]</p> </li> <li><strong>date</strong> : date of the measurement, DD/MM/YYYY</li> <li><strong>week</strong> : week number of the measurement date</li> </ul> <h3>In situ measurements</h3> <ul> <li> <p><strong>volume</strong> : Transect water volume [m^3]. Transect length (500m) multiplied by sampler surface area. Used as survey effort.</p> </li> <li> <p><strong>heading</strong> : compass heading (direction) for the transect, in [degrees].</p> </li> <li> <p><strong>SumKUHA</strong> : total pikeperch (<em>Sander lucioperca</em>, kuha in Finnish) larvae count in each transect [scalar]</p> </li> </ul> <h2>Expert assessments</h2> <p>The raster contains assessments from 10 local experts encoded as separate raster layers (Expert_1, Expert_2, ..., Expert_10). Raster resolution is 50m x 50m and the planar coordinates are based on the same coordinate reference system as the transect observations (UTM zone 35).</p> <p>The assessments are coded as integers with values between 1 and 4, with smaller values corresponding to higher probabilities.</p> <h2>Covariate raster example</h2> <p>This raster has the same spatial dimensions and uses the same coordinate reference system as the expert assessment raster and has three layers, one for each covariate. The covariate values are generated based on the spatial coordinates such that each covariate has similar spatial gradient as the original covariate. The covarites have the same names as in the original covariate data (<strong>dptLUKE</strong>, <strong>dist10m</strong> and <strong>lined3km</strong>).</p> <h1>Creators</h1> <p>Transect data collected and curated by Sanna Kuningas.</p> <p>Original covariate rasters curated by Sanna Kuningas from data sets collected by the Finnish Environment Institute and the Natural Resources Institute Finland.</p> <p>Expert assessments originally digitized and rasterized by Jussi Mäkinen. Additional refinement to assessment rasters by Karel Kaurila.</p> <p>Preparation for publishing on Zenodo for all of the data sets by Karel Kaurila.</p> <h2>Change log</h2> <ul> <li> 2025 Jan 31: Included columns <strong>date</strong> and <strong>week</strong> for <strong>transect_data.csv</strong>.</li> </ul>
Data from: Integrated SDM database: Enhancing the relevance and utility of species distribution models in conservation management
<p><span>1. Species' ranges are changing at accelerating rates. Species distribution models (SDMs) are powerful tools that help rangers and decision-makers prepare for reintroductions, range shifts, reductions, and/or expansions by predicting habitat suitability across landscapes. Yet, range-expanding or -shifting species in particular face other challenges that traditional SDM procedures cannot quantify, due to large differences between a species' currently-occupied range and potential future range. The realism of SDMs is thus lost and not as useful for conservation management in practice. Here, we address these challenges with an extended assessment of habitat suitability through an <i>integrated SDM database (iSDMdb)</i>.</span></p> <p><span>2. The<i> iSDMdb</i> is a spatial database of predicted sites in a species' prediction range, derived from SDM results, and is a single spatial feature that contains additional, user-friendly data fields that synthesise and summarise SDM predictions and uncertainty, human impacts, restoration features, novel preferences in novel spaces, and management priorities. To illustrate its utility<i>,</i> we used the endangered New Zealand sea lion (<i>Phocarctos hookeri</i>). We consulted with wildlife rangers, decision-makers, and sea lion experts to supplement SDM predictions with additional, more realistic, and applicable information for management. </span></p> <p><span>3. Almost half the data fields included in this database resulted from engaging with these end-users during our study. The SDM found 395 predicted sites. However, the <i>iSDMdb</i>'s additional assessments showed that the actual suitability of most sites (90%) was questionable due to human impacts. >50% of sites contained unnatural barriers (fences, grazing grasslands), and 75% of sites had roads located within the species' range of inland movement. Just 5% of the predicted sites were mostly (>80%) protected.</span></p> <p><span>4. Integrating SDM results with supplemental assessments provides a way to address SDM limitations, especially for range-expanding or -shifting species. SDM products for conservation applications have been critiqued for lacking transparency and interpretation support, and ineffectively communicating uncertainty. The <i>iSDMdb</i> addresses these issues and enhances the practical relevance and utility of SDMs for stakeholders, rangers, and decision-makers. We exemplify how to build an <i>iSDMdb</i> using open-source tools, and how to make diverse, complex assessments more accessible for end-users.</span></p>
Fig. 1 in The impact of land use on species composition and habitat structure in Sudanian savannas - A modelling study in protected areas and agricultural lands of southeastern Burkina Faso
Fig. 1. − Study area including the Pama reserve and neighbouring PAs of the western WAPO complex. The Pama, Tindangou and Madjoari areas are enclaves where agriculture is allowed. The small country map in the lower right shows the position of the study area within Burkina Faso.
Fig. 3 in The impact of land use on species composition and habitat structure in Sudanian savannas - A modelling study in protected areas and agricultural lands of southeastern Burkina Faso
Fig. 3. − Maps of mean maximum plant size (calculated as average of maximum plant size of all species predicted as present within a grid cell). A. Grasses (Poaceae) (30-360 cm); B. Woody species (3-25 m). The color coding stretches from light yellow for the lowest values via orange and red to violet for the highest values.
Fig. 2 in The impact of land use on species composition and habitat structure in Sudanian savannas - A modelling study in protected areas and agricultural lands of southeastern Burkina Faso
Fig. 2. − Maps of species richness. A. All plant species (2-211 spp.); B. Graminoids (0-50 spp.); C. Forbs (0-86 spp.); D. Woody species (0-52 spp.); E. Weedy species (0-48 spp.); F. Non-weedy species (0-140 spp.). The color coding stretches from light yellow for the lowest values via orange and red to violet for the highest values.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.