Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
487
datasets available to search
ShareScore release 0.7.1
Dataset results
487 results for “species distribution modeling”
Data from: Modelling habitat distributions for multiple species using phylogenetics
In this paper, we describe an empirical approach to model community structure using phylogenetic signals. That approach combines information about the species (i.e. traits and phylogeny) with information about the habitat (i.e. environmental conditions and spatial distribution of sampling sites) and their interactions to predict the species responses (e.g. the local densities). As an application, we use the approach to model fish densities in rivers. In the model, the different species and size classes were described using a functional trait, body length, and phylogenetic eigenvectors maps whereas the sites were described using water velocity, depth, substrate composition, macrophyte cover, degree-days, total phosphorus, and spatial eigenvector maps. The model (estimated using a regularised Poisson-family Generalised Linear Modelling approach) fitted the data well (likelihood-based R2adj=0.512) and showed fair predictive power (likelihood-based cross-validation R2=0.283) to predict the density of fish pertaining to 48 species totalling 143 combinations of species and size classes in 15 unregulated Canadian rivers. Using the model as a baseline to estimate the effect of flow regulation on community composition, we found that, with few exceptions, the densities of most fish species were lower in regulated than in unregulated rivers. Phylogenetics have been proposed to study community structure, but this is, to our knowledge, the first time phylogenetic information is used explicitly for numerical habitat modelling. We expect that models of that type will be in increasing demand now that development projects are routinely assessed through impact studies.
Data from: Model-based inference for estimating shifts in species distribution, area occupied and centre of gravity
Changing climate is already impacting the spatial distribution of many taxa, including bees, plants, birds, butterflies and fishes. A common goal is to detect range shifts in response to climate change, including changes in the centre of the population's distribution (the centre of gravity, COG), population boundaries and area occupied. Conventional estimators, such as the abundance-weighted average (AWA) estimator for COG, confound range shifts with changes in the spatial distribution of available survey data and may be biased when the distribution of survey data shifts over time. AWA also does not estimate the standard error of COG in individual years and cannot incorporate data from multiple survey designs. To explicitly account for changes in the spatial distribution of survey effort, we propose an alternative species distribution function (SDF) estimator. The SDF approach involves calculating distribution metrics, including COG, population boundary and area occupied, directly from the predicted species distribution or density function. We illustrate the SDF approach using a spatiotemporal model that is available as an r package. Using simulated data, we confirm that the SDF substantially decreases bias in COG estimates relative to the AWA estimator. We then illustrate the method by analysing data from two data sets spanning 1977–2013 for 18 marine fishes along the U.S. West Coast. In our case study, the SDF estimator shows significant northward shifts for six of 18 species (with southward shifts for only 2), where two species (darkblotched and greenstriped rockfishes) have both a northward shift and a decreased area occupied. Pelagic species (e.g. Pacific hake and spiny dogfish) have more variable distribution than bottom-associated species. We also find substantial differences between AWA and SDF estimates of COG that are likely caused by shifts in sampling distribution (which affect the AWA but not the SDF estimator). We caution that common estimators for range shift can yield inappropriate inference whenever sampling designs have shifted over time. We conclude by suggesting further improvements in model-based approaches to analysing climate impacts, including methods addressing the impact of local and regional temperature changes on species distribution.
Quantifying niche similarity among new world seed plants--Species Distribution Models (SDMs) & associated metadata
<p>Niche shift and conservatism are often framed as mutually exclusive. However, both processes could contribute to biodiversity patterns. We tested this expectation by quantifying the degree of climatic niche similarity among New World seed plants.</p> <p>To incorporate the biological reality that species experience varied abiotic conditions across their range, we assembled distribution models and used these to characterize temperature, precipitation, and elevation niches for species as continuously-valued distributions. We then quantified niche similarity (distributional overlap) and identified statistically significant differences compared to a randomized null.</p> <p>The degree of niche similarity differed among climate variables, plant lineages, and at different phylogenetic scales. For example, ~17% of all seed plants were significantly different in elevational niche from their closest relative(s), whereas for precipitation, this value was only ~4%. Average niche similarity decreased with increasing phylogenetic distance, consistent with niche conservatism; however, variance in niche similarity among close relatives was large, such that there always existed niche differences equaling those among distantly related species.</p> <p>Our results suggest researchers should incorporate both niche shift and conservatism as important, scale-dependent factors shaping biodiversity patterns as these processes are not mutually exclusive, nor do they contribute equally to patterns among different plant lineages or niche variables.</p>
Do the predicted suitability scores from species distribution models correlate with species performance on-ground?
<p>Species distribution models are a very popular statistical tool for inferring potential distribution range of species across space and time and are thought to be a good predictor for habitat suitability. Some studies have suggested that if these models are reliable, predicted habitat suitability (PHS) should relate to species traits visualization, growth potential, body size, abundance. We validated this hypothesis by estimating association between the PHS and species abundance for 17 avian species endemic to the Western Ghats - Sri Lanka biodiversity hotspot. Additionally, we compared the PHS of sites where species were detected in both seasons (wet and dry) against sites where they were detected in the dry season alone. As a proxy for abundance, we estimated single-season occupancy estimates (ψ) using detection/non-detection data from multiple visits to the survey sites. We report significant and positive PHS-ψ correlation, though the strength of this association varied across species and models. Half of the species showed higher suitability scores for the sites where they were detected year round. The results presented here suggest that the predictive models can be used as a proxy for habitat quality, in addition to inferring the potential distribution.</p>
Figure 3 in Potential geographic distribution niche modeling based on bioclimatic variables of three species of Temnomastax Rehn and Rehn, 1942 (Orthoptera: Eumastacidae)
Figure 3. Potential geographic distribution predicted by DOMAIN model to Temnomastax ricardoi Descamps, 1973 (blue), and Temnomastax tigris (Burr, 1899) (green). Darkest regions represent higher probabilities of occurrence than clearest regions. (■) Temnomastax ricardoi Descamps, 1973 records; (▲) Temnomastax tigris (Burr, 1899) records.
Figure 2 in Potential geographic distribution niche modeling based on bioclimatic variables of three species of Temnomastax Rehn and Rehn, 1942 (Orthoptera: Eumastacidae)
Figure 2. Potential geographic distribution predicted by DOMAIN model to Temnomastax hamus Rehn and Rehn, 1942 (green). Darkest regions represent higher probabilities of occurrence than clearest regions. On the left is marked the Andes in red, orange and yellow. (●) species records.
Figure 1 in Potential geographic distribution niche modeling based on bioclimatic variables of three species of Temnomastax Rehn and Rehn, 1942 (Orthoptera: Eumastacidae)
Figure 1. Male specimens of some studied species. (a) Temnomastax hamus Rehn and Rehn, 1942 from Minas Gerais, Brazil; (b) Temnomastax ricardoi Descamps, 1973 and (c) Temnomastax tigris (Burr, 1899) from Mato Grosso do Sul, Brazil (photos used with permission of the authors: Marcos Cesar Campis (a) and Paulo Robson de Souza (c).
Resource selection functions based on hierarchical generalized additive models provide new insights into individual animal variation and species distributions
<p>Habitat selection studies are designed to generate predictions of species distributions or inference regarding general habitat associations and individual variation in habitat use. Such studies frequently involve either individually indexed locations gathered across limited spatial extents and analyzed using resource selection functions (RSF), or spatially extensive locational data without individual resolution typically analyzed using species distribution models. Both analytical methodologies have certain desirable features, but analyses that combine individual- and population-level inference with flexible non-linear functions may provide improved predictions while accounting for individual variation. Here, we describe how RSFs can be fit using hierarchical generalized additive models (HGAMs) using widely available software, providing a means to explore individual variation in habitat associations and to generate species distribution maps. We used GPS tracking data from Golden Eagles (Aquila chrysaetos) from across eastern North America with four environmental predictors to generate monthly distribution models. We considered three model structures that assumed different amounts of individual variation in the functional relationship between predictors and habitat use and used k-fold cross-validation to compare model performance. Models accounting for individual variability in shape and smoothness of functional responses performed best. Eagles exhibited the least amount of individual variation in response to land cover variables during winter months, with most individuals more closely adhering to the population-level trend. During summer months, eagles exhibited more substantial individual variation in shape and smoothness of the functional relationships, suggesting some need to account for individual variation in eagle habitat use for both inferential and predictive purposes, during this time of year. Because they allow users to blend flexible functions with random effects structures and are well-supported by a variety of software platforms, we believe that HGAMs provide a useful addition to the suite of analyses used for modeling habitat associations or predicting species distributions.</p>
FIGURE 6 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 6. Beta diversity (β) of Phanaeus within each dominion, segmented by its components (β + β ).
FIGURE 2 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 2. Mean environmental conditions (points) and standard deviation (lines) within each Phanaeus species distribution model sorted by mean altitudinal predicted occurrence.
FIGURE 5 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 5. Occurrence of Phanaeus species in the resulting regionalization, predicted richness and co-occurrence in each dominion. The circle size is the percentage of the predicted species' distribution in each dominion.
FIGURE 7 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 7. Pairwise comparison of Beta diversity of Phanaeus between dominions. The upper panel shows the relative size of β segmented by its components (β + β ). The lower panel shows the value of β . total repl rich total
FIGURE 3 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 3. Potential richness of Phanaeus species obtained by stacking each species Maxent's distribution model, (a) at 30 arc second or by a spatial query (b) at 1° hexagonal cells. This hexagonal grid was used for the regionalization and beta diversity analyses.
FIGURE 4 in Distribution, Regionalization, and Diversity of the dung beetle genus Phanaeus MacLeay (Coleoptera: Scarabaeidae) using Species Distribution Models
FIGURE 4. Regionalization of Phanaeus distribution: Mexican Transition Zone (North American, Mexican, and Mesoamerican dominions) and Neotropical region (Mesoamerican, Pacific, Brazilian and Chacoan dominions).This was obtained from a UPGMA cluster analysis to the result, to produce a dendrogram of the relationship between cells (a) that produced a regionalization (b).
Data from: Microclimate-based species distribution models in complex terrain indicate widespread cryptic refugia under climate change
<p class="MsoNoSpacing"><i>Aim: </i>Species' climatic niches may be poorly predicted by regional climate estimates used in species distribution models (SDMs) due to microclimatic buffering of local conditions. Here, we compare SDMs generated using a locally validated below-canopy microclimate model to those based on interpolated weather station data at two spatial scales to determine the effects of scale, topography, and forest cover on potential future ground-level warming and species distributions.</p> <p class="MsoNoSpacing"><i>Location:</i> Great Smoky Mountains National Park (2090 km<sup>2</sup>; NC, TN, USA)</p> <p class="MsoNoSpacing"><i>Time period: </i>1970 – 2006</p> <p class="MsoNoSpacing"><i>Major taxa:</i> Vascular plant species of the Southern Appalachians</p> <p class="MsoNoSpacing"><i>Methods:</i> We compared the fit and predictions of SDMs generated using a database of plant occurrences and three climate models: macroclimate (1 km, WorldClim), fine-scale (30 m) interpolation of macroclimate with elevation, and fine-scale below-canopy microclimate from a ground-level sensor network.</p> <p class="MsoNoSpacing"><i>Results: </i>We found that, although SDM fit was similar across models, microclimate-derived SDMs predicted substantially greater species persistence with 4 °C of regional warming, with a difference of 50% of the species pool in some areas. Microclimate SDMs predicted higher stability of mid-elevation species, particularly in thermally buffered areas near streams, and critically, less change in species composition at high elevation. In contrast, predictions of macroclimate and interpolation models were similar despite improved resolution.</p> <p class="MsoNoSpacing"><i>Main conclusions:</i> Our results demonstrate that careful selection of climate drivers, including local near-ground validation rather than interpolation, is critical for projecting distributions. They also suggest that some species at risk from climate change might persist, even with 4 °C of macroclimate warming, in cryptic refugia buffered by microclimate, pointing to the roles of forest cover and topography in explaining slower-than-expected changes in understory communities. However, certain species, such as those currently occurring on low-elevation ridges that are sensitive to atmospheric changes, may be at more risk than macroclimate or interpolated SDMs suggest.</p> <p class="MsoNoSpacing"> </p>
FIGURE 1 in Species distribution modelling of Hylarana Species (Anura, Ranidae) and the problem of accurate species identification
FIGURE 1. Distribution maps for Hylarana species. (A) H. erythraea. (B) H. taipehensis. (C) H. tytleri. (D) H. macrodactyla. Each dot represents an occurrence, schematized as following: black round dot, A-quality data; purple square, B-quality data; red triangle, C-quality data. The background shows the topography and water bodies (essentially rivers) of the study area.
FIGURE 2 in Species distribution modelling of Hylarana Species (Anura, Ranidae) and the problem of accurate species identification
FIGURE 2. Species distribution modelling of Hylarana species as estimated by Maxent for present-day conditions, using A (A, D, G, J), B (B, E, H, K) and A+B (C, F, I, L) quality data. (A–C) H. erythraea. (D–E) H. taipehensis. (G–I) H. tytleri. (J–L) H. macrodactyla. Black round dot, A (A, D, G, J), B (B, E, H, K) and A+B (C, F, I, L) quality data.
Data from: Integrated species distribution models fitted in INLA are sensitive to mesh parameterisation
<p class="MsoNormal">The ever-growing popularity of citizen science, as well as recent technological and digital developments, have allowed the collection of data on species' distributions at an extraordinary rate. In order to take advantage of these data, information of varying quantity and quality needs to be integrated. Point process models have been proposed as an elegant way to achieve this for estimates of species distributions. These models can be fitted efficiently using Bayesian methods based on integrated nested Laplace approximations (INLA) with stochastic partial differential equations (SPDE). This approach uses an efficient way to model spatial autocorrelation using a Gaussian random field and a triangular mesh over the spatial domain. The mesh is constructed by user-defined variables, so effectively represents a free parameter in the model. However, there is a lack of understanding about how to set these mesh parameters, and their effect on model performance. Here, we assess how mesh parameters affect predictions and model fit to estimate the distribution of the serotine bat, <em><span>Eptesicus serotinus</span></em><span>,<em> </em>in Great Britain. A Bayesian INLA model was fitted using five meshes of varying densities to a dataset comprising both structured observations from a national monitoring programme and opportunistic records. We demonstrate that mesh density impacted spatial predictions with a general loss of accuracy with increasing mesh coarseness</span>. However, we also show that the finest mesh was unable to overcome spatial biases in the data. In addition, the magnitude of the covariate effects differed markedly between meshes. This confirms that mesh parameterisation is an important and delicate process with implications for model inference. We discuss how species distribution modellers might adapt their use of INLA in light of these findings.</p>
Figure 2 in A two-species distribution model for parapatric newts, with inferences on their history of spatial replacement
Figure 2. Two-species distribution model derived from Triturus cristatus and Triturus marmoratus records over France along with a suite of environmental variable (for details, see main text), extrapolated over neighbouring areas. The colours show the probability for any eligible locality to be occupied by T. cristatus (P c), from deep red for T. cristatus to deep blue for T. marmoratus. Intermediate colours, such as orange and green, represent intermediate probabilities (see the colour scale, which ranges from P c at zero to P c at unity). Areas in black have an elevation of> 1500 m a.s.l. A, model with forestation as documented. B, C, the mutual species distribution under the assumption that western Europe would be completely forested (full forest; B) and devoid of forestation (zero forest; C). The white line approximates the mutual species border as modelled in A. Note that large areas in the south-east of France are devoid of Triturus newts (cf. Fig. 1) and that Italy has another crested newt species (Triturus carnifex), but that a parapatric contact zone is being modelled nevertheless.
Figure 1. A in A two-species distribution model for parapatric newts, with inferences on their history of spatial replacement
Figure 1. A, the outer range borders of the crested newt, Triturus cristatus (c; southern border shown by continuous line) and the marbled newt, Triturus marmoratus (m; northern and eastern border shown by dashed line) in continental France, after Castanet & Guyetant (1989) and Lescure & De Massary (2012). Departments mentioned in the text are as follows: DS, Deux-Sevres; M, Mayenne; V, Vienne. The Lower Rhône T. cristatus population is indicated by LR. The base map was downloaded from MapsLand (https://www.mapsland.com), under a Creative Commons Attribution-ShareAlike 3.0 Licence. B, the area of T. cristatus–T. marmoratus range overlap in Mercator projection, with the generalized species border as inferred from a two-species distribution model (see Fig. 2). The open circles represent documented species occurrences that strongly contradict the model, for T. cristatus (probability of occurrence, Pc ≤ 0.2, in red) and T. marmoratus (Pc ≥ 0.8, in blue). Large symbols represent multiple observations at close range. The drawings of animals, with T. cristatus at the top and T. marmoratus at the bottom, are by Bas Blankevoort, Naturalis Biodiversity Center.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.