Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

487

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

487 results for “species distribution models”

Learn how ShareScore rates datasets ↗
dryad36/100

Maxent species distribution modelling of 10 cetacean species in the northeastern Pacific from citizen science occurrence records

Open the record for dataset details and reuse information.

publicAug 2025View details →
dryad36/100

Data from: The challenge of modeling niches and distributions for data-poor species: a comprehensive approach to model complexity

Open the record for dataset details and reuse information.

publicJul 2017View details →
dryad36/100

Testing whether ensemble modelling is advantageous for maximising predictive performance of species distribution models

Open the record for dataset details and reuse information.

publicMar 2020View details →
dryad36/100

Accommodating the role of site memory in dynamic species distribution models

Open the record for dataset details and reuse information.

publicMay 2021View details →
dryad36/100

Derived variables and coordinates to assess the ecological relevance of multiscale bathymetry for coral species distribution modelling across the Great Barrier Reef

Open the record for dataset details and reuse information.

publicApr 2025View details →
dryad36/100

Presettlement tree distributions and forest types of northeast Ohio, USA, mapped with species distribution models

Open the record for dataset details and reuse information.

publicNov 2024View details →
dryad36/100

Challenges and opportunities of species distribution modelling of terrestrial arthropod predators

Open the record for dataset details and reuse information.

publicOct 2021View details →
dryad36/100

Data from: Combining past and contemporary species occurrences with ordinal species distribution modeling to investigate responses to climate change

Open the record for dataset details and reuse information.

publicJan 2025View details →
dryad36/100

Soil chemical variables improve models of understory plant species distributions

Open the record for dataset details and reuse information.

publicMay 2022View details →
zenodo32/100

Provisioning forest and conservation science with European tree species distribution models under climate change

<p>Estimating shifts in the current range of forest tree species is crucial for formulating adaptive management strategies such as assisted migration. Ecological niche models have been the most widely used tools to estimate the potential climatic suitability of species worldwide. The reliability of such estimations depends on the model algorithm and the input data such as climate and species occurrence. We developed a dataset of the potential distribution of seven ecologically and economically important tree species of Europe in terms of their climatic suitability with an ensemble approach while accounting for uncertainty due to model algorithms. The distribution models shall be the basis for follow-up studies in forest and conservation science.</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

FIGURE­5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated. in Redescription and distributional range extension of the Speckled Siphonfish, Siphamia guttulata (Pisces: Apogonidae)

FIGURE­5. Maximum likelihood tree based on the Kimura 2-parameter model of the COI sequences from the Siphamia species with P. kauderni as the outgroup. Tree shown here has the highest log likelihood following 10 000 replications. The percentage of trees in which the associated taxa clustered together is shown next to the branches, branch lengths are measured in the number of substitutions per site and all positions containing gaps and missing data have been eliminated.

opennotspecifiedApr 2020View details →
zenodo32/100

High-resolution future climate data for species distribution models in Europe

<p><strong>Description</strong></p> <p>This dataset contains a set of 13 climatological variables (<code>Variable</code>, <code>VariableName</code>) at a spatial resolution of 1x1km for Europe (nx = 13147, ny = 6071) for historical (<code>ClimatePeriod</code>) and future climate conditions. These variables are a subset of the so-called bioclimatic variables that are often part of global gridded datasets (e.g. <a href="https://worldclim.org/data/bioclim.html">WorldClim</a>, <a href="http://chelsa-climate.org/bioclim/">CHELSA</a>) that have been specifically developed for species distribution modelling and ecological applications.</p> <p>The climatological data correspond to 35-year (<code>Startyear_Endyear</code> = <code>1971_2005</code>) and 30-year (<code>Startyear_Endyear</code> = <code>2041_2070</code>) mean values representing respectively historical and future climate conditions. To account for the future climate conditions, three possible emission scenarios of greenhouse gases as defined by the <a href="https://www.ipcc.ch/">Intergovernmental Panel on Climate Change (IPCC)</a> are used (<code>ClimatePeriod</code> = <code>rcp26</code>, <code>rcp45</code>, <code>rcp85</code>).</p> <p>The complete set of variables (var[1-13]) for which historical and future climate data layers are produced are given below.</p> <p>The source data for the climate layers were assembled from the <a href="https://cordex.org/data-access/">EURO-CORDEX archive</a> (Kotlarski et al., 2014). More specifically, we have used the regional climate model simulations for Europe at a spatial resolution of 12.5x12.5km on which a three-step statistical downscaling approach has been applied:</p> <ol> <li><strong>Processing</strong> (averaging, totals, &hellip;) of all available time series of the EURO-CORDEX model experiments (<code>ClimatePeriod</code> = evaluation, historical, rcp) for the climatological variables.</li> <li><strong>Interpolation</strong> of the data layers from the 12.5x12.5km EURO-CORDEX grid to a 1x1km spatial <a href="http://chelsa-climate.org/">CHELSA</a> (Karger et al., 2017) reference grid (see files <code>lat_1km.csv</code> and <code>lon_1km.csv</code>).</li> <li><strong>Calculate differences</strong> between the 1x1km-interpolated variables (<code>Variable</code> = only for var[1-9]) from the evaluation model experiments (or <code>ClimatePeriod</code>) and the corresponding reference bioclimatic CHELSA variables. In order to account for possible biases present in the EURO-CORDEX climate models, these differences (or biases) are then subtracted from the respective 1x1-km-interpolated variables for the historical and rcp model experiments (<code>ClimatePeriod</code>).</li> </ol> <p>The dimensions of the 1x1km grid (excl. the first row and column):</p> <ul> <li>y-dimension = number of columns = 6071</li> <li>x-dimension = number of rows = 13147</li> </ul> <p>The longitudes and latitudes of respectively the southwest and northeast corner of the grid are:</p> <ul> <li>longitude -44.592; latitude 21.991 (southwest corner)</li> <li>longitude 64.967; latitude 72.583 (northeast corner)</li> </ul> <p>The climatological variables are used as input data for the species distribution modelling of Invasive Alien Species for the <a href="https://osf.io/7dpgr/">Tracking Invasive Alien Species (TrIAS)</a> project.</p> <p><strong>Variables</strong></p> <ul> <li><strong>Variable</strong> (VariableName): Unit</li> <li><strong>var1</strong> (AnnualMeanTemperature): &deg;C</li> <li><strong>var2</strong> (AnnualAmountPrecipitation): mm year<sup>-1</sup></li> <li><strong>var3</strong> (AnnualVariationPrecipitation): coefficient of variation</li> <li><strong>var4</strong> (AnnualVariationTemperature): stdev</li> <li><strong>var5</strong> (MaximumTemperatureWarmestMonth): &deg;C</li> <li><strong>var6</strong> (MinimumTemperatureColdestMonth): &deg;C</li> <li><strong>var7</strong> (TemperatureAnnualRange): &deg;C</li> <li><strong>var8</strong> (PrecipitationWettestMonth): mm</li> <li><strong>var9</strong> (PrecipitationDriestMonth): mm</li> <li><strong>var10</strong> (30yrMeanAnnualCumulatedGDDAbove5degreesC): &deg;C days</li> <li><strong>var11</strong> (AnnualMeanPotentialEvapotranspiration): mm day<sup>-1</sup></li> <li><strong>var12</strong> (AnnualMeanSolarRadiation): W m<sup>-2</sup></li> <li><strong>var13</strong> (AnnualVariationSolarRadiation): stdev</li> </ul> <p><strong>Files</strong></p> <ul> <li><strong>varX_VariableName_ClimatePeriod_Startyear_Endyear.csv</strong>:&nbsp;climatological data layers for the 13 variables listed above</li> <li><strong>lon_1km.csv</strong>: longitudes for the&nbsp;1x1km grid</li> <li><strong>lat_1km.csv</strong>: latitudes for the&nbsp;1x1km grid</li> </ul>

opencc-zeroApr 2020View details →
zenodo32/100

Supplementary material 1 from: Datta A, Schweiger O, Kühn I (2020) Origin of climatic data can determine the transferability of species distribution models. NeoBiota 59: 61-76. https://doi.org/10.3897/neobiota.59.36299

Variable selection using cluster analsys based on Spearman's rank corellation and UPGMA method for agglomeration

opencc-zeroAug 2020View details →
dryad32/100

Data from: Evaluating presence-only species distribution models with discrimination accuracy is uninformative for many applications

Aim: Species distribution models are used across evolution, ecology, conservation, and epidemiology to make critical decisions and study biological phenomena, often in cases where experimental approaches are intractable. Choices regarding optimal models, methods, and data are typically made based on discrimination accuracy: a model's ability to predict subsets of species occurrence data that were withheld during model construction. However, empirical applications of these models often involve making biological inferences based on continuous estimates of relative habitat suitability as a function of environmental predictor variables. We term the reliability of these biological inferences "functional accuracy." We explore the link between discrimination accuracy and functional accuracy. Methods: Using a simulation approach we investigate whether models that make good predictions of species distributions correctly infer the underlying relationship between environmental predictors and the suitability of habitat. Results: We demonstrate that discrimination accuracy is only informative when models are simple and similar in structure to the true niche, or when data partitioning is geographically structured. However, the utility of discrimination accuracy for selecting models with high functional accuracy was low in all cases. Main conclusions: These results suggest that many empirical studies and decisions are based on criteria that are unrelated to models' usefulness for their intended purpose. We argue that empirical modeling studies need to place significantly more emphasis on biological insight into the plausibility of models, and that the current approach of maximizing discrimination accuracy at the expense of other considerations is detrimental to both the empirical and methodological literature in this active field. Finally, we argue that future development of the field must include an increased emphasis on simulation; methodological studies based on ability to predict withheld occurrence data may be largely uninformative about best practices for applications where interpretation of models relies on estimating ecological processes, and will unduly penalize more biologically informative modeling approaches.

opencc-zeroAug 2020View details →
dryad32/100

Integrating univariate niche dynamics in species distribution models: a step forward for marine research on biological invasions

<p>Aim The development of approaches to predict the distribution and potential expansion of invasive species is still an open challenge. Here our goal is to improve the modelling procedure for marine invaders by coupling Species Distribution Models (SDMs) with an analysis of their univariate niche dynamics. In particular, we tested for the first time whether choosing model predictors among the stable niche dimensions was effective in improving predictions of invasive species expansion.<br> Location Mediterranean Sea<br> Taxon Dusky spinefoot, Siganus luridus.<br> Methods We analysed the univariate niche dynamics for S. luridus across its native and invaded ranges, by applying a standardized framework that allowed the identification of cases of niche stability or shift. We compared inter-range transferability of SDMs fitted with different combinations of labile or stable predictors. Finally, we evaluated interactions in SDM settings (calibration area, model technique and predictors set) on models' predictive ability, using independent data from the most recent phase of invasion.<br> Results We detected a pattern of niche stability for several variables, especially salinity and bathymetry, which positively influenced model inter-ranges transferability: when the models calibrated in the native range include only stable niche axes, predictive ability is improved. We also identified a shift toward lower surface temperatures in the introduced range, which were almost never experienced by the species before invasion. The model calibrated within the combined ranges was the most ecologically congruent. Also, models calibrated in the invaded range allowed a correct prediction of range expansion, with the predicted suitable areas only slightly underestimated.<br> Main conclusions We provide the first evidence that using conserved predictors in SDMs improves inter-range projections of expanding invasive species. Variable selection, calibration area and modelling technique all matter when modelling invasive species, with important interaction effects. We provide guidelines on how to improve SDMs applications in biological invasion research.</p>

opencc-zeroOct 2020View details →
dryad32/100

Avian point-counts from Rhode Island and Connecticut used to test species distribution models

<p>Spatial-biases are a common feature of presence-absence data from citizen scientists. Spatial thinning can mitigate errors in species distribution models (SDMs) that use these data. When detections or non-detections are rare, however, SDMs may suffer from class imbalance or low sample size of the minority (i.e. rarer) class. Poor predictions can result, the severity of which may vary by modeling technique. To explore the consequences of spatial bias and class imbalance in presence-absence data, we used eBird citizen science data for 102 bird species from the northeastern USA to compare spatial thinning, class balancing, and majority-only thinning (i.e., retaining all samples of the minority class). We created SDMs using two parametric or semi-parametric techniques (generalized linear models and generalized additive models) and two machine-learning techniques (random forest and boosted regression trees). We tested the predictive abilities of these SDMs using an independent and systematically collected reference dataset with a combination of discrimination (area under the receiver operator characteristic curve; true skill statistic; area under the precision-recall curve) and calibration (Brier score; Cohen's kappa) metrics. We found large variation in SDM performance depending on thinning and balancing decisions. Across all species, there was no single best approach, with the optimal choice of thinning and/or balancing depending on modeling technique, performance metric, and the baseline sample prevalence of species in the data. Spatially thinning all the data was often a poor approach, especially for species with baseline sample prevalence &lt; 0.1. For most of these rare species, balancing classes improved model discrimination between presence and absence classes, but hindered model calibration. Baseline sample prevalence, sample size, modeling approach, and the intended application of SDM output – whether discrimination or calibration – should guide decisions about how to thin or balance data, given the considerable influence of these methodological choices on SDM performance.  For prognostic applications requiring good model calibration (vis-à-vis discrimination), the match between sample prevalence and true species prevalence may be the overriding feature and warrants further investigation.</p>

opencc-zeroOct 2020View details →
dryad32/100

Data from: Modelling species distributions limited by geographic barriers: a case study with African and American primates

<p><strong>Aim:</strong> The boundaries of species distributions are often shaped by natural barriers such as mountains and rivers, but species distribution models usually fail to include these constraints. We tested several approaches that include barriers as explanatory variables in species distribution models.</p> <p><strong>Location:</strong> Africa and South America.</p> <p><strong>Time period:</strong> Current</p> <p><strong>Major taxa studied:</strong> Primates</p> <p><strong>Methods:</strong> We modelled the ranges of pairs of species separated by a river taking into account three explanatory components: the environment (ecosystems, topo-hydrography, climate, human pressure), the spatial structure shaped by history and population dynamics (using a trend-surface approach), and rivers as naturals barriers to dispersal (using a binary cis-trans variable that describes both sides of the river). To assess how the addition of a spatial structure and the barrier could improve distribution models, we used a nested approach by comparing models based on: a) the environment; b) the environment and the spatial structure; and c) the environment, the spatial structure and the river. These models were constructed using the favourability functions.</p> <p><strong>Results:</strong> There was a decreased occurrence of high-favourability values in the opposite side of the rivers in models that included the spatial structure of distributions, compared to models based on environment alone. This decrease was more marked when the description of the spatial structure was made more flexible. However, model performance was significantly improved by the inclusion of cis-trans variables that identified areas on the opposite side as totally unfavourable.</p> <p><strong>Main conclusions:</strong> The performance of distribution models can improve by the use of approaches that describe barriers. Although adding the location of geographic units in relation to a river appears to be the most accurate way to define the presence of a barrier, defining this variable may be challenging. A suitable alternative is to analyse the spatial structure of distributions using a flexible approach.</p>

opencc-zeroNov 2020View details →
zenodo32/100

Combining Satellite Remote Sensing and Climate Data in Species Distribution Models to Improve the Conservation of Iberian White Oaks (Quercus L.)

<p>The Iberian Peninsula hosts a high diversity of oak species, being a hot-spot for the&nbsp; conservation of European White Oaks (Quercus) due to their environmental heterogeneity and its&nbsp;critical role as a phylogeographic refugium. Identifying and ranking the drivers that shape the&nbsp;distribution of White Oaks in Iberia requires that environmental variables operating at distinct&nbsp;scales are considered. These include climate, but also ecosystem functioning attributes (EFAs)&nbsp;related to energy&ndash;matter exchanges that characterize land cover types under various environmental&nbsp;settings, at finer scales. Here, we used satellite-based EFAs and climate variables in species&nbsp;distribution models (SDMs) to assess how variables related to ecosystem functioning improve our&nbsp; understanding of current distributions and the identification of suitable areas for White Oak species&nbsp;in Iberia. We developed consensus ensemble SDMs targeting a set of thirteen oaks, including both&nbsp;narrow endemic and widespread taxa. Models combining EFAs and climate variables obtained a&nbsp;higher performance and predictive ability (true-skill statistic (TSS): 0.88, sensitivity: 99.6, specificity:&nbsp;96.3), in comparison to the climate-only models (TSS: 0.86, sens.: 96.1, spec.: 90.3) and EFA-only&nbsp;models (TSS: 0.73, sens.: 91.2, spec.: 82.1). Overall, narrow endemic species obtained higher&nbsp;predictive performance using combined models (TSS: 0.96, sens.: 99.6, spec.: 96.3) in comparison to&nbsp;widespread oaks (TSS: 0.80, sens.: 92.6, spec.: 87.7). The Iberian White Oaks show a high dependence&nbsp;on precipitation and the inter-quartile range of Normalized Difference Water Index (NDWI) (i.e.,&nbsp;seasonal water availability) which appears to be the most important EFA variable. Spatial&nbsp;projections of climate&ndash;EFA combined models contribute to identify the major diversity hotspots for&nbsp;White Oaks in Iberia, holding higher values of cumulative habitat suitability and species richness.&nbsp;We discuss the implications of these findings for guiding the long-term conservation of IberianWhite Oaks and provide spatially explicit geospatial information about each oak species (or set of&nbsp;species) relevant for developing biogeographic conservation frameworks.</p>

opencc-by-4.0Dec 2020View details →
zenodo32/100

Supplementary material 1 from: Bustamante RO, Alves L, Goncalves E, Duarte M, Herrera I (2020) A classification system for predicting invasiveness using climatic niche traits and global distribution models: application to alien plant species in Chile. NeoBiota 63: 127-146. https://doi.org/10.3897/neobiota.63.50049

Table S1. Exotic species located in Quadrant 1 (see Figure 3) and impacts on biodiversity, agriculture and cattle raisng

opencc-zeroDec 2020View details →
dryad32/100

Combining conservation status and species distribution models for planning assisted colonisation under climate change

<p>Effects of climate change are particularly important in the Mediterranean Biodiversity hotspot where rising temperatures and drought are negatively affecting several plant taxa, including endemic species. Assisted Colonisation (AC) represents a useful tool for reducing the effect of climate change on endemic plant species threatened by climate change.</p> <p>We combined SDMs for 188 taxa endemic to Italy with the IUCN red listing range loss threshold under criterion A (30%) to define: a) the number of AC (measured as 2×2 km grid cells that should be occupied by new populations, that is grid cells = new populations) required to fully compensate for predicted range loss and to halt the decline below the 30% of range loss; b) The number of cells necessary to compensate for range loss was calculated as the number of currently occupied cells lost under future climate due to unsuitable conditions. We used two Representative Concentration Pathways, +2.6 and +8.5 W/m2, optimistic and pessimistic scenarios, respectively. Availability of suitable areas for AC was also assessed within the current species distribution and within protected areas.</p> <p>Under the optimistic scenario, no taxa would lose more than 30% of their range and AC would not be required. Under the pessimistic scenario, roughly 90% of taxa showed a cell loss higher than 30%. Eight taxa were predicted to lose &gt;95% of their range. For these species, AC was required from 13 to 16 new populations (= 13 to 16 grid cells) per taxon to cap the range loss at 30%. For currently VU or EN species, an average number of 32 to 35 AC attempts would be necessary to fully compensate for their range loss under a pessimistic scenario. Suitable recipient sites within protected areas falling in their projected range were identified, allowing for short-distance AC.</p> <p>Synthesis. Combining SDMs and red listing thresholds under Criterion A has enabled the strategic planning of multiple-species AC minimising the effort in terms of new populations to be created and maximising the conservation benefit in terms of range loss compensation.</p>

opencc-zeroJan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record