Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,119
datasets available to search
ShareScore release 0.7.1
Dataset results
8,119 results for “species distribution”
Figure 12. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 12. - Distribution of P. confusa (Maui, Hawaii) and P. quadrisetosa (Kauai).
Figure 14. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 14. - Distribution of P. wirthi (Oahu and Kauai)
Figure 11. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 11. - Distribution of P. bifurcata (Kauai, Oahu, Molokai and Maui).
Figure 7. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 7. - Distribution of P. hardyi (Kauai).
Figure 3. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 3. - Habitus of the holotype female of P. hardyi in dorsal view.
Figure 1. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 1. - Habitus of a paratype male of P. hardyi in lateral view.
Figure 2. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 2. - Habitus of the holotype female of P. hardyi in lateral view.
Figure 13. from: Studies in Hawaiian Diptera III: New Distributional Records for Canacidae and a New Endemic Species of Procanace - Biodiversity Data Journal 4: e5611 (08 April 2016) https://doi.org/10.3897/BDJ.4.e5611
Figure 13. - Distribution of P. constricta (Molokai, Maui, Hawaii) and P. williamsi (Oahu).
Plant species distribution survey and its explanatory variables
<p>Ten common herbaceous species were selected based on a survey of a 6.4-km² upstream catchment named « Bourdic » in southern France. Approximately 74 % of the catchment is agricultural (mainly vineyards), and 26 % is semi-natural (mainly woodlands and shrubs). The catchment has a Mediterranean climate with heavy rainfalls causing significant Hortonian runoff. The mean annual temperature is 14°C, and precipitation ranges from 600 to 800 mm per year with a drier period from March to October. Annual potential evapotranspiration is about 1100 mm. The altitude ranges from 55 a.s.l. at the outlet at the northeast to 128 m a.s.l. at the northwest.</p> <p>The surveys were conducted in July-August 2013 according to a non-destructive sampling procedure using GPS with an Android self-developed application; this enabled a location accuracy of 2 m. Agricultural ditches, including roadside ditches, were part of the study. Thirty-five kilometres of the drainage network (46%) were surveyed for presence/absence of the species. The remaining ditches were excluded from the analysis because surveying them was impractical or because recent management practices impaired species identification. After the survey, the georeferenced data were exported in a shapefile data format with line features.</p> <p>Also, explanatory variables of the dataset were reported, such as the geomorphological variables at the landscape scale that included the distance to the outlet (<strong>Doutlet</strong>), the drained surface area (<strong>Drain</strong>), the Multiresolution Index of Valley Bottom Flatness (<strong>Mrvbf</strong>), and the sun exposure of the slopes (<strong>Northness</strong>). The geomorphological variables at the local (ditch) scale were the slope (<strong>Slope</strong>) and solar radiation (<strong>Solar</strong>). All these variables are derived from a Digital Elevation Model (MNT) and a Digital Surface Model (MNS) taken in 2001 using an aerial lidar.</p> <p>We added also the distance to natural lands (<strong>Dnat</strong>) and distance to roads (<strong>Droad</strong>) on the basis of the manual classification of an orthophoto of the area taken in 2012.</p>
Data and scripts for "Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"
<p>The README explains how to reproduce the analyses presented in the paper <strong>"Joint species-trait distribution modelling: The role of intraspecific trait variation in community assembly"</strong> by Abrego et al.</p> <p>The input data for the script pipeline is the file “Kilpisjarvi_plant_data.csv”. This file includes the data on the plants and their traits in the long format. Hence, each row of the data matrix corresponds to measurements on one plant species in one study plot. The joint species-trait distribution modelling (JSTDM) pipeline that analyses these data consists of the following R-scripts.</p> <p>· <strong>S1_define_JSTDM_models.R</strong>. This script defines the JSDTM models (null model and environmental model) that include five response types for each species: the presence-absence, abundance conditional on presence, and the plot-level trait values of specific leaf area (SLA), leaf area (LA) and mean height (MH). The model is defined in the Hierarchical Modelling of Species Communities (HMSC) framework utilizing the R-package Hmsc. The models are saved in the file “unfitted_models.RData”.</p> <p>· <strong>S2_fit_models.R. </strong>This script loads the unfitted models and fits them using the posterior sampling methods implemented in the R-package Hmsc. The models are fitted with increasing thinning until thin=100, which value was used to generate the results of the paper. The fitted models are saved in the file "models_thin_100_samples_250_chains_4.Rdata".</p> <p>· <strong>S3_plot_Omega_matrices.R. </strong>This script loads the fitted models and plots the association matrices (Fig. 2 of the paper). The csv file containing the values used to construct Fig. 2 is also given (figure2Cdata.csv and figure2Ddata.csv).</p> <p>· <strong>S4_show_VP_Beta_Gamma.R. </strong>This script loads the fitted models and extracts information on the variance partitionings (VP; Figs. S3 and S4 of the paper), the relationships between response types and environmental predictors (beta; Fig. S2 of the paper), and the relationships between response types and species-level traits (gamma; Fig. S5 of the paper). The csv file containing the values used to construct Fig. S2 (figureS2Adata.csv and figureS2Bdata.csv), Fig. S3 (figureS3data.csv), Fig. S4 (figureS4data.csv) and Fig S5 (figureS5data.csv) are also given.</p> <p>· <strong>S5_conditional_cross_validation.R</strong>. This script performs 10-fold cross validation to the data to test the predictive power related to the modelled plant traits. The script performs both regular (unconditional) cross-validation where all data are masked for the test fold, and conditional cross-validation where only the trait data (but not the abundance data) are masked for the test fold.</p> <p>· <strong>S6_show_conditional_cross_validation_results.R. </strong>This script plots the results of cross-validation (Fig. 3 of the paper). The csv file containing the values used to construct Fig. 3 is also given (figure3Adata.csv and figure3Bdata.csv).</p> <p>· <strong>S7_scenario_predictions.R</strong>. This script performs the scenario simulations described and shown in Fig. 4 of the paper. The csv file containing the values used to construct Fig. 4 is also given (figure4Bdata.csv and figure4Cdata.csv).</p>
Australian Frog Atlas: Fine-scale species distribution maps informed by the FrogID dataset
<p>This is version three of the Australian Frog Atlas: shapefiles and KML files of the most up to date, accurate and detailed set of Australian frog species distribution maps available, and both a shapefile and detailed PDF map of Australian frog species richness. All of these are available as a free resource under the Creative Commons Attribution (CC-BY) 4.0 License. When using any material from the Australian Frog Atlas, please refer to and cite Cutajar et al. (2022), outlined in full below, which introduces the dataset and describes it in detail.</p> <p>Cutajar. T. P., Portway, C. D., Gillard, G. L., and Rowley, J. J. L. (2022) Australian Frog Atlas: Fine-scale species distribution maps informed by the FrogID dataset. <em>Technical Reports of the Australian Museum Online</em>. 36: 1-48. <a href="https://doi.org/10.3853/j.1835-4211.36.2022.1789">https://doi.org/10.3853/j.1835-4211.36.2022.1789</a></p> <p><strong>Caveat:</strong> The occurrence data used in the development of these maps are from a range of sources. Every effort has been made to ensure accuracy and completeness, including vetting of data using published literature. However, no guarantee is given, nor responsibility taken by the authors or their institutions for errors or omissions, nor in respect of any information or advice given in relation to, or as a consequence of, the maps and data provided herein. These maps are indicative only and aim to capture the known and presumed distributions of individual frog species and frog species richness within Australia. A taxonomic revision of Australo-Papuan tree frogs published in June 2025 proposes the reclassification of more than 120 Australian tree frog species (in the genera <em>Litoria</em> and <em>Cyclorana</em>) into 22 genera. This proposed change has not been adopted in this version of the Australian Frog Atlas. </p> <p> </p> <p> </p>
Data from: How far can I extrapolate my species distribution model? Exploring Shape, a novel method
<p>Species distribution and ecological niche models (hereafter SDMs) are popular tools with broad applications in ecology, biodiversity conservation, and environmental science. Many SDM applications require projecting models in environmental conditions non-analog to those used for model training (extrapolation), giving predictions that may be statistically unsupported and biologically meaningless. We introduce a novel method, Shape, a model-agnostic approach that calculates the extrapolation degree for a given projection data point by its multivariate distance to the nearest training data point. Such distances are relativized by a factor that reflects the dispersion of the training data in environmental space. Distinct from other approaches, Shape incorporates an adjustable threshold to control the binary discrimination between acceptable and unacceptable extrapolation degrees. We compared Shape's performance to five extrapolation metrics based on their ability to detect analog environmental conditions in environmental space and improve SDMs suitability predictions. To do so, we used 760 virtual species to define different modeling conditions determined by species niche tolerance, distribution equilibrium condition, sample size, and algorithm. All algorithms had trouble predicting species niches. However, we found a substantial improvement in model predictions when model projections were truncated independently of extrapolation metrics. Shape's performance was dependent on extrapolation threshold used to truncate models. Because of this versatility, our approach showed similar or better performance than the previous approaches and could better deal with all modeling conditions and algorithms. Our extrapolation metric is simple to interpret, captures the complex shapes of the data in environmental space, and can use any extrapolation threshold to define whether model predictions are retained based on the extrapolation degrees. These properties make this approach more broadly applicable than existing methods for creating and applying SDMs. We hope this method and accompanying tools support modelers to explore, detect, and reduce extrapolation errors to achieve more reliable models.</p>
Mapping multiscale breeding bird species distributions across the United States and evaluating their conservation applications
<p>Species distribution models are vital to management decisions that require understanding habitat use patterns, particularly for species of conservation concern. However, the production of distribution maps for individual species is often hampered by data scarcity, and existing species maps are rarely spatially validated due to limited occurrence data. Furthermore, community-level maps based on stacked species distribution models lack important community assemblage information (e.g., competitive exclusion) relevant to conservation. Thus, multispecies, guild, or community models are often used in conservation practice instead. To address these limitations, we aimed to generate fine-scale, spatially-continuous, nationwide maps for species represented in the North American Breeding Bird Survey (BBS) between 1992-2019. We generated ensemble models for each species at three spatial resolutions – 0.5, 2.5, and 5 km – across the conterminous United States. We also compared species richness patterns from stacked single-species models with those of 19 functional guilds developed using the same data to assess the similarity between predictions. We successfully modeled 192 bird species at 5-km resolution, 160 species at 2.5-km resolution, and 80 species at 0.5-km resolution. However, the species we could model represent only 28-56% of species found in the conterminous US BBS surveys across resolutions owing to data limitations. We found stacked maps and guild maps generally had high correlations across resolutions (median = 84%), but spatial agreement varied regionally by resolution and was most pronounced between the East and West at the 5-km resolution. The spatial differences between our stacked maps and guild maps illustrate the importance of spatial validation in conservation planning. Overall, our species maps are useful for single-species conservation and can support fine-scale decision-making across the United States, and can also support community-level conservation when used in tandem with guild maps. However, there are still data scarcity issues for many species of conservation concern when using the BBS for single-species models.</p>
Fig. 4 in Species Diversity And Distribution Of Synanthropic Acarid Mites (Acariformes, Acaridia) In Transcarpathia
Fig. 4. Correlation between the studied parameters.
Fig. 3 in Species Diversity And Distribution Of Synanthropic Acarid Mites (Acariformes, Acaridia) In Transcarpathia
Fig. 3. Indices of species biodiversity of acarid mites in Transcarpathia.
Data for: The meta-analysis of the effects of spatial sampling bias correction on presence only species distribution models
<p>This dataset contains information extracted from 70 studies identified through a systematic review of the peer-reviewed literature (Web of Science and SCOPUS databases both searched on the 13/02/2023) to evaluate the effect of spatial sampling bias correction methods in presence-only species distribution models.</p>
Patterns in the genetic structure of 49 lowland rain forest tree species co-distributed on opposite sides of the Northern Andes
<p>The Andes are a major dispersal barrier for lowland rain forest plants and animals, yet hundreds of lowland tree species are distributed on both sides of the Northern Andes, raising questions about how the Andes influenced their biogeographic histories and population genetic structure. To explore these questions, we generated standardized datasets of thousands of SNPs from paired populations of 49 tree species co-distributed in rain forest tree communities located in Panama and Amazonian Ecuador and calculated genetic diversity (<em>π</em>) and absolute genetic divergence (<em>d</em><sub>XY</sub>) within and between populations, respectively. We predicted (1) higher genetic diversity in the ancestral source region (east or west of the Andes) for each taxon, and (2) correlation of genetic statistics with species attributes, including elevational range and life-history strategy. We found that genetic diversity was higher in putative ancestral source regions, possibly reflecting founder events during colonization. We found little support for a relationship between genetic divergence and species attributes except that species with higher elevational range limits exhibited higher <em>d</em><sub>XY</sub>, implying older divergence times. One possible explanation for this pattern is that dispersal through mountain passes declined in importance relative to dispersal via alternative lowland routes as the Andes experienced uplift. We found no difference in mean genetic diversity between populations in Central America and the Amazon. Overall, our results suggest that dispersal across the Andes has left enduring signatures in the genetic structure of widespread rain forest trees. We outline additional hypotheses to be tested with species-specific case studies.</p>
The global distribution of plants used by humans datasets: list of utilised species, occurrence data and model outputs at 10 arc-minutes spatial resolution
<p>Datasets and model outputs used to map the global distribution of utilised plants by humans. The folder is composed of two subfolders <em>raw_data</em> and <em>processed_data</em> containing respectively the list of utilised plant species modelled -<em>utilised_plants_species_list.csv</em>-, and their occurrence data -<em>occurrence_data.zip-</em> and predicted distribution -<em>species_proba_per_cell.rds-.</em></p> <p> </p> <ul> <li>The file <em>utilised_plants_species_list.csv</em> in the <em>raw_data</em> folder contains a<strong> </strong>list of 35687 plant species (and hybrids) used by humans and 10 plant use categories with the following 14 fields:</li> </ul> <p><strong>plant_ID:<em> </em></strong>plant identifier number ranging from between 1-35687</p> <p><strong>binomial_acc_name:</strong> binomial accepted name of the plant species</p> <p><strong>author_acc_name</strong>: name of the author(s)</p> <p><strong>is_hybrid:</strong> logical TRUE or FALSE indicating whether the species is an hybrid or not.</p> <p><strong>AnimalFood:</strong> forage and fodder for vertebrate animals only.</p> <p><strong>EnvironmentalUses:</strong> examples include intercrops and nurse crops, ornamentals, barrier hedges, shade plants, windbreaks, soil improvers, plants for revegetation and erosion control, wastewater purifiers, indicators of the presence of metals, pollution, or underground water.</p> <p><strong>Fuels:</strong> charcoal, petroleum substitutes, fuel alcohols, etc. Given the importance of energy plants for people, those were distinguished from Materials.</p> <p><strong>GeneSources:</strong> wild relatives of major crops which may possess traits associated with biotic or abiotic resistance and may be valuable for breeding programs.</p> <p><strong>HumanFood:</strong> food for humans only, including beverages and food additives.</p> <p><strong>InvertebrateFood:</strong> plants consumed by invertebrates used by humans, such as bees, silkworms, lac insects and edible grubs.</p> <p><strong>Materials:</strong> woods, fibers, cork, cane, tannins, latex, resins, gums, waxes, oils, lipids, etc. and their derived products.</p> <p><strong>Medicines:</strong> both human and veterinary.</p> <p><strong>Poisons:</strong> plants which are poisonous to both vertebrates and invertebrates, both accidentally and intentionally, e.g., for hunting and fishing, molluscicides, herbicides, insecticides.</p> <p><strong>SocialsUses:</strong> plants used for social purposes, which cannot be defined as food or medicine, for instance, masticatories, smoking materials, narcotics, hallucinogens and psychoactive drugs, and plants with ritual or religious significance.</p> <p><strong>Totals:</strong> total number of uses recorded for a species</p> <p> </p> <ul> <li>The zipfile <em>occurrence_data.zip</em> in the <em>processed_data</em> folder contains 35687 Comma Separated Values (CSV) files, one for each species, containing curated geographic occurrence records used to build species distribution models with the following 14 fields:</li> </ul> <p><strong>Species:</strong> the binomial accepted name of the species</p> <p><strong>Fullname:</strong> same as species</p> <p><strong>decimalLongitude:</strong> the geographic longitude of the occurrence records of the species in decimal degrees</p> <p><strong>decimalLatitude:</strong> the geographic latitude of the occurrence records of the species in decimal degrees</p> <p><strong>countryCode:</strong> a three-letter standard abbreviation for the country of the occurrence locality</p> <p><strong>coordinateUncertaintyinMeters</strong>: indicator for the accuracy of the coordinate location, described as the radius of a circle around the stated point location</p> <p><strong>year:</strong> year of the observation of the occurrence record of the species</p> <p><strong>individualCount:</strong> the number of individuals present at the time of the observation</p> <p><strong>gbifID:</strong> unique identifier number for the occurrence from the original database</p> <p><strong>basisOfRecords:</strong> the type of the individual record, e.g. observation, physical specimen, fossil, living ex-situ, culture collection specimen</p> <p><strong>institutionCode</strong>: the name of the institution or organization listed as the data publisher on GBIF</p> <p><strong>establishmentMeans:</strong> statement about whether an organism has been introduced to a given place and time through the direct or indirect activity of modern humans</p> <p><strong>is_cultivated_observation:</strong> whether or not an organism is cultivated</p> <p><strong>sourceID:</strong> name of the source database</p> <p> </p> <ul> <li>The file <em>species_proba_per_cell.rds</em> in the <em>processed_data</em> folder is<em> a R Data Serialization </em>(RDS) file containing a data.table object with the following 3 fields:</li> </ul> <p><strong>plant_ID:</strong><em> </em>plant identifier number ranging from between 1-35687</p> <p><strong>proba:</strong> species occurrence probability</p> <p><strong>cell:</strong><em> </em>raster grid cell number between 1-2251762</p> <p>This object can be used in combination with a raster layer to reconstruct the modelled distribution of each species or retrieve species richness and endemism.</p>
Figure 2 in Selymbria Stål, 1861 (Hemiptera: Cicadidae: Tibicininae): description of a new species with notes on the genus taxonomy and distribution
Figure 2. Selymbria amazonensis sp. nov., paratype female (A-F): (A) Habitus in dorsal view; (B) Head and pronotum in dorsal view; (C) Head in ventral view; (D) Operculum in latero-ventral view; (E) Terminalia in ventral view (F) Terminalia in lateral view. Scale bars: A = 1 cm; B, C, E, F = 2 mm; D = 1 mm.
Evaluation data for: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study
<p>All of the evaluation data for the simulations in the paper: Adaptive sampling by citizen scientists improves species distribution model performance: a simulation study. We considered the impact of five adaptive sampling methods on the performance of species distribution models (SDMs), please see the paper for more information. Contained in this repository are the evaluation metrics (AUC, mean square error (MSE) and correlation) for SDMs before and after adaptive sampling has taken place. The MSE and correlation evaluation metrics were calculated against the true distributions of the species. These files are those with "combined_outputs" in the titles. The repository also contains the observations of all the species in the simulations both before and after adaptive sampling (the files with "all_observations" in the title.</p> <p>These datasets are to be used with the plotting and evaluation scripts in the GitHub repository associated with the paper.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.