Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
403
datasets available to search
ShareScore release 0.7.1
Dataset results
403 results for “occurrence data”
Spatiotemporal data for spotted lanternfly occurrence in the US
<p>An aggregated data set containing anonymized, spatiotemporal occurrence data for the spotted lanternfly (<em>Lycorma delicatula</em>, White 1845) in the United States. More details on the data, and additional tools to visualize it, can be found here: https://github.com/ieco-lab/lydemapr, and in the related publication.</p>
Data in support of 'ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension'
<p>Data in support of 'Chandler M, Sprintall J, Zilberman NV. (2025). ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension. <em>Journal of Geophysical Research: Oceans</em>. <a href="https://doi.org/10.1029/2025JC022899" target="_blank" rel="noopener">https://doi.org/10.1029/2025JC022899</a>'</p> <p> </p> <p>There are 2 netCDF files:</p> <ol> <li>p40tem1211_2312.nc</li> <li>synthetic_T_10day_px40_kuroshio_chandler2024.nc</li> </ol> <p><strong>p40tem1211_2312.nc </strong>contains the temperature sections from <a href="https://www-hrx.ucsd.edu/px40.html">HR-XBT transect PX40</a> objectively mapped onto a 10-m depth grid and a 0.1° longitudinal grid. <em>[LONGITUDE; LATITUDE; DEPTH; TIME; TEM]</em></p> <p><strong>synthetic_T_10day_px40_kuroshio_chandler2024.nc</strong> contains the synthetic temperature anomaly time series between the surface and 800-m deep at the western end of transect PX40 over the period from January-1993 to April-2023, as well as the temperature annual cycle needed for reconstructing the full synthetic temperature time series. <em>[time; depth; longitude; latitude; T_prime; T_ann]</em></p> <p> </p> <p>There is 1 MATLAB file:</p> <ol> <li>px40_synthetic_T.m</li> </ol> <p><strong>px40_synthetic_T.m</strong> is the MATLAB script used to produce the synthetic temperature anomaly time series saved in synthetic_T_10day_px40_kuroshio_chandler2024.nc.</p> <p> </p> <p>There is 1 Julia file:</p> <ol> <li>px40_synthetic_T_julia.jl</li> </ol> <p><strong>px40_synthetic_T_julia.jl</strong> is a Julia implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p>There is 1 R file:</p> <ol> <li>px40_synthetic_T_R.R</li> </ol> <p><strong>px40_synthetic_T_R.R</strong> is an R implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p><code>Version history:</code><br><code>v1.0.0 First uploaded (25-November-2024)</code><br><code>v1.0.1 Julia script uploaded (18-January-2025)</code><br><code>v1.0.2 R script uploaded (28-January-2025)</code><br><code>v1.1.0 Updated description of synthetic_T_10day_px40_kuroshio_chandler2024.nc to include reference to accepted publication (21-August-2025)</code></p>
Data for the paper 'Reducing networks of ethnographic codes co-occurrence in anthropology'
<p>Pseudonymized data supporting the paper "Reducing networks of ethnographic codes co-occurrence in anthropology", published in "Advances in Quantitative Ethnography. Fourth International Conference on Quantitative Ethnography (ICQE 2022), Copenhagen, Denmark, October 15–19, 2022, Proceedings", and edited by Amanda Barany and Crina Damsa. The paper is part of the POPREBEL project. The data were gathered in the spring and summer of 2021, as a part of a larger research project on populism in Central and Eastern Europe, to be completed by the end of 2022. They consist of 17 semi-structured interviews with Polish-speaking Internet users, who used social media to seek and share information about health against the backdrop of the COVID-19 pandemic. Research participants were asked about their opinion on the current state of affairs in their respective countries, and their political choices over the years and at present.</p> <p><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data).</p>
Grasshopper species occurrence data in Mt. Kilimanjaro
<p>Species occurence data of grasshopper in 60 plots on the southern slope of Mt. Kilimanjaro.</p> <p>Orthoptera assemblages were recorded on all study sites by repeatedly walking for 1.5 h on parallel tracks (distance between transects ca. 1-1.5 m) and recording all sighted species. In forested study sites, trees and bushes in the understory vegetation were shaken for approximately 1.5 h. Insects falling from the vegetation were gathered on white canvas laid on the forest floor. Species which could not be identified during visits were collected and later identified. Study sites were also visited at night where Ensifera were registered acoustically. Additionally, two rounds of sweep net sampling were conducted on study sites to collect small species which may have remained undetected during transect walks. One round was conducted during the cool dry season (July to October) and one during the warm dry season (December to March). During each sweep netting round, 100 sweeps with a 30-cm diameter sweep were taken and all collected specimens were identified in the laboratory. Species accumulation curves for Caelifera and Ensifera on Mt. Kilimanjaro were published in , showing that more than 90% of the grasshopper, locust and bushcricket fauna for Mt. Kilimanjaro have been registered.</p> <p>The KiLi project (2010-2018) is a German Science Foundation (DFG) funded research unit (DFG research unit FOR1246) that focuses on biodiversity and ecosystem processes along altitudinal and disturbance gradients on Mt. Kilimanjaro (Tanzania, Africa), capitalizing on its world-wide unique range of climatic and vegetation zones. The research unit comprises 2 central projects and 7 subprojects from various disciplines. On a total of 60 study sites in both natural and human-disturbed ecosystems biodiversity (e.g. plants, soil arthropods, ants, bees, frogs, lizards, bats, birds), related ecosystem processes (decomposition, seed dispersal, pollination, herbivory, predation), and biogeochemical processes and properties of ecosystems (climate, soil properties and nutrient status, regulation of water and carbon fluxes, trace gas emissions, primary productivity, functional diversity) are analyzed.</p>
Global taxonomic occurrence grids using GBIF data for species distribution models.
<p>To achieve large geographic coverage, species occurrence databases that are composed of ad hoc species data collections such as that provided by the Global Biodiversity Information Facility (GBIF) are often used. A drawback to using these data is their geographic sampling bias, in which some regions are more intensively sampled than others, while other areas have very little to none reported sampling effort. Uneven sampling effort can mislead conclusions about biodiversity patterns and species distributions (Gotelli & Colwell, 2001; Lobo, 2008).</p> <p>Here we provide taxonomic occurrence grids to help mitigate the effects of sampling bias in species distribution modeling. These grids can be used to exclude areas of (a custom-defined) low sampling effort from the background when sampling for pseudo-absences’ (Phillips et al., 2009; Barbet-Massin et al.,2012). The occurrence grids have a 1 degree spatial resolution using WGS 84 as the geographic coordinate system. Each 1 degree grid cell contains the number of records present in GBIF corresponding to a specific taxonomic group: plants, mammals, reptiles, amphibians, birds and molluscs.</p> <p>To construct the occurrence grids, we used the 1- by 1-degree world latitude and longitude vector grid provided by ESRI (Redlands, California). It has a custom license which permits it reuse as long as ESRI is cited. It was downloaded from : <a href="https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7">https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7</a></p> <p>To map spatial sampling effort, the number of georeferenced occurrences corresponding to each taxonomic group contained by each 1- by 1-degree grid cell were counted. The grids were then converted to GeoTIFFs. The raster values correspond to the number of occurrences reported for the grid cells. For the purposes of the <a href="https://osf.io/7dpgr/">TrIAS project</a>, grid cells with fewer than 5 occurrences were removed. The TrIAS taxonomic occurrence grids are used as inputs to the TrIAS risk modelling and mapping workflow: https://github.com/trias-project/risk-modelling-and-mapping. Full (with all grid cells containing at least one occurrence) taxonomic occurrence grids are also provided.</p> <p>GBIF data for each taxonomic group were downloaded using the following criteria: “Basis of Record”: Observation, Machine Observation, Human Observation, Specimen, Material sample, Literature Occurrence, Unknown evidence., "HasCoordinate is true", "HasGeospatialIssue is false", "TaxonKey is Amphibia", "Year 1975-2005".</p> <p><strong>Raster Attributes</strong></p> <table> <tbody> <tr> <td> <p>Attribute</p> </td> <td> <p>Description</p> </td> </tr> <tr> <td> <p>OID</p> </td> <td> <p>numeric row ID</p> </td> </tr> <tr> <td> <p>Value</p> </td> <td> <p>the number of records contained in the grid cell</p> </td> </tr> <tr> <td> <p>Count</p> </td> <td> <p>the number of times the value appears in the raster</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p>The extent of each taxonomic occurrence grid:</p> <ul> <li> <p>longitude -180.0; latitude -90.0 (southwest corner)</p> </li> <li> <p>longitude 180.0; latitude 90.0 (northeast corner)</p> </li> </ul> <p> </p> <p><strong>Files:</strong></p> <p>TrIAS taxonomic occurrence grids</p> <p>amphib_1deg_min5.tif</p> <p>birds_1deg_min5.tif</p> <p>mammals_1deg_min5.tif</p> <p>molluscs_1deg_min5.tif</p> <p>reptiles_1deg_min5.tif</p> <p> </p> <p>Raw taxonomic occurrence grids</p> <p>amphib_1deg_grid.tif</p> <p>birds_1deg_grid.tif</p> <p>mammals_1deg_grid.tif</p> <p>molluscs_1deg_grid.tif</p> <p>reptiles_1deg_grid.tif</p> <p><br> </p> <p> </p> <p> </p>
Data from: Heterogeneity in habitat and nutrient availability facilitate the co-occurrence of N2 fixation and denitrification across wetland - stream - lake ecotones of Lakes Superior and Huron
Great Lakes coastlines are mosaics of wetland, stream, and lake habitats, characterized by a high degree of spatial heterogeneity that may facilitate the co-occurrence of seemingly incompatible biogeochemical processes due to variation in environmental factors that favor each process. We measured nutrient limitation and rates of N2 fixation and denitrification along transects in 5 wetland - stream - lake ecotones with different nutrient loading in Lakes Superior and Huron and hypothesized that rates of both processes would be related to nutrient limitation status, habitat type, and environmental characteristics including temperature, nutrient concentrations, and organic matter quality. This data package includes information on sampling sites, dates and locations; rates of N fixation and denitrification measured at each site, date and transect location; and biomass information from nutrient diffusing substrates deployed on the study transects.
Mammal occurrence data derived from camera traps in grassland-shrubland ecotones at 24 sites in the Jornada Basin, southern New Mexico, USA, 2014-ongoing
The objective of this ongoing study is to investigate how abundance, distribution, and activity of mammals (>= 1 kg) vary across grassland to shrubland ecotones in the northern Chihuahuan Desert. This dataset includes animal occurrence data derived from camera trap images captured in 24 grassland-to-shrubland ecotone sites in the Jornada Basin, Dona Ana County, New Mexico, USA. The data set contains occurrence records from 14 mammal species with the date and time a species was detected. Also included are the number of individuals in a photo, operational dates and number of functional camera days for cameras, total number of trap nights a camera was active, and geographical coordinates of camera trap locations. Sampling is ongoing and occurs during the monsoon season from July-November. Sampling has occurred annually since 2014.
Occurrence data used to create species distribution models and apply an evaluation method
<p>These two files containing a table with three columns: species names, longitude, latitude. Each row of the tables represents a georeferenced presence record for the corresponding species. The original presence data were downloaded from the GBIF database and after going through a cleaning process, we ended with these records that passed all the tests.</p> <p>These datasets were used to create species distribution models (SDMs) that were then used to apply a new method to evaluate the performance of different SDMs. Jiménez & Soberón (2020)</p>
Risk assessment on Glycoalkaloids in feed and food: Occurrence data in food and feed submitted to EFSA and dietary exposure assessment for humans
<p><strong>UPDATE to version 2 of this upload:</strong></p> <p>Also the raw (no data cleaning applied to it) occurrence dataset as extracted from EFSA DWH is provided <em>in csv format</em>. This dataset is compliant with EFSA SSD model and contains two additional columns documenting issues identified in the cleaning process (column: issue) and the action taken (column: action) to address the issue (e.g. delete record or update values in specific fields).</p> <p><strong>Description - Version 1</strong></p> <p><strong>Annex: Tables on GAs on occurrence data in food and feed, and dietary exposure assessment for humans</strong></p> <p>Table A.1. Dietary surveys used for the estimation of acute dietary exposure to GA</p> <p>Table A.2. Number of results and samples per food category submitted to EFSA through the continuous call for data</p> <p>Table A.3. Analytical results excluded from the final dataset used to estimate dietary exposure and the criteria applied for exclusion</p> <p>Table A.4. Occurrence of alpha-chaconine and alpha-solanine (UB mg/kg) in the samples included in the final dataset (left censored results highlighted in yellow)</p> <p>Table A.5. European Starch Association data on feed and potatoes for starch</p> <p>Table A.6. Details acute assessment across surveys (consumption days only)</p> <p>Table A.7. Comparison of exposure summary results obtained using the uniform vs the normal distribution for reduction factors</p>
Porto Santo landscape features and endemic lichens occurrence data
<p>Landscape features of Porto Santo island of and observation data of endemic lichens belonging to Sparrius et al. 2017, Bryologist.</p>
Research Data and Code for "Interdisciplinarity in the 17th Century? A Co-Occurrence Analysis of Early Modern German Dissertation Titles"
<p>This dataset documents results and code for the paper "Interdisciplinarity in the 17th Century? A Co-Occurrence Analysis of Early Modern German Dissertation Titles" by Stefan Heßbrüggen-Walter, forthcoming in *Synthese*. The data to be processed are contained in four files, derived from a larger dataset related to German dissertations and sourced from the national bibliography of 17th century German prints *VD 17* that will be released at a later date. More information can be found in the file `README.md`. </p>
Supplementary Data: Fungal biostarter and bacterial occurrence of dry-aged beef: the sensory quality and volatile aroma compounds after 21 days of aging
<p>This dataset contains data generated during realisation of the project Tango-IV-C/0005/2019: Biostarters development for dry aged beef production, funded by National Centre for Research and Development (Poland). These data were used to prepare paper entitled "Fungal biostarter and bacterial occurrence of dry-aged beef: the sensory quality and volatile aroma compounds after 21 days of aging" by Wiesław Przybylski , Danuta Jaworska, Paweł Kresa, Grzegorz Michał Ostrowski, Magdalena Płecha, Dorota Korsak, Dorota Derewiaka, Lech Adamczak, Urszula Siekierko, Julia Pawłowska.</p>
Spatial clustering of Neobuccinum eatoni occurrence data for potential distribution modeling
<p>The occurrence dataset for <em>Neobuccinum eatoni</em> was compiled through filtration process, starting with records from the Global Biodiversity Information Facility (GBIF) and supplemented by museum specimens and additional sources like SOMBASE, iBOL, NIWA, ANTABIF, and SCAR-AntOBIS. Further data were sourced from the National Museum of Natural History in Paris, the University of Vigo, and recent fieldwork in Antarctica, Heard Island, and Kerguelen Island. Records were meticulously screened to remove misidentified specimens, inaccurate locations, duplicates, and outdated entries, ensuring accuracy and relevance. To address spatial autocorrelation, clustering methods divided the data into distinct geographic clusters, producing a refined dataset used to model <em>N. eatoni</em>'s potential distribution with enhanced predictive reliability by reducing spatial autocorrelation effects.</p>
Blair et al. 2020: Machine learning identification of ground beetles (repackaging of occurrences published by the NEON Biorepository Data Portal)
Blair, J.; Weiser, M. D.; Kaspari, M.; Miller, M.; Siler, C.; Marshall, K. E. 2020. Robust and simplified machine learning identification of pitfall trap-collected ground beetles at the continental scale. Ecology and Evolution 10 (23): 13143-13153. https://doi.org/10.1002/ece3.6905 Additional NEON samples (not yet archived at the Biorepository) were used in this research: full list of occurrences used.
Stachewicz et al. 2021: Trait correlation, phylogenetic signal in Carabidae morphology (repackaging of occurrences published by the NEON Biorepository Data Portal)
Stachewicz JD, Fountain-Jones NM, Koontz A, Woolf H, Pearse WD, Gallinat AS. 2021. Strong trait correlation and phylogenetic signal in North American ground beetle (Carabidae) morphology bioRxiv 02.12.431029; doi: https://doi.org/10.1101/2021.02.12.431029 Many NEON samples and specimens used in this work resulted from NEON prototype data and will not be archived in the Biorepository. See the appendices in the above linked article for a full list of NEON samples and specimens and their associated collection data. Additionally, see appendices of above linked article for specimen-level morphological trait measurements and genetic sequence data.
Staines & Staines 2021: The Geadephaga (Coleoptera: Carabidae and Rhysodidae) of SERC (repackaging of occurrences published by the NEON Biorepository Data Portal)
Linked records of Carabidae and Rhysodidae specimens from the SERC site, collected between 2015-2018, included in the following publication: Staines, C.L. & Staines, S. L. The Geadephaga (Coleoptera: Carabidae and Rhysodidae) of the Smithsonian Environmental Research Center, Maryland. 2021. Banisteria 55: 75-100. Original abstract: "An inventory of the Geadephaga (Coleoptera) at the Smithsonian Environmental Research Center, Anne Arundel County, Maryland is being conducted. Pitfall traps were placed and monitored from 2015 to 2018. From 2017 to 2020 directed collecting efforts were made to document the Geadephaga of the facility. A total of 111 Geadephaga species was collected: Carabidae - 110, Rhysodidae - 1." Research article available for download: https://virginianaturalhistorysociety.com/banisteria/pdf-files/ban55/Staines_SERC_Geadephaga.pdf
Decomissioned Site: Arthur Brook's (D01 ARTH) Legacy and Prototype Aquatic invertebrates (repackaging of occurrences published by the NEON Biorepository Data Portal)
This collection contains legacy and prototype NEON aquatic invertebrates from the decommissioned Arthur Brook's site in Worcester County, Massachusets. These samples and specimens were collected using NEON protocol NEON.DP1.20120 and are summarized in NEON Prototype dataset e7d152a1-1181-4c7e-ac75-ae707ffc7299. This data is available here.
Decomissioned Site: Ichawanochaway Creek (D03 ICHA) Legacy and Prototype Aquatic invertebrates (repackaging of occurrences published by the NEON Biorepository Data Portal)
This collection contains legacy and prototype NEON aquatic invertebrates from the decommissioned Ichawanochaway Creek site in Baker County, Georgia. These samples and specimens were collected using NEON protocol NEON.DP1.20120 and are summarized in NEON Prototype dataset 1599af29-fad4-4721-9b09-f8324b672d50. This data is available here.
Decommissioned Site: Mameyes (D04 MAME) Belowground Plant Biomass (Megapit) (repackaging of occurrences published by the NEON Biorepository Data Portal)
These samples are associated with NEON prototype dataset: NEON Decommissioned Site data: Root sampling, chemistry, and isotopes (Megapit) from D04 MAME site (a36e6ce2-d845-13e9-c881-c1d4e5881a53) DOI: 10.48443/j1hp-f273 NEON (National Ecological Observatory Network). NEON Decommissioned Site data: Root sampling, chemistry, and isotopes (Megapit) from D04 MAME site, v1 (10.48443/j1hp-f273). https://doi.org/10.48443/j1hp-f273. Dataset accessed from https://data.neonscience.org on July 7, 2021
Identified invertebrate bycatch from beetle pitfall traps at SJER and SOAP, 2017 - 2018 (repackaging of occurrences published by the NEON Biorepository Data Portal)
California permit requirements necessitated a more thorough identification of beetle pitfall samples than is typical of this protocol. These invertebrate bycatch samples therefore have occurrence associations that indicate their contents in both the NEON Biorepository and main NEON data portals. See NEON prototype dataset 9bc959c-148b-aaad-aa35-2d0805327428 available here.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.