Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,015
datasets available to search
ShareScore release 0.7.1
Dataset results
3,015 results for “Occurrence”
Occurrence of animals along five transects at the Jornada Basin LTER site from 1989-1994
This dataset contains data on the occurrence of rabbits, birds, and lizards observed along the Jornada Basin LTER (II) animal transects in southern New Mexico, USA. Five, 1 km transects were established, each in a different vegetation zone, near the current NPP study locations C-CALI, G-IBPE, M-NORT, P-COLL, and T-EAST. An observer walked each transect once every two weeks from early 1989 through 1994 recording animals observed along the transects. The data consists of species names, numbers of individuals, and perpendicular distances observed from transects, as well as weather and other context observations. Observation history and species codes are described in additional files. This study is complete.
Red knot occurrence, prey density, island morphology, and climate change in the Virginia Barrier Islands (2009-2023)
Global climate change is reshaping dynamic coastal ecosystems, with uncertain consequences for migratory shorebirds such as the federally threatened red knot (Calidris canutus rufa) that rely on coastal staging sites during migration. Understanding how sea-level rise and changing climate drivers affect red knot foraging ecology is critical for informing conservation and management at coastal staging sites. We integrated long-term biological, geomorphological, and climatological data to examine the direct and indirect pathways influencing red knots and their prey at intertidal foraging sites on the Virginia Barrier Islands during spring migration (May 21 - 28, 2009-2023). Using piecewise structural equation modeling, we tested hypothesized two causal networks linking 1) red knot occurrence and 2) densities of their main invertebrate prey to habitat characteristics, island morphology, geomorphic change, and climate drivers of ecosystem change. Red knots were indirectly affected by geomorphic change and climate drivers through bottom-up effects on invertebrate communities mediated by island morphology. Accelerated shoreline change narrowed islands, reducing invertebrate density and richness and indirectly decreasing red knot occurrence. Storms interacted with global climate oscillations to drive erosion or accretion of beaches, with variable effects on invertebrate density and red knot occurrence. Invertebrate responses were taxon-specific: shoreline change directly increased blue mussel density but indirectly reduced coquina clam and crustacean densities by narrowing island width, while storms impacts on crustacean density were mediated by beach width. Our findings suggest that accelerated ecosystem change under future climate scenarios may alter foraging conditions for red knots and other migratory shorebirds in the Virginia Barrier Islands, with broader implications for long-term population resilience.
Van Allen Probes Occurrence Rates of Electromagnetic Ion Cyclotron (EMIC) Waves with Rising Tones
<p>CSV files with the values for the occurrence rates of electromagnetic ion cyclotron (EMIC) waves with rising tones observed by the Van Allen Probes from 2012-09-07 to 2016-07-01 from the paper</p><p>Sigsbee, K., Kletzing, C. A., Faden, J., & Smith, C. W. (2023). Occurrence rates of electromagnetic ion cyclotron (EMIC) waves with rising tones in the Van Allen Probes data set. Journal of Geophysical Research: Space Physics, 128, e2022JA030548. https://doi.org/10.1029/2022JA030548 </p><p>The below files contain the values from Figures 5 and 6. The first row of each file gives the lower value of each L shell bin (0.0, 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.5, 6.0, 7.0, 7.5). The first column of each file gives the magnetic local time (MLT) values (0-23) for each bin. </p><p>rbspab_lshellmlt_minutes_20120907_to_20160701.csv gives the number of minutes spent by the Van Allen Probes in each bin of L shell and MLT.</p><p>rbspab_emic_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes all EMIC waves were observed in each bin of L shell and MLT.</p><p>rbspab_h_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes H+ band EMIC waves were observed in each bin of L shell and MLT.</p><p>rbspab_hr_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes H+ band EMIC waves with rising tones were observed in each bin of L shell and MLT.</p><p>rbspab_he_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes He+ band EMIC waves were observed in each bin of L shell and MLT.</p><p>rbspab_her_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes He+ band EMIC waves with rising tones were observed in each bin of L shell and MLT.</p><p>rbspab_o_lshellmlt_pcnt_20120907_to_20160701.csv gives the percentage of minutes O+ band EMIC waves with rising tones were observed in each bin of L shell and MLT.</p><p>The below files contain the values from Figures 7-13. The first row of each file gives the lower value of each bin of the radial distance RXY in the XY SM plane (0.0, 0.5, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.5, 6.0, 7.0, 7.5) in Earth radii (RE). The first column of each file gives the lower value of each bin of Z SM in RE (-2.0, -1.75, -1.5, -1.25, -1.0, 0.0, 1.0, 1.25, 1.50, 1.75). Separate files are provided for four MLT sectors: midnight (21 MLT to 3 MLT), dawn (3 MLT to 9 MLT), noon (9 MLT to 15 MLT), and dusk (15 MLT to 21 MLT).</p><p>Number of minutes spent by the Van Allen Probes in bins of RXY and Z SM (Figure 7):</p><p>rbspab_rxyzsm_minutes_midnight_20120907_to_20160701.csv, rbspab_rxyzsm_minutes_dawn_20120907_to_20160701.csv, rbspab_rxyzsm_minutes_noon_20120907_to_20160701.csv, rbspab_rxyzsm_minutes_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes all EMIC waves were observed in bins of RXY and Z SM (Figure 8):</p><p>rbspab_emic_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_emic_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_emic_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_emic_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes H+ band EMIC waves were observed in bins of RXY and Z SM (Figure 9):</p><p>rbspab_h_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_h_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_h_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_h_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes He+ band EMIC waves were observed in bins of RXY and Z SM (Figure 10):</p><p>rbspab_he_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_he_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_he_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_he_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes O+ band EMIC waves were observed in bins of RXY and Z SM (Figure 11):</p><p>rbspab_o_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_o_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_o_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_o_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes H+ band EMIC waves with rising tones were observed in bins of RXY and Z SM (Figure 12):</p><p>rbspab_hr_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_hr_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_hr_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_hr_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p><p>Percentage of minutes He+ band EMIC waves with rising tones were observed in bins of RXY and Z SM (Figure 13):</p><p>rbspab_her_rxyzsm_pcnt_midnight_20120907_to_20160701.csv, rbspab_her_rxyzsm_pcnt_dawn_20120907_to_20160701.csv, rbspab_her_rxyzsm_pcnt_noon_20120907_to_20160701.csv, rbspab_her_rxyzsm_pcnt_dusk_20120907_to_20160701.csv </p>
Occurrence cubes for non-native taxa in Belgium and Europe
<p>This package contains aggregated occurrence data ("occurrence cubes") for non-native taxa in Belgium and Europe. These occurrence cubes were generated by grouping species occurrence data from the <a href="https://www.gbif.org/">Global Biodiversity Information Facility (GBIF)</a> by year (year), 1x1km spatial <a href="https://www.eea.europa.eu/en/datahub/datahubitem-view/3c362237-daa4-45e2-8c16-aaadfb1a003b">EEA reference grid</a> cell (eea_cell_code) and taxon (taxonKey or classKey). For each grouping, the number of occurrences found in GBIF (n) and the minimum <a href="http://rs.tdwg.org/dwc/terms/coordinateUncertaintyInMeters">coordinateUncertaintyInMeters</a> (min_coord_uncertainty) are provided. The provided coordinateUncertaintyInMeters of an occurrence is taken into account when assigning it to a grid cell (see <a href="https://github.com/trias-project/occ-cube-alien/blob/20201201/src/europe/2_assign_grid.Rmd#L463-L481">this code</a>). The occurrence cubes have been used as input data for indicators and risk modelling/mapping for the <a href="http://trias-project.be/">Tracking Invasive Alien Species (TrIAS)</a> project and are now used for monitoring the effectiveness of the early detection and rapid eradication of emerging Invasive Alien Species (IAS) for the <a href="https://www.riparias.be/">LIFE RIPARIAS</a> project.</p> <p>The occurrence cubes are built on open science principles and intended to be completely reproducible:</p> <ul> <li>The input data are publicly available on GBIF, with the download DOIs listed in the related identifiers of this package.</li> <li>The code to process the data to cubes is publicly available on GitHub at <a href="https://github.com/trias-project/occ-cube-alien">https://github.com/trias-project/occ-cube-alien</a> (version <a href="https://github.com/trias-project/occ-cube-alien/releases/tag/20240118">20240118</a>).</li> </ul> <h2>Files</h2> <ul> <li><strong>be_alientaxa_cube.csv</strong>: occurrence cube of alien taxa listed by the Global Register of Introduced and Invasive Species - Belgium (Desmet et al. 2019) (GRIIS) and limited to occurrences in Belgium (country=BE).</li> <li><strong>be_alientaxa_info.csv</strong>: taxonomic information for taxa in be_alientaxa_cube.csv.</li> <li><strong>be_classes_cube.csv</strong>: occurrence cube of all <a href="http://rs.tdwg.org/dwc/terms/class">classes</a> found in Belgium (country=BE), used to assess sampling effort bias in be_alientaxa_cube.csv.</li> <li><strong>eu_modellingtaxa_cube.csv</strong>: occurrence cube of <a href="https://github.com/trias-project/occ-cube-alien/blob/2ada0ded33c034946380b02a28cb9a8d2884d54a/references/modelling_species.tsv">selected modelling species</a> in Europe (bounding box).</li> <li><strong>eu_modellingtaxa_info.csv</strong>: taxonomic information for taxa in eu_modellingtaxa_cube.csv.</li> </ul> <h2>Acknowledgements</h2> <p>This work has been funded under the Belgian Science Policies Brain program (BelSPO BR/165/A1/TrIAS), the European Union's LIFE program (LIFE19 NAT/BE/000953 - LIFE RIPARIAS) and the European Union's Horizon Europe Research and Innovation Programme (ID No 101059592 - Biodiversity Building Blocks for Policy).</p>
Occurrence Record Dataset from "Annotated checklist of the bees of Bonaire, with a focus on host plants"
<p>This is the occurrence dataset created for the publication "Annotated checklist of the bees of Bonaire, with a focus on host plants" (<a href="https://natuurtijdschriften.nl/pub/1026875" target="_blank" rel="noopener">https://natuurtijdschriften.nl/pub/1026875</a>).</p> <p>Observation and specimen data were assembled for this dataset, with the majority of records obtained during the Bonaire Estafette Expeditie (BEE). All citizen science records from Observation.org and iNaturalist.org up to December 2023 have been critically reviewed.<br>A project was created (<a href="https://www.inaturalist.org/projects/flower-visitors-and-pollinators-of-the-caribbean" target="_blank" rel="noopener">Flower visitors and pollinators of the Caribbean</a>) to improve standardized data collecting of plant-pollinator interactions and on <a href="https://observation.org/">observation.org</a> the standardized fields for interactions were used.<br>Records from passive trapping methods are not included. All bees were either observed or collected by hand or insect net. The majority of specimens will be accessible in the collection of Naturalis Biodiversity Center (RMNH), Leiden (the Netherlands). A synoptic collection is retained at the University of Tartu Zoological Collections in Tartu, Estonia (TUZ).</p> <p>The occurrence dataset (Version 1.4 and later) is:</p> <ul> <li>conform Darwin Core (DwC): <a href="https://dwc.tdwg.org/terms/">https://dwc.tdwg.org/terms</a></li> <li>in the data format CSV (tab delimited values) and UTF-8 encoded</li> </ul> <p> </p> <p><strong>DwC terms (Column labels) used in the dataset with their description:</strong></p> <table> <tbody> <tr> <td><strong>Column label</strong></td> <td><strong>Column description</strong></td> </tr> <tr> <td>occurrenceID</td> <td>Unique identifier or URI (GUID) for each record, mainly unique URLs generated by the web-based data holder.</td> </tr> <tr> <td>catalogNumber</td> <td>Unique code derived from URI in occurrenceID. Each specimen bears a label with this identifier and multimedia are tagged with this identifier.</td> </tr> <tr> <td>recordNumber</td> <td>Sample field ID used to manage data of preserved specimen occurrence records.</td> </tr> <tr> <td>otherCatalogNumbers</td> <td>Other unique identifiers used on specimen labels, but not derived from an URI.</td> </tr> <tr> <td>scientificName</td> <td>The scientific name of the lowest taxonomic rank to which the individual(s) was identified.</td> </tr> <tr> <td>scientificNameAuthorship</td> <td>The author name and year of publication in accordance with ICZN rules.</td> </tr> <tr> <td>verbatimIdentification</td> <td>The original identification, including qualifiers if needed.</td> </tr> <tr> <td>individualCount</td> <td>The number of individuals present at the time of the occurrence.</td> </tr> <tr> <td>sex</td> <td>The sex of the individual(s). The values female, male or unknown are used, if a mixed group is observed multiple values are listed.</td> </tr> <tr> <td>lifeStage</td> <td>The life stage of the individual(s).</td> </tr> <tr> <td>basisOfRecord</td> <td>The specific nature of the data record at the time of the identification (e.g. PreservedSpecimen).</td> </tr> <tr> <td>identifiedBy</td> <td>The name of the person who made the identification in the field or based on collected evidence (e.g. specimen or photo).</td> </tr> <tr> <td>identificationQualifier</td> <td>In case the identification could be given only to a species group 'cf.' is recorded.</td> </tr> <tr> <td>dateIdentified</td> <td>The year when the identification was made.</td> </tr> <tr> <td>previousIdentifications</td> <td>The scientific name originally given to the observed or collected individual(s).</td> </tr> <tr> <td>order</td> <td>The name of the order (e.g. Hymenoptera).</td> </tr> <tr> <td>family</td> <td>The name of the family (e.g. Apidae).</td> </tr> <tr> <td>genus</td> <td>The name of the genus (e.g. Apis).</td> </tr> <tr> <td>subgenus</td> <td>The name of the subgenus (e.g. Apis).</td> </tr> <tr> <td>specificEpithet</td> <td>The name of the species, epithet as given in dwc:scientificName.</td> </tr> <tr> <td>taxonRank</td> <td>The taxonomic rank of the most specific name in dwc:scientificName.</td> </tr> <tr> <td>eventDate</td> <td>The date-time when the event was observed and recorded. The event date uses the ISO 8601-1:2019 standard, with the following formatting being used: format YYYY-MM-DD, or YYYY if only the year is known. If time of capture is known, then format is YYYY-MM-DDTHH:MM, with HH:MM the local time.</td> </tr> <tr> <td>year</td> <td>The year in which the event was observed and recorded.</td> </tr> <tr> <td>month</td> <td>The month in which the event was observed and recorded.</td> </tr> <tr> <td>day</td> <td>The day in which the event was observed and recorded.</td> </tr> <tr> <td>eventTime</td> <td>The time or interval during which the event occurred.</td> </tr> <tr> <td>samplingProtocol</td> <td>The name or description of the collecting or recording method used.</td> </tr> <tr> <td>behavior</td> <td>A description of the behavior shown by the individual(s) recorded in this occurrence.</td> </tr> <tr> <td>decimalLatitude</td> <td>The geographic latitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>decimalLongitude</td> <td>The geographic longitude in decimal degrees recorded by a GPS device (WGS84) when observing and recording the occurrence.</td> </tr> <tr> <td>geodeticDatum</td> <td>The ellipsoid, geodetic datum, or spatial reference system (SRS) upon which the geographic coordinates given in dwc:decimalLatitude and dwc:decimalLongitude is based.</td> </tr> <tr> <td>verbatimLocality</td> <td>The original textual description of the place.</td> </tr> <tr> <td>island</td> <td>The name of the island.</td> </tr> <tr> <td>countryCode</td> <td>The standard ISO 3166-1 alpha-2 country code for the country.</td> </tr> <tr> <td>coordinateUncertaintyInMeters</td> <td> <p>The horizontal distance (in meters) from the given dwc:decimalLatitude and dwc:decimalLongitude describing the smallest circle containing the actual location, usually the EPE (Estimated Position Error) from the GPS device. The EPE is here measured as the horizontal position error in meters.</p> </td> </tr> <tr> <td>recordedBy</td> <td>A person, group, or organization observing and recording the occurrence.</td> </tr> <tr> <td>associatedTaxa</td> <td>The type of association and the scientific name of the host taxon is recorded that is associated/has relationship with the taxon in dwc:scientificName. The association/relationship is recorded using the format as in the following example: "floral host":"Lantana sp."</td> </tr> <tr> <td>occurrenceRemarks</td> <td>Comments or notes about the dwc:Occurrence.</td> </tr> <tr> <td>associatedSequences</td> <td>A list (concatenated and separated) of identifiers (publication, global unique identifier, URI) of genetic sequence information.</td> </tr> <tr> <td>typeStatus</td> <td>A list (concatenated and separated) of nomenclatural types (type status, typified scientific name, publication) applied to the subject.</td> </tr> <tr> <td>collectionCode</td> <td>The name, acronym, coden, or initialism identifying the collection or data set from which the record was derived.</td> </tr> <tr> <td>identificationRemarks</td> <td>Comments or notes about the identification.</td> </tr> <tr> <td>identificationReferences</td> <td>A reference or list of references (publication, global unique identifier, URI) used for the identification.</td> </tr> <tr> <td>nameAccordingTo</td> <td>A reference to the checklist or publication that was followed to record the name in dwc:scientificName.</td> </tr> <tr> <td>samplingEffort</td> <td>The amount of effort, expressed in minutes or hours, to obtain and record the occurrences.</td> </tr> <tr> <td>occurrenceStatus</td> <td>A statement about the presence or absence of a taxon during the time of an event.</td> </tr> <tr> <td>disposition</td> <td>The current state of a specimen with respect to a collection.</td> </tr> <tr> <td>language</td> <td>The language of the record using ISO 639-1 codes, e.g. en</td> </tr> </tbody> </table>
Software and suspect database for: "A large scale multi-laboratory suspect screening of pesticide metabolites in human biomonitoring: From tentative annotations to verified occurrences"
<p>This upload contains the pesticide suspect list aggregated among the laboratories of work package 16 of the HBM4EU (https://www.hbm4eu.eu) project for a large-scale pesticide suspect screening and the resolving search templates for each pesticide. Additionally, we provide the used software version of MetAlign applied in this screening.</p>
Spatiotemporal data for spotted lanternfly occurrence in the US
<p>An aggregated data set containing anonymized, spatiotemporal occurrence data for the spotted lanternfly (<em>Lycorma delicatula</em>, White 1845) in the United States. More details on the data, and additional tools to visualize it, can be found here: https://github.com/ieco-lab/lydemapr, and in the related publication.</p>
Predicted occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution
<p>The dataset contains predictions of occurrence probability for ticks in Great Britain (2014 to 2021) at 1 km spatial resolution + all covariate layers used for modeling. Over seven million electronic health records (EHRs), among which 11,741 EHRs reported tick attachment, were used to evaluate climate, environmental and animal host factors affecting the risk of tick attachment in cats and dogs in Great Britain (GB). The tick presence/absence EHRs for dogs and cats were further overlaid with spatiotemporal time-series of climatic, vegetation, human influence, hydrological and terrain variables (slope, wetness index) to produce a spatiotemporal regression matrix; an Ensemble Machine Learning framework was used to fine-tune hyperparameters for Random Forest (classif.ranger), Gradient boosting (classif.xgboost) and GLM-net (classif.glmnet) algorithms, which were then used to produce a final ensemble meta-learner that predicts the probability of occurrence of ticks across GB with monthly intervals.</p> <ul> <li>gb1km_covariates.zip contains ALL covariate layers as GeoTIFFs (time-series) used for modeling ticks dynamics;</li> <li>data_1km_2014_M01.rds = contains all covariates for January 2014 prepared as SpatialGridDataFrame (R data object);</li> </ul> <p>Codes of files indicate e.g.:</p> <ul> <li>"monthly.tick.prob_savsnet.mar_p_1km_s_2014_2021" = monthly occurrence probability for January based on the training data from 2014 to 2021;</li> <li>"monthly.tick.prob_savsnet.oct_md_1km_s_20211001_20211031" = monthly prediction (model) error derived as the standard deviation from multiple base learners;</li> </ul> <p>The dataset is described in detail in the following publication:</p> <ul> <li>Arsevska, E., Hengl, T., Singelton, D. et al. (2023?) <strong>Risk factors for tick attachment in companion animals in Great Britain: a spatiotemporal analysis covering 2014–2021</strong>. Submitted to Parasites & Vectors (in review).</li> </ul> <p>The model summary shows:</p> <pre><code>Call: stats::glm(formula = f, family = "binomial", data = getTaskData(.task, .subset), weights = .weights, model = FALSE) Deviance Residuals: Min 1Q Median 3Q Max -1.4749 -0.0557 -0.0471 -0.0430 3.7611 Coefficients: Estimate Std. Error z value Pr(>|z|) (Intercept) -7.64495 0.02095 -364.957 < 2e-16 *** classif.ranger 4.95061 0.63615 7.782 7.13e-15 *** classif.xgboost 189.75543 5.53109 34.307 < 2e-16 *** classif.glmnet 140.24208 5.05375 27.750 < 2e-16 *** --- Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1 (Dispersion parameter for binomial family taken to be 1) Null deviance: 170604 on 7303013 degrees of freedom Residual deviance: 162571 on 7303010 degrees of freedom AIC: 162579 Number of Fisher Scoring iterations: 9</code></pre> <p><em>Acknowledgements</em>: We are grateful to data providers in veterinary practice (VetSolutions, Teleos, CVS, and other practitioners). We are grateful to the INRAE MIGALE bioinformatics facility (MIGALE, INRAE, 2020. Migale Bioinformatics Facility, doi: <a href="https://entrepot.recherche.data.gouv.fr/dataverse/migale">10.15454/1.5572390655343293E12</a>) for providing computing resources. We are also grateful for<br> the help and support provided by <a href="https://www.liverpool.ac.uk/savsnet/">SAVSNET team members</a> Bethaney Brant, Susan Bolan and Steven Smyth.<br> This study was funded mainly by a grant from the <strong>Biotechnology and Biological Sciences Research Council</strong>,<br> BB/NO19547/1 and <strong>British Small Animal Veterinary Association</strong> (BSAVA). The research was partly funded by the National Institute for <strong>Health Research Health Protection Research Unit</strong> (NIHR HPRU) in Emerging and Zoonotic Infections at the <strong>University of Liverpool</strong> in partnership with <strong>Public Health England</strong> (PHE) and <strong>Liverpool School of Tropical Medicine</strong> (LSTM). This work has been partially funded by the <em>“Monitoring outbreak events for disease surveillance in a data science context"</em> (MOOD) project from the European Union’s Horizon 2020 research and innovation program under grant agreement No. 874850 (<a href="https://mood-h2020.eu/">https://mood-h2020.eu/</a>). The views expressed are those of the authors and not necessarily those of the NHS, the NIHR, the Department of Health or Public Health England.</p>
PIE LTER salt marsh vegetation frequency of occurrence from regularly monitored transects in Rowley, MA
Marsh vegetation presence-absence data along transects in salt marsh sites in Rowley, MA. The transects are intended to study long term changes in marsh vegetation. Four sites ((12 transects) were originally set up to study the impact of salt marsh haying. Two of these sites (labeled McH and EPH) were regularly hayed until 2002. The other two (PUH and CC) were reference sites. Two additional sites labeled RM and RR (8 transects) were originally set up to track invasion by Phragmites australis.
Data in support of 'ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension'
<p>Data in support of 'Chandler M, Sprintall J, Zilberman NV. (2025). ENSO influences subsurface marine heatwave occurrence in the Kuroshio Extension. <em>Journal of Geophysical Research: Oceans</em>. <a href="https://doi.org/10.1029/2025JC022899" target="_blank" rel="noopener">https://doi.org/10.1029/2025JC022899</a>'</p> <p> </p> <p>There are 2 netCDF files:</p> <ol> <li>p40tem1211_2312.nc</li> <li>synthetic_T_10day_px40_kuroshio_chandler2024.nc</li> </ol> <p><strong>p40tem1211_2312.nc </strong>contains the temperature sections from <a href="https://www-hrx.ucsd.edu/px40.html">HR-XBT transect PX40</a> objectively mapped onto a 10-m depth grid and a 0.1° longitudinal grid. <em>[LONGITUDE; LATITUDE; DEPTH; TIME; TEM]</em></p> <p><strong>synthetic_T_10day_px40_kuroshio_chandler2024.nc</strong> contains the synthetic temperature anomaly time series between the surface and 800-m deep at the western end of transect PX40 over the period from January-1993 to April-2023, as well as the temperature annual cycle needed for reconstructing the full synthetic temperature time series. <em>[time; depth; longitude; latitude; T_prime; T_ann]</em></p> <p> </p> <p>There is 1 MATLAB file:</p> <ol> <li>px40_synthetic_T.m</li> </ol> <p><strong>px40_synthetic_T.m</strong> is the MATLAB script used to produce the synthetic temperature anomaly time series saved in synthetic_T_10day_px40_kuroshio_chandler2024.nc.</p> <p> </p> <p>There is 1 Julia file:</p> <ol> <li>px40_synthetic_T_julia.jl</li> </ol> <p><strong>px40_synthetic_T_julia.jl</strong> is a Julia implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p>There is 1 R file:</p> <ol> <li>px40_synthetic_T_R.R</li> </ol> <p><strong>px40_synthetic_T_R.R</strong> is an R implementation of the MATLAB script px40_synthetic_T.m.</p> <p> </p> <p><code>Version history:</code><br><code>v1.0.0 First uploaded (25-November-2024)</code><br><code>v1.0.1 Julia script uploaded (18-January-2025)</code><br><code>v1.0.2 R script uploaded (28-January-2025)</code><br><code>v1.1.0 Updated description of synthetic_T_10day_px40_kuroshio_chandler2024.nc to include reference to accepted publication (21-August-2025)</code></p>
Occurrence cubes at species level for European countries
<p>This package contains aggregated occurrence data ("occurrence cubes") at species level for European countries. These occurrence cubes were generated by grouping species occurrence data from the <a href="https://www.gbif.org/">Global Biodiversity Information Facility (GBIF)</a> by year (year), 1x1km spatial <a href="https://www.eea.europa.eu/en/datahub/datahubitem-view/3c362237-daa4-45e2-8c16-aaadfb1a003b">EEA reference grid</a> cell (eea_cell_code) and species (speciesKey). For each grouping, the number of occurrences found in GBIF (n) and the minimum <a href="http://rs.tdwg.org/dwc/terms/coordinateUncertaintyInMeters">coordinateUncertaintyInMeters</a> (min_coord_uncertainty) are provided. The provided coordinateUncertaintyInMeters of an occurrence is taken into account when assigning it to a grid cell (see <a href="https://github.com/trias-project/occ-cube/blob/master/src/3_assign_grid.Rmd#L198-L234">this code</a>). The occurrence cubes can be used as input data for indicators, mapping and species distribution modelling.</p> <p>The occurrence cubes are built on open science principles and intended to be completely reproducible:</p> <ul> <li>The input data are publicly available on GBIF, with the download DOIs listed in the related identifiers of this package.</li> <li>The code to process the data to cubes is publicly available on GitHub at <a href="https://github.com/trias-project/occ-cube-alien">https://github.com/trias-project/occ-cube</a> (version <a href="https://github.com/trias-project/occ-cube/tree/20240124">20240124</a>).</li> </ul> <h2>Files</h2> <ul> <li><strong>Occurrence cubes at species level per country</strong>: filename format countrycode_species_cube.csv.</li> <li><strong>Taxonomic information for species in a cube</strong>: filename format countrycode_species_info.csv.</li> </ul> <h2>Included countries</h2> <ul> <li><strong>Belgium</strong> (BE): based on <a href="https://doi.org/10.15468/dl.9qx3ba">https://doi.org/10.15468/dl.9qx3ba</a></li> <li><strong>Italy</strong> (IT): based on <a href="https://doi.org/10.15468/dl.jghpm5">https://doi.org/10.15468/dl.jghpm5</a></li> <li><strong>Lithuania</strong> (LT): based on <a href="https://doi.org/10.15468/dl.duegx2">https://doi.org/10.15468/dl.duegx2</a></li> <li><strong>Slovenia</strong> (SI): based on <a href="https://doi.org/10.15468/dl.9eky98">https://doi.org/10.15468/dl.9eky98</a></li> <li><strong>Romania</strong> (RO): based on <a href="https://doi.org/10.15468/dl.b7z5vw">https://doi.org/10.15468/dl.b7z5vw</a></li> <li><strong>Portugal</strong> (PT): based on <a href="https://doi.org/10.15468/dl.b89nr4">https://doi.org/10.15468/dl.b89nr4</a></li> </ul> <p>Occurrence cubes are added on demand. To include occurrence cubes for other European countries, <a href="https://github.com/trias-project/occ-cube/issues">leave an issue</a> or contact the main author, or generate your own cube using the code in <a href="https://github.com/trias-project/occ-cube">this repository</a>.</p>
Alaska Lake and Pond Occurrence Dataset
<p>The Alaska Lake and Pond Occurrence Dataset (ALPOD) is a spatially explicit map of seasonal fluctuations in lake and pond extent across Alaska. ALPOD is constructed from Sentinel-2 imagery and its lake occurrence raster has a 10 m resolution, with each pixel recording the percentage of observations classified as open water during the 2016-2021 ice-free seasons at near-daily temporal resolution. The number of water bodies depends on the occurrence threshold, but a conservative estimate is that ALPOD maps over 800,000 lakes and ponds larger than 0.001 square km. The data format includes lake occurrence rasters tiled by UTM zone and latitude, as well as a vector product defined based on a 25% occurrence threshold. ALPOD was created using a U-Net lake identification model and manual inspection to produce a maximum possible lake extent mask. This mask serves as the region of interest for an adaptive NDWI threshold water classification algorithm written in Google Earth Engine, which was used to classify open water within the maximum lake mask in every available cloud and ice free Sentinel-2 image during the study period. The resulting dataset is suitable for investigations of individual water bodies as well as lake and pond patterns across Alaska.</p> <p>Geotiffs included here are 10 m resolution lake occurrence rasters with each pixel recording the percentage of observations classified as open water during the 2016-2021 ice-free seasons. They are tiled by UTM zone above 60 degrees north. Regions below 60 degrees north are separated into two geotiffs.</p>
Occurrences records of Herichthys labridens (Cichliformes: Cichlidae), with associated habitat information, in the Media Luna spring, San Luis Potosí, Mexico
<h2><strong>Introduction</strong></h2> <blockquote> <p>Occurrence records of the endemic cichlid <em>Herichthys labridens</em>, by adult and juvenile life stages, during three summer events (years of 1999, 2009, and 2019), in the Media Luna spring, San Luis Potosí Mexico. </p> </blockquote> <h2><strong>Material and Methods </strong></h2> <blockquote> <p>The occurrence records, ordered by adult and juvenile life stages, were obtained from two sources. For the summer of 1999, data were downloaded from the literature (Palacio-Núñez et al., 2010). For subsequent events, we recorded new data from 66 underwater transects distributed among 14 sectors (S1 to S14) in the Media Luna spring. We followed the method of Palacio-Núñez (2007), which maintained the transect location and sector boundaries of the summer of 1999 (Fig. 1a). The 20 m² transects were placed transversely to the current, from the edge to the central part of the canal (Fig. 1b). This sampling design was selected to meet two basic assumptions for studies of spatial distribution and habitat suitability: (1) the observations within the area are true and, (2) these observations delimit the initial position of the recorded individuals (Buckland & Elston, 1993). The analysis of the spatial information of the sectors, the underwater transects, and the delimitation of the water surface was performed using the QGIS® software version 3.4.8 (Menke, 2019).</p> <p> </p> <p><strong>Figure 1</strong>. <a href="https://zenodo.org/api/records/14231104/draft/files/Sector%20boundary_Transect%20location%20and%20sampling_Media%20Luna%20spring.jpeg/content" target="_blank" rel="noopener noreferrer">Sector boundary_Transect location and sampling_Media Luna spring.jpeg</a>. (a) Location of the transects in the Media Luna spring, Mexico. (b) Design scheme of the sampling transect; a CPVC pipe was used to give width to the edges of the transect and a nylon rope was attached to each side of the pipes to demarcate the length of the transect. Floating rubber buoys were added to the transects (at the edge towards the center of the canal) to prevent them from sinking into the sediment and to locate them among the vegetation. Transect scheme: Jorge Palacio-Núñez.</p> <p><br>In the summer events where we worked in field, we recorded the spatial location (i.e., GPS coordinates) of each individual and its life stage by direct observation with snorkel equipment and using a Garmin etrex device. The recorded information included the data of water depth and related underwater coverage. It is important to mention that, to prevent a repeat observation of the same organism or to ommit any individual, the transect was swaped slowly and in one direction only (i.e., from the center of the canal to the shore). We also used underwater cameras to validate the information. In adittion, the characterization of <em>H. labridens</em> individuals by life stage was performed by approximate size. For this purpose, previous studies on the life history and biology of the species were reviewed (Miller et al., 2005; De La Maza-Benignos & Lozano-Vilano, 2013). It is worth mentioning that, during fieldwork, we avoided manipulation, damage, or unnecessary capture of the fish (e.g., Prchalová et al., 2009).</p> <p><br>The databases by life stage were organized for each summer event, where, each observation record was included along with the associated habitat conditions. Subsequently, we depurated each database to remove atypical spatial data, data without information, incomplete data, or data with duplicate coordinates (García-Roselló et al., 2014). Then, we performed spatial filtering of the remaining records to validate those that were within the study area, and to prevent that two or more points were within 0.1 m of each other. These steps of our analysis were performed using the software Qgis® version 3.28.4 and Rstudio® (Rstudio team, 2020). Subsequently, with the data set that included fish records, water depth, and underwater coverage variables, we performed a final environmental filter to rule out atypical records. This exploration was performed in Rstudio ® using the outliers function, starting from the lowest and highest quantiles.</p> </blockquote> <h2><strong>Results</strong></h2> <blockquote> <p>The final filtered databases were organized by life stage and summer event:</p> <p><strong>Adult: </strong></p> <table> <tbody> <tr> <td>Summer event</td> <td>Database</td> </tr> <tr> <td>1999</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_1999_Palacio-N%C3%BA%C3%B1ez%20et%20al.,%202010.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_1999_Palacio-Núñez et al., 2010.csv</a></td> </tr> <tr> <td>2009</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_2009_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_2009_Field work.csv</a></td> </tr> <tr> <td>2019</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Adult_2019_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Adult_2019_Field work.csv</a></td> </tr> </tbody> </table> <p><strong> Juvenile:</strong></p> <table> <tbody> <tr> <td>Summer event</td> <td>Database</td> </tr> <tr> <td>1999</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_1999_Palacio-N%C3%BA%C3%B1ez%20et%20al.,%202010.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_1999_Palacio-Núñez et al., 2010.csv</a></td> </tr> <tr> <td>2009</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_2009_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_2009_Field work.csv</a></td> </tr> <tr> <td>2019</td> <td><a href="https://zenodo.org/api/records/14231104/draft/files/Occurrences_records_H_labridens_Juvenile_2019_Field%20work.csv/content" target="_blank" rel="noopener noreferrer">Occurrences_records_H_labridens_Juvenile_2019_Field work.csv</a></td> </tr> </tbody> </table> </blockquote> <p> </p> <blockquote> <p>These occurrence records for <em>H. labridens </em>are ready to be used in ecological niche modeling and spatial distribution studies. Also, these records can be used for other ecological and spatial studies, because each record (i.e., individual) included geoespatial coordinates, sector, location, and transect number. Also, we recorded information about the conditions of underwater coverage and water depth, which were asociated to each ocurrence record.</p> <p>For more information about several R codes where the previous databases can be used, visit the following repository URL: <a href="https://doi.org/10.5281/zenodo.7603557">https://doi.org/10.5281/zenodo.7603557</a>.</p> <p>Also, to download the UC and WDp variables to run the spatial and ecological modeling, visit the following repository URL: <a href="https://doi.org/10.5281/zenodo.7603890">https://doi.org/10.5281/zenodo.7603890</a>.</p> </blockquote>
Global GFED-based monthly burned area time series (1996-2016) at 1 km and ESA CCI MODIS-based long-term monthly P90 burned area occurrence at 500 m
<p>Contains two separate datasets:</p> <ol> <li>Global <a href="https://www.globalfiredata.org/data.html">GFED-based monthly burned area</a> (in ha) <a href="https://youtu.be/kBJcP8mL2Qs">time series (1996-2016)</a> at 1 km (downscaled using cubic-splines from 25 km);</li> <li>Global burned area long term (2000-2012) P90 (quantile probability = 0.9) based on the <a href="http://maps.elie.ucl.ac.be/CCI/viewer/index.php">ESA CCI burned area accumulated weekly product</a>;</li> </ol> <p>Original GFED monthly data is provided as HDF4 files (ftp.fuoco.geog.umd.edu/data/GFED/GFED4). Dataset is described in detail in <a href="https://doi.org/10.1002/jgrg.20042">Giglio et al. (2013)</a>. Processing steps are available <a href="https://gitlab.com/openlandmap/global-layers/tree/master/input_layers/GFED"><strong>here</strong></a>. Antarctica is not included.</p> <p>To access and visualize global datasets use: <a href="https://openlandmap.org"><strong>https://openlandmap.org</strong></a> or watch <a href="https://youtu.be/kBJcP8mL2Qs"><strong>this video</strong></a>.</p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> </ul> <p>All files provided as Cloud-Optimized GeoTIFFs / internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>nhz = theme: natural hazards,</li> <li>monthly.burned.ha = variable: estimated monthly burned area in ha,</li> <li>gfed = data source GFED data,</li> <li>m = mean value,</li> <li>1km = spatial resolution / block support: 1 km,</li> <li>s0..0cm = vertical reference: land surface,</li> <li>2000.02 = time reference aggregated: month Feb of year 2000,</li> <li>v4 = version number: GFEDv4,</li> </ul>
Data for the paper 'Reducing networks of ethnographic codes co-occurrence in anthropology'
<p>Pseudonymized data supporting the paper "Reducing networks of ethnographic codes co-occurrence in anthropology", published in "Advances in Quantitative Ethnography. Fourth International Conference on Quantitative Ethnography (ICQE 2022), Copenhagen, Denmark, October 15–19, 2022, Proceedings", and edited by Amanda Barany and Crina Damsa. The paper is part of the POPREBEL project. The data were gathered in the spring and summer of 2021, as a part of a larger research project on populism in Central and Eastern Europe, to be completed by the end of 2022. They consist of 17 semi-structured interviews with Polish-speaking Internet users, who used social media to seek and share information about health against the backdrop of the COVID-19 pandemic. Research participants were asked about their opinion on the current state of affairs in their respective countries, and their political choices over the years and at present.</p> <p><a href="https://edgeryders.eu/t/long-term-ssna-data-storage-documentation-manual/12786">Data export and documentation process</a> (contains links to the code used to export the data).</p>
Grasshopper species occurrence data in Mt. Kilimanjaro
<p>Species occurence data of grasshopper in 60 plots on the southern slope of Mt. Kilimanjaro.</p> <p>Orthoptera assemblages were recorded on all study sites by repeatedly walking for 1.5 h on parallel tracks (distance between transects ca. 1-1.5 m) and recording all sighted species. In forested study sites, trees and bushes in the understory vegetation were shaken for approximately 1.5 h. Insects falling from the vegetation were gathered on white canvas laid on the forest floor. Species which could not be identified during visits were collected and later identified. Study sites were also visited at night where Ensifera were registered acoustically. Additionally, two rounds of sweep net sampling were conducted on study sites to collect small species which may have remained undetected during transect walks. One round was conducted during the cool dry season (July to October) and one during the warm dry season (December to March). During each sweep netting round, 100 sweeps with a 30-cm diameter sweep were taken and all collected specimens were identified in the laboratory. Species accumulation curves for Caelifera and Ensifera on Mt. Kilimanjaro were published in , showing that more than 90% of the grasshopper, locust and bushcricket fauna for Mt. Kilimanjaro have been registered.</p> <p>The KiLi project (2010-2018) is a German Science Foundation (DFG) funded research unit (DFG research unit FOR1246) that focuses on biodiversity and ecosystem processes along altitudinal and disturbance gradients on Mt. Kilimanjaro (Tanzania, Africa), capitalizing on its world-wide unique range of climatic and vegetation zones. The research unit comprises 2 central projects and 7 subprojects from various disciplines. On a total of 60 study sites in both natural and human-disturbed ecosystems biodiversity (e.g. plants, soil arthropods, ants, bees, frogs, lizards, bats, birds), related ecosystem processes (decomposition, seed dispersal, pollination, herbivory, predation), and biogeochemical processes and properties of ecosystems (climate, soil properties and nutrient status, regulation of water and carbon fluxes, trace gas emissions, primary productivity, functional diversity) are analyzed.</p>
Landsat bands (cloud free), tree cover (2000, 2010), bare-ground and surface water occurrence at 250 m based on GlobalForestWatch and USGS
<p>Landsat bands (cloud free) and tree cover (2000) based on Hansen et al. (2013), global surface water occurrence based on Pekel at al. (2016), and tree cover and bare-ground cover (2010) based the USGS land cover mapping projects (University of Maryland, Department of Geographical Sciences and USGS). All layers resampled to spatial resolution 1/480 d.d. (about 250 m) using gdalwarp with "average" resampling. Antarctica is not included. Original layers are available at 30 m resolution.</p> <p>If you discover a bug, artifact or inconsistency in the maps, or if you have a question please use some of the following channels:</p> <ul> <li>Technical issues and questions about the code: <a href="https://gitlab.com/openlandmap/global-layers/issues">https://gitlab.com/openlandmap/global-layers/issues</a> </li> <li>General questions and comments: <a href="https://disqus.com/home/forums/landgis/">https://disqus.com/home/forums/landgis/</a></li> </ul> <p>All files internally compressed using "COMPRESS=DEFLATE" creation option in GDAL. File naming convention:</p> <ul> <li>lcv = theme: land cover,</li> <li>bareground = variable: occurrence of bareground,</li> <li>landsat.usgs = determination method: Landsat landcover at 30 m resolution project (https://landcover.usgs.gov/glc/),</li> <li>p = probability or fraction,</li> <li>250m = spatial resolution / block support: 250 m,</li> <li>s0..0cm = vertical reference: land surface,</li> <li>2010..2010 = time reference: year 2010,</li> <li>v1.0 = version number: 1.0,</li> </ul>
ZooBase: A global synthesis of marine zooplankton species occurrences.
<p><em><strong>Description of the methods used to implement the present ZooBase dataset (extarct from Section A.2 from the Methods of Benedetti et al., 2021).</strong></em></p> <p>A new dataset of global zooplankton species occurrences was compiled in a comparable fashion to that put together for phytoplankton (Righetti et al., 2020). Prior to retrieving the occurrence data online, we first identified the phyla (Order/Class/Family) that comprise the bulk of extant oceanic zooplankton communities: Copelata (i.e. appendicularians), Ctenophora, Cubozoa (i.e. box jellyfish), Euphausiidae (i.e. krill), Foraminifera, Gymnosomata (i.e. sea angels, pteropods), Hydrozoa (i.e. jellyfish), Hyperiidea (i.e. amphipods), Myodocopina (i.e. ostracods), Mysidae (i.e. small pelagic shrimps resembling krill), Neocopepoda, Podonidae and <em>Penilia</em> <em>avirostris</em> (i.e. cladocerans), Sagittoidea (i.e. chaetognaths), Scyphozoa (i.e. jellyfish), Thaliacea (i.e. salps, doliolids and pyrosomes), Thecosomata (i.e. pteropods), and four families of pelagic Polychaeta (i.e. worms) that are often found in the zooplankton and whose species are known to display holoplanktonic lifecycles (Tomopteridae, Alciopidae, Lopadorrhynchidae, Typhloscolecidae). The presence data associated with species belonging to these groups were retrieved from OBIS and GBIF between the 12/04/2018 and the 18/04/2018 using online queries via the R packages RPostgreSQL, robis and rgbif. Since the Neocopepoda infra-class comprise several thousands of benthic and parasitic taxa (https://copepodes.obs-banyuls.fr/en/), a preliminary selection of the non-parasitic planktonic species had to be carried out prior to the online downloading using the species list of Razouls et al. (https://copepodes.obs-banyuls.fr/en/) as a reference. The spatial distributions of the groups cited above were first inspected using GBIF’s and OBIS’s online mapping tools to evaluate the potential number of overlapping observations between the two databases. As a result of their relatively low contributions to total observations/diversity, and very high overlap between databases, the occurrences of Cladocera and Polychaeta were retrieved from OBIS only (which usually harbours more occurrences). On top of the data collected from OBIS and GBIF, the copepod occurrences from Cornils et al. (2018) and the pteropod occurrences from the MAREDAT initiative (Buitenhuis et al., 2013) were added to the dataset. We discarded records that: (i) presented at least one missing spatial coordinate, (ii) were associated with an incomplete sampling date (d/m/y), (iii) were associated with a year of collection older than 1800, (iv) were not associated with any sampling depth, (v) were not identified down to the species level. Occurrences associated with grid cells shallower than 10m were removed (bathymetry data from ETOPOv1 at a 15min resolution, downloaded using the 'marmap' R package). Finally, every species name was then carefully examined and compared to the taxonomic reference list of the World Register of Marine Species (WoRMS; <a href="http://www.marinespecies.org">http://www.marinespecies.org</a>) for all taxa. The AphiaID and the Status were retreived from WoRMS based on the ScientificName. To remove the duplicate occurrences due to the highly overlapping source archives (GBIF and OBIS), a unique occurrenceID was given to each record based on rounded spatial coordinates (closest 0.1°x0.1°), rounded depth layer (10m depth layers), month and year of the occurrence and the acccpted species name (e.g., AphiaID).</p> <p><strong>This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 862923. This output reflects only the author’s view, and the European Union cannot be held responsible for any use that may be made of the information contained therein.</strong></p>
Harmonized Tree Species Occurrence Points for Europe
<p>This data set is a harmonized collection of existing data from GBIF, the EU-Forest project and the LUCAS survey. It has about 3 million observations and is supplemented by variables (e.g. location accuracy, land cover type, canopy height, etc.) which enable precise filtering for specific user applications.</p> <p>The <em>RDS </em>file is created from an sf-object and suitable for fast reading in the R-programming environment. The <em>CSV.GZ</em> file contains records as a table with Easting and Northing in Coordinate Reference System ETRS89 / LAEA Europe (= EPSG code 3035) and can be fed in a GIS after being unzipped.</p> <p><strong>The code producing this data set is <a href="https://gitlab.com/openlandmap/eu-forest-tree-point-data">publicly available on GitLab.</a></strong></p> <p>Data sets were last updated in September 2021.</p> <p>Variables:</p> <ul> <li><strong>id</strong> = unique point identifier</li> <li><strong>easting</strong> = x coordinate</li> <li><strong>northing </strong>= y coordinate</li> <li><strong>country </strong>= ISO country code</li> <li><strong>species </strong>= Latin species name</li> <li><strong>genus </strong>= genus name</li> <li><strong>scientific_name </strong>= long species name</li> <li><strong>gbif_taxon_key </strong>= taxon key from GBIF</li> <li><strong>gbif_genus_key </strong>= genus key from GBIF</li> <li><strong>taxon_rank </strong>= species or genus</li> <li><strong>year </strong>= year of observation</li> <li><strong>accessed_through </strong>= database through which data was accessed (GBIF, LUCAS, EU-Forest)</li> <li><strong>dataset_info </strong>= data set name (individual sub-data-set)</li> <li><strong>citation </strong>= DOI citation of the individual data set</li> <li><strong>license </strong>= distribution license</li> <li><strong>location_accuracy </strong>= spatial accuracy of observation (meters)</li> <li><strong>flag_location_issue </strong>= known location issues present</li> <li><strong>flag_date_issue </strong>= known date issues present</li> <li><strong>eoo </strong>= Extent of occurrence (applying the concept of natural geographical range used for the EU-Forest data set (<a href="https://www.nature.com/articles/sdata2016123">Mauri et al., 2017</a>) to all other data points. 1 = point inside species range; 0 = point outside; NA = EOO polygon not available for this species)</li> <li><strong>dbh </strong>= Diameter Breast Height (only recorded for observations from the EU-Forest data set (<a href="https://www.nature.com/articles/sdata2016123">Mauri et al., 2017</a>))</li> <li><strong>lc1 </strong>= <a href="https://ec.europa.eu/eurostat/web/lucas/data/primary-data/2018">LUCAS</a> land cover type 1 (only recorded for observations from LUCAS data)</li> <li><strong>lc2 </strong>= <a href="https://ec.europa.eu/eurostat/web/lucas/data/primary-data/2018">LUCAS </a>land cover type 2 (only recorded for observations from LUCAS data)</li> <li><strong>landmask_country </strong>= land mask overlay 30 meters (NA = not on land)</li> <li><strong>corine </strong>= <a href="https://land.copernicus.eu/pan-european/corine-land-cover/clc2018">CORINE 2018</a> land cover type (extracted from the 100 meter raster data set)</li> <li><strong>nightlights</strong> = <a href="https://eogdata.mines.edu/download_dnb_composites.html">light pollution</a> observed by VIIRS (proxy for remoteness / distance to human structures)</li> <li><strong>canopy_height </strong>= <a href="https://zenodo.org/record/4057883#.X3sx4-2xW9I">canopy height</a> derived from GEDI waveform LiDAR point data</li> <li><strong>natura_2000 </strong>= Natura 2000 site code (if a point falls inside a protected area (<a href="https://www.eea.europa.eu/data-and-maps/data/natura-11">GIS-layer</a>) this variable contains the site identification code; all sites can be explored on an <a href="https://natura2000.eea.europa.eu/">interactive map</a>)</li> <li><strong>freq_location </strong>= number of points with identical location (in some cases one location has multiple observation, differing in species and/or year. This may lead to difficulties in certain modeling tasks)</li> <li><strong>geometry </strong>= point geometry in ETRS89 / LAEA Europe</li> </ul> <p>See <strong><a href="https://docs.google.com/spreadsheets/d/1WM0BIaVEKxTsCISEaF76RJ8F1iWiZlDfuyQCBiF2Sxw/edit?usp=sharing">this detailed documentation</a></strong> for more insights into each variable and individual GBIF data set citations.</p> <p>If you would like to know more about the creation of this data set, see</p> <ol> <li>the R-Markdown documenting the process (<a href="https://gitlab.com/openlandmap/eu-forest-tree-point-data">GitLab repository</a>)</li> <li>the talk at OpenGeoHub Summer School 2020 (<a href="https://www.youtube.com/watch?v=5HhmLGcqXLs&list=PLXUoTpMa_9s0Ea--KTV1OEvgxg-AMEOGv&index=40">Youtube</a>)</li> </ol> <p>Some advice: This data set is a puzzle with pieces from many different sources. Take some time to explore before including it in your work. Use summary statistics to see which variables have NAs and how many. Choose your filtering criteria wisely. For example, some points with the highest location accuracy have no record for the year of observations. You would exclude these, if "year > 1990" was your criteria.</p> <p> </p> <p>This work has received funding from the European Union's the Innovation and Networks Executive Agency (INEA) under Grant Agreement Connecting Europe Facility (CEF) Telecom project 2018-EU-IA-0095 (<a href="https://ec.europa.eu/inea/en/connecting-europe-facility/cef-telecom/2018-eu-ia-0095">https://ec.europa.eu/inea/en/connecting-europe-facility/cef-telecom/2018-eu-ia-0095</a>).</p>
Global taxonomic occurrence grids using GBIF data for species distribution models.
<p>To achieve large geographic coverage, species occurrence databases that are composed of ad hoc species data collections such as that provided by the Global Biodiversity Information Facility (GBIF) are often used. A drawback to using these data is their geographic sampling bias, in which some regions are more intensively sampled than others, while other areas have very little to none reported sampling effort. Uneven sampling effort can mislead conclusions about biodiversity patterns and species distributions (Gotelli & Colwell, 2001; Lobo, 2008).</p> <p>Here we provide taxonomic occurrence grids to help mitigate the effects of sampling bias in species distribution modeling. These grids can be used to exclude areas of (a custom-defined) low sampling effort from the background when sampling for pseudo-absences’ (Phillips et al., 2009; Barbet-Massin et al.,2012). The occurrence grids have a 1 degree spatial resolution using WGS 84 as the geographic coordinate system. Each 1 degree grid cell contains the number of records present in GBIF corresponding to a specific taxonomic group: plants, mammals, reptiles, amphibians, birds and molluscs.</p> <p>To construct the occurrence grids, we used the 1- by 1-degree world latitude and longitude vector grid provided by ESRI (Redlands, California). It has a custom license which permits it reuse as long as ESRI is cited. It was downloaded from : <a href="https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7">https://www.arcgis.com/home/item.html?id=f11bcdc5d484400fa926dcce68de3df7</a></p> <p>To map spatial sampling effort, the number of georeferenced occurrences corresponding to each taxonomic group contained by each 1- by 1-degree grid cell were counted. The grids were then converted to GeoTIFFs. The raster values correspond to the number of occurrences reported for the grid cells. For the purposes of the <a href="https://osf.io/7dpgr/">TrIAS project</a>, grid cells with fewer than 5 occurrences were removed. The TrIAS taxonomic occurrence grids are used as inputs to the TrIAS risk modelling and mapping workflow: https://github.com/trias-project/risk-modelling-and-mapping. Full (with all grid cells containing at least one occurrence) taxonomic occurrence grids are also provided.</p> <p>GBIF data for each taxonomic group were downloaded using the following criteria: “Basis of Record”: Observation, Machine Observation, Human Observation, Specimen, Material sample, Literature Occurrence, Unknown evidence., "HasCoordinate is true", "HasGeospatialIssue is false", "TaxonKey is Amphibia", "Year 1975-2005".</p> <p><strong>Raster Attributes</strong></p> <table> <tbody> <tr> <td> <p>Attribute</p> </td> <td> <p>Description</p> </td> </tr> <tr> <td> <p>OID</p> </td> <td> <p>numeric row ID</p> </td> </tr> <tr> <td> <p>Value</p> </td> <td> <p>the number of records contained in the grid cell</p> </td> </tr> <tr> <td> <p>Count</p> </td> <td> <p>the number of times the value appears in the raster</p> </td> </tr> </tbody> </table> <p> </p> <p> </p> <p>The extent of each taxonomic occurrence grid:</p> <ul> <li> <p>longitude -180.0; latitude -90.0 (southwest corner)</p> </li> <li> <p>longitude 180.0; latitude 90.0 (northeast corner)</p> </li> </ul> <p> </p> <p><strong>Files:</strong></p> <p>TrIAS taxonomic occurrence grids</p> <p>amphib_1deg_min5.tif</p> <p>birds_1deg_min5.tif</p> <p>mammals_1deg_min5.tif</p> <p>molluscs_1deg_min5.tif</p> <p>reptiles_1deg_min5.tif</p> <p> </p> <p>Raw taxonomic occurrence grids</p> <p>amphib_1deg_grid.tif</p> <p>birds_1deg_grid.tif</p> <p>mammals_1deg_grid.tif</p> <p>molluscs_1deg_grid.tif</p> <p>reptiles_1deg_grid.tif</p> <p><br> </p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.