Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3,015
datasets available to search
ShareScore release 0.7.1
Dataset results
3,015 results for “occurrence”
Model outputs for update of occurrence and hunting yield-based data models for wild boar at European scale: new approach to handle the bioregion effect, May 2020 update
<p>These maps are models obtained in intermediate phases of the ENETWILD project based on available information. There are frequent updates in order to improve the results.<br> <br> Objectives:<br> <br> - Incorporate additional data to provide new maps of wild boar suitability with a resolution of 2x2 km >>> file 3_June_2020_suitability_2x2.tif<br> - New model based on hunting yield with different approaches to handle the biorregion effect >>> files 1_June_2020_HY_nut01_10x10_twostep.tif & 2_June_2020_HY_nut01_10x10_pca.tif<br> <br> Model settings and predictors: <br> - Hunting yield modeling including biorregion effect as bioclimatic PCA scores<br> - Hunting yield addressing biorregion effect in a two-step procedure with independent parametrization for each bioregion<br> <br> Conclusions guiding future methodological steps:<br> - For wild boar suitability maps at 2x2 km, additional data on survey effort is critical in the southern bioregion<br> - Hunting yield model predictions at 10x10 km grids overestimated the hunting bag numbers obtained from the external datasets<br> - HY model with independent parametrization for each bioregion performed better that previous and new strategies<br> <br> For further details and methodological approach see the paper:<br> ENETWILD-consortium, P. Acevedo, S .Croft, G C Smith, J. A. Blanco-Aguiar, J. Fernandez-Lopez, M. Scandura, M. Apollonio, E.Ferroglio, Oliver Keuling, M. Sange, S. Zanet, F. Brivio, T. Podgórski, K.Petrović, Soriguer, J. Vicente (2020) update of occurrence and hunting yield-based data models for wild boar at European scale: new approach to handle the bioregion effect. EFSA supporting publication 2020 TO BE COMPLETED<br> <br> Permission for reuse hunting yield outputs is granted under the terms indicated by EFSA.</p>
Species occurrence and occupancy in protected areas of the Natura2000 network in Belgium
<p><strong>Context</strong></p> <p>Invasive alien species have been pointed out as an important driver of biodiversity loss. Many policy responses are being developed to address this threat. Protected areas often represent and preserve hotspots of biological diversity and ensure the maintenance of ecosystem services crucial to human livelihoods. The impact of biological invasions can be particularly severe in protected areas and their occurrence and impact in such areas is an important element of the risk they pose. To address this, there is a need for data on the occurrence and extent of alien species invasions in protected areas.</p> <p><strong>Description</strong></p> <p>This dataset contains species occurrence and occupancy in protected areas of the Natura2000 network in Belgium (Special Conservation Areas sensu Habitat Directive and Special Protection Areas sensu Bird Directive). The dataset was generated using the <a href="https://doi.org/10.5281/zenodo.3637911">Belgian occurrence cube at species level</a> and the <a href="https://doi.org/10.5281/zenodo.3635510">Belgian occurrence cube for non-native taxa</a> (both containing GBIF data aggregated using Oldoni et al. 2020), the 1x1km <a href="https://www.eea.europa.eu/data-and-maps/data/eea-reference-grids-2">EEA reference grid</a> and the <a href="https://www.eea.europa.eu/data-and-maps/data/natura-11/natura-2000-spatial-data/natura-2000-shapefile-1">Natura2000 protected areas shapefiles</a> from the European Environment Agency.</p> <p>Data are grouped by protected area (<code>SITECODE</code>), year (<code>year</code>) and (infra)species (<code>taxonKey</code>, <code>speciesKey</code>). For each group, it provides the number of occurrences found in GBIF (<code>n</code>), the area of occupancy (<code>aoo</code>: number of 1 km<sup>2</sup> squares), the coverage (<code>coverage</code>: % of 1 km<sup>2</sup> squares), the minimum <a href="http://rs.tdwg.org/dwc/terms/coordinateUncertaintyInMeters">coordinateUncertaintyInMeters</a> (<code>min_coord_uncertainty</code>), and the alien status (<code>is_alien</code>) based on the <a href="https://doi.org/10.15468/xoidmd">Global Register of Introduced and Invasive Species - Belgium</a>. For infraspecific taxa in the latter, the <a href="https://github.com/trias-project/indicators/blob/00e1ae72df3fb98b2a215c3af8769e53fbcd0182/reference/species_of_infraspecific_alien_taxa.tsv">alien status of the species</a> is looked up and included.</p> <p>The dataset is built on open science principles and intended to be completely reproducible:</p> <ul> <li>The input data are publicly available on Zenodo, with the download DOIs listed in the related identifiers of this dataset package.</li> <li>The <a href="https://trias-project.github.io/indicators/10_species_observations_occupancy_in_protected_areas.html">code</a> to process the data is publicly available and documented on GitHub.</li> </ul> <p><strong>Files</strong></p> <ul> <li><strong>protected_areas_species_occurrence.csv</strong>: number of occurrences (<code>n</code>), area of occupancy (<code>aoo</code>) and <code>coverage</code> of taxa (<code>taxonKey</code>) in Natura2000 areas of Belgium (<code>SITECODE</code>). Other columns included: <code>speciesKey</code> (for species is <code>speciesKey</code> = <code>taxonKey</code>), <code>SITETYPE</code> containing the site type of the Natura2000 area (one of <code>A</code>, <code>B</code> or <code>C</code>), <code>min_coord_uncertainty</code> with the lowest coordinate uncertainty in meters, <code>is_alien</code> containing the alien status (<code>TRUE</code> or <code>FALSE</code>) and <code>remarks</code> containing, if present, the infraspecific alien taxa whose occurrences contribute to the calculated <code>aoo</code> (only for species).</li> <li><strong>protected_areas_species_info.csv</strong>: taxonomic information of taxa in <code>protected_areas_species_occurrence.csv</code> as retrieved from <a href="https://www.gbif.org/dataset/d7dddbf4-2cf0-4f39-9b2a-bb099caae36c">GBIF Backbone Taxonomy</a>. Columns: <code>taxonKey</code>, <code>speciesKey</code>, <code>scientificName</code>, <code>kingdom</code>, <code>phylum</code>, <code>order</code>, <code>class</code>, <code>genus</code>, <code>family</code>, <code>species</code>, <code>rank</code> and <code>includes</code>. The latter contains the infraspecific taxa and synonyms whose occurrences contribute to the number of occurrences at species level.</li> <li><strong>protected_areas_metadata.csv</strong>: protected area information for areas included in <code>protected_areas_species_occurrence.csv</code>. Columns: <code>SITECODE</code> as in <code>protected_areas_species_occurrence.csv</code> (<code>BE*******</code>), <code>SITENAME</code> containing the name of the protected area, <code>SITETYPE</code> as in <code>protected_areas_species_occurrence.csv</code>, <code>flanders</code>, <code>wallonia</code> and <code>brussels</code> containing whether the area is situated respectively in Flanders, Wallonia or Brussels-Capital Region (<code>TRUE</code> or <code>FALSE</code>). Field codes are in line with <a href="https://www.eea.europa.eu/data-and-maps/data/natura-11/natura-2000-tabular-data-12-tables">EEA element definitions</a> for Natura 2000 sites.</li> </ul> <p><strong>Potential use of the dataset</strong></p> <p>Currently, there is no comprehensive reporting system for invasive alien species in Natura 2000 sites. This dataset provides a baseline as to which species occur in which protected area. We envisage this dataset can be an interesting starting point for various types of analyses on alien species in protected areas in Belgium, but that it can also be used in complement to other data on alien species in protected areas to study more general patterns. Some examples of research questions:</p> <ul> <li>Which protected areas are most invaded by alien species</li> <li>Which alien species are most distributed in protected areas and which traits do they have</li> <li>How does the proportion of alien species in protected areas change in time</li> <li>How does the occurrence/occupancy of alien species in protected areas match lists of regulated species (e.g. Union List, EPPO lists)</li> <li>To what extent can the network of protected areas contribute to providing safe refuge to native species from the impacts of invasive alien species</li> <li>How widespread are the impacts of certain alien species on protected areas</li> </ul> <h2>Acknowledgements</h2> <p>This work has been funded under the Belgian Science Policies Brain program (BelSPO BR/165/A1/TrIAS), the European Union's LIFE program (LIFE19 NAT/BE/000953 - LIFE RIPARIAS).</p>
Sentiment analysis of tech media articles using VADER package and co-occurrence analysis (01.2016-12.2019)
<p>Sentiment analysis of tech media articles using VADER package and co-occurrence analysis</p> <p>Sources with weights:</p> <ul> <li>Euractiv 5%</li> <li>The Conversation 5%</li> <li>Politico Europe 5 %</li> <li>IEEE Spectrum 5 %</li> <li>Techforge 5%</li> <li>Fastcompany 5%</li> <li>The Guardian (Tech) 12%</li> <li>Arstechnica 5%</li> <li>Reuters 5%</li> <li>Gizmodo 9%</li> <li>ZDNet 9%</li> <li>The Register 12%</li> <li>The Verge 9%</li> <li>TechCrunch 9%</li> </ul> <p>Methodology</p> <p>The sentiment analysis has been prepared using VADER*, an open-source lexicon and rule-based sentiment analysis tool. VADER is specifically designed for social media analysis, but can be also applied for other text sources. The sentiment lexicon was compiled using various sources (other sentiment data sets, Twitter etc.) and was validated by human input. The advantage of VADER is that the rule-based engine includes word-order sensitive relations and degree modifiers.</p> <p>As VADER is more robust in the case of shorter social media texts, the analysed articles have been divided into paragraphs. The analysis have been carried out for the social issues presented in the co-occurrence exercise.</p> <p>The process included the following main steps:</p> <ul> <li>The 100 most frequently co-occurring terms are identified for every social issue (using the co-occurrence methodology)</li> <li>The articles containing the given social issue and co-occurring term are identified</li> <li>The identified articles are divided into paragraphs</li> <li>Social issue and co-occurring words are removed from the paragraph</li> <li>The VADER sentiment analysis is carried out for every identified and modified paragraph</li> <li>The average for the given word pair is calculated for the final result</li> </ul> <p>Therefore, the procedure has been repeated for 100 words for all identified social issues.</p> <p>The sentiment analysis resulted in a compound score for every paragraph. The score is calculated from the sum of the valence scores of each word in the paragraph, and normalised between the values -1 (most extreme negative) and +1 (most extreme positive). Finally, the average is calculated from the paragraph results. Removal of terms is meant to exclude sentiment of the co-occurring word itself, because the word may be misleading, e.g. when some technologies or companies attempt to solve a negative issue. The neighbourhood's scores would be positive, but the negative term would bring the paragraph's score down.</p> <p> </p> <p>*Hutto, C.J. & Gilbert, E.E. (2014). VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text. Eighth International Conference on Weblogs and Social Media (ICWSM-14). Ann Arbor, MI, June 2014.</p>
Figure 1 in Co-occurrence of three Aristolochia-feeding Papilionids (Archon apollinus, Zerynthia polyxena and Zerynthia cerisy) in Greek Thrace
Figure 1. Position of the study area in northeastern Greece (dark dot on the map) and mutual positions of the three study subsites, with the mosaics of individual biotopes. The longest single moves of three study species: Aa, Archon apollinus; Zc, Zerynthia cerisy; Zp, Zerynthia polyxena.
Figure 4 in Co-occurrence of three Aristolochia-feeding Papilionids (Archon apollinus, Zerynthia polyxena and Zerynthia cerisy) in Greek Thrace
Figure 4. Results of model for eggs and larval records. (A) Interaction plot showing average egg batch sizes for individual butterfly species on individual species of Aristolochia plants; (B) box-plots (medians and quartiles) showing the amount of canopy closure (variable Trees10: see Material and methods) above Aristolochia plants bearing eggs of the respective butterflies; (C) numbers of larvae of the three studied butterfly species recorded during searches for larvae, note the unbalanced scale on the x-axis; (D) interaction plot showing average number of larvae of the three studied butterfly species in individual instars. For panels (A, C, D) dotted line, Archon apollinus; dashed line, Zerynthia cerisy; full line, Zerynthia polyxena.
Figure 2 in Co-occurrence of three Aristolochia-feeding Papilionids (Archon apollinus, Zerynthia polyxena and Zerynthia cerisy) in Greek Thrace
Figure 2. Adults of the studied butterflies: (A) Archon apollinus; (B) Zerynthia cerisy; (C) Zerynthia polyxena (the small inserts stand for host plant species used by the respective species at the study locality); and drawings of their Aristolochia host plants (D) Aristolochia pallida; (E) Aristolochia rotunda; (F) Aristolochia clematitis; (G) Aristolochia hirta; (H) Dissected subterranean Aristolochia hirta flower with A. apollinus first-instar larvae; (I) Habitat mosaic at the Greek Thrace study site, showing a field in the front, and scrub with open forest in the background; (J) Silk-woven Aristolochia hirta leaves with A. apollinus larvae.
Figure 3 in Co-occurrence of three Aristolochia-feeding Papilionids (Archon apollinus, Zerynthia polyxena and Zerynthia cerisy) in Greek Thrace
Figure 3. Estimates of the adult daily population sizes based on mark–recapture data: year 2010, when only data for Archon apollinus (most of flight period) and Zerynthia cerisy (late tail of flight period) allowed the estimation; year 2011, A. apollinus, Z. cerisy, Zerynthia polyxena. The error lines present standard errors of estimates, see Table 3 for model parameters.
Accounting for environmental variation in co‐occurrence modelling reveals the importance of positive interactions in root‐associated fungal communities
<p>Understanding the role of interspecific interactions in shaping ecological communities is one of the central goals in community ecology. In fungal communities, measuring interspecific interactions directly is challenging because these communities are composed of large numbers of species, many of which are unculturable. An indirect way of assessing the role of interspecific interactions in determining community structure is to identify the species co-occurrences that are not constrained by the environmental conditions. In this study, we investigated co-occurrences among root-associated fungi, asking whether fungi co-occur more or less strongly than expected based on the environmental conditions and the host plant species examined. For this purpose, we generated molecular data on root-associated fungi of five plant species evenly sampled along an elevational gradient at a high Arctic site. We analysed the data using a joint species distribution modelling approach that allowed us to identify those co-occurrences that could be explained by the environmental conditions and the host plant species, as well as those co-occurrences that remained unexplained and thus more likely reflect interactive associations. Our results indicate that positive interactions play an important role in shaping microbial communities in arctic plant roots. In particular, we found that mycorrhizal fungi are especially prone to positively co-occur with other fungal species. Our results bring new understanding to the structure of arctic interaction networks by suggesting that interactions among root-associated fungi are predominantly positive.</p>
Data and code for the manuscript: "Varying richness need not imply non-random species co-occurrence: implications for specifying null models"
<p>Data and R code for the manuscript "Varying richness need not imply non-random species co-occurrence: implications for specifying null models".</p>
Aeronautical Occurrences from the Brazilian Civil Aviation (2010/2019) Dataset
<p>This dataset is originated from the dataset of Aeronautical Occurrences managed by Centro de Investigação e Prevenção de Acidentes Aeronáuticos (CENIPA) and available at the Portal Brasileiro de Dados Abertos website (https://dados.gov.br/dataset/ocorrencias-aeronauticas-da-aviacao-civil-brasileira), which was processed and integrated into this dataset. The dataset was used on the Accidents Severity Classifier Model v1.0 (http://doi.org/10.5281/zenodo.4298791) for a Data Science Project. Metadata is available at: <https://drive.google.com/file/d/1Dpp3uEYI-9B_Gf36j5CwARKrWgtgyDOt/view?usp=sharing>.</p>
Fig. 3 in A case of intersex occurrence in Steindachneridion parahybae (Steindachner, 1877) (Siluriformes: Pimelodidae) under captivity condition: a cytogenetic and morphological study
Fig. 3. Micrographs of the ovotestes of Steindachneridion parahybae juveniles in captivity showing the different germ cell types. a) unrestricted spermatogonial testes with anastomosing tubular type, composed by numerous tubules (arrowhead); b) many spermatozoa inside the tubular lumen (arrowhead), separated by interstitial tissue (asterisk); c) gonads section showing perinucleolar oocyte (arrow) close to testicular tissue (arrowhead), containing spermatozoa inside the duct and many intratubular cists; d) ovotestes showing perinucleolar oocyte (arrow) close to spermatozoa inside the duct (arrowhead). Periodic-Acid-Schiff (PAS)/Weigert's Haematoxylin/Metanil Yellow staining. Scale bars: 74 µm (a); 19 µm (b, d); 37 µm (c).
Fig. 4 in A case of intersex occurrence in Steindachneridion parahybae (Steindachner, 1877) (Siluriformes: Pimelodidae) under captivity condition: a cytogenetic and morphological study
Fig. 4. Karyotype in Giemsa in male (a), detail in (b), in female (c), detail in (d) and in intersex animal (e), detail in (f) of Steindachneridion parahybae from Unidade de Hidrobiologia e Aquicultura, Companhia Energética de São Paulo. See the text for more information about karyotype characteristics. Scale bars: 37 μm (a,c,e); 7,4 μm (b,d,f).
Fig. 2 in A case of intersex occurrence in Steindachneridion parahybae (Steindachner, 1877) (Siluriformes: Pimelodidae) under captivity condition: a cytogenetic and morphological study
Fig. 2. Micrographs of the testes of Steindachneridion parahybae juveniles in captivity showing the different germ cell types. a) unrestricted spermatogonial testes with anastomosing tubular type, composed by numerous tubules (arrowhead), which contained the germinal compartment (GC, indicate with arrow) and separated by interstitial tissue; b) isolated spermatogonia cell (arrow) surrounded by Sertoli cells (arrowhead), separated by interstitial tissue (asterisk); c) many spermatocytes and spermatids within the cysts (arrowhead), separated by interstitial tissue (asterisk); d) cysts of spermatocytes (arrowhead), separated by interstitial tissue (asterisk); e) cysts of spermatids (arrowhead), separated by interstitial tissue (asterisk); f) many spermatozoa inside the tubular lumen (arrowhead). Periodic-Acid-Schiff (PAS)/Weigert's Haematoxylin/Metanil Yellow staining. Scale bars: 74µm (a); 37 µm (f); 19 µm (c,e); 14 µm (b); 7.4 µm (d).
SPIKEPIPE: A metagenomic pipeline for the accurate quantification of eukaryotic species occurrences and intraspecific abundance change using DNA barcodes or mitogenomes
<p>The accurate quantification of eukaryotic species abundances from bulk samples remains a key challenge for community ecology and environmental biomonitoring. We resolve this challenge by combining shotgun sequencing, mapping to reference DNA barcodes or to mitogenomes, and three correction factors: (a) a percent‐coverage threshold to filter out false positives, (b) an internal‐standard DNA spike‐in to correct for stochasticity during sequencing, and (c) technical replicates to correct for stochasticity across sequencing runs. The SPIKEPIPE pipeline achieves a strikingly high accuracy of intraspecific abundance estimates (in terms of DNA mass) from samples of known composition (mapping to barcodes R<sup>2</sup> = .93, mitogenomes R<sup>2</sup> = .95) and a high repeatability across environmental‐sample replicates (barcodes R<sup>2</sup> = .94, mitogenomes R<sup>2</sup> = .93). As proof of concept, we sequence arthropod samples from the High Arctic, systematically collected over 17 years, detecting changes in species richness, species‐specific abundances, and phenology. SPIKEPIPE provides cost‐efficient and reliable quantification of eukaryotic communities.</p>
FIGURE 2 in First occurrence of the genus Paraleucilla (Calcarea, Porifera) in the Atlantic Ocean: P. m a g n a sp. nov.
FIGURE 2. External morphology of Paraleucilla magna sp. nov. A — Sponge in situ; B — Preserved holotype; C — Detail of osculum and atrial cavity of holotype.
FIGURE 1 in First occurrence of the genus Paraleucilla (Calcarea, Porifera) in the Atlantic Ocean: P. m a g n a sp. nov.
FIGURE 1. Distribution of Paraleucilla magna sp. nov. 1 — Type locality, Praia Vermelha; 2 — Cagarras Archipelago; 3 — Arraial do Cabo (Rio de Janeiro State); 4 — Alcatrazes Archipelago (São Paulo State).
FIGURES 14 – 16 in The occurrence of Taraxitrichia Flint & Harris, 1992 (Trichoptera: Hydroptilidae) in Brazil, with description of the final larval stage
FIGURES 14 – 16. Taraxitrichia sp. larva and pupa. 14, lateral view of larval case and larva; 15, lateral view of larval thorax; 16, view of pupal cases fixed on a sponge colony.
FIGURES 1 – 7 in The occurrence of Taraxitrichia Flint & Harris, 1992 (Trichoptera: Hydroptilidae) in Brazil, with description of the final larval stage
FIGURES 1 – 7. Taraxitrichia sp. larva. 1, dorsal view of thorax and abdomen (scale = 1 mm): a, setae with spinules; b. lateral projection abdominal of segment II; 2, thorax ventral view (PR, MS and MT), and first abdominal segment (I): a. sclerite of pronotum; b, hyaline tubular structures; c – d, sclerite of meso and metanotum; e, protuberances, not sclerotized of I abdominal segment; 3, head dorsal view; 4, prothoracic leg (arrow indicates the trochantin); 5, mesothoracic leg. 6, metathoracic leg; 7, anal proleg, lateral view (scale = 0.012 mm).
FIGURE 1 in On the occurrence of caligids (Copepoda: Siphonostomatoida) in the marine plankton: a review and checklist
FIGURE 1. Map showing geographical distribution (38 localities) of caligid species collected in plankton samples. Details of species with known hosts are summarized in Table 1; those known exclusively from planktonic records are listed in Table 2. Localities marked are: 1. Black Sea, 2. San Francisco, 3. Newport Beach, 4. Sonora, 5. Off Laguna de Tamiahua, 6. Texas, 7. Miami, 8. Chelem, 9. Costa Rica, 10. Venezuela, 11. Easter Islands, 12. Falkland Islands, 13. Rio Grande do Sul, 14. Gulf of Guinea, 15. off Trivandrum, 16. off Kayankulam, 17. Gulf of Thailand, 18. east of New Zealand, 19. Mindanao, 20. Amoy, 21. Qingdao (= Tsingtao), 22. Yantai (= Chefoo), 23. Yellow River estuary, 24. Mankyong River, 25. Seomjin River, 26. Iheyashima, 27. Amami Island, 28. Akuseki-shima, 29. Ariake Sea, 30. Hirado Island, 31. Suo-nada (Seto Inland Sea), 32. Ube, 33. Mukaijima, 34. Takamatsu, 35. Nagasu, 36. Osaka Port, 37. Tokushima, and 38. Uwajima.
Species co-occurrences from EuPMC articles related to pines
<p>A dataset containing info on matches from full text searches by ContentMine tools, that can be mapped to Wikidata. See README.md.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.