Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
28,952
datasets available to search
ShareScore release 0.7.1
Dataset results
28,952 results for “Distributed”
Reference data set for a Norwegian medium voltage power distribution system
<p>This reference data set describes a representative Norwegian radial, medium voltage (MV) electric power distribution system operated at 22 kV. The data set is developed in the Norwegian research centre CINELDI and will in brief be referred to as the CINELDI MV reference system.</p> <p>Data for a real Norwegian distribution system were provided by a distribution grid company. The data have been anonymized and processed to obtain a simplified but still realistic grid model with 124 nodes. The data set consists of the following three parts:<br> 1. Grid data files: describe the base version of the reference system that represents the present-day state of the grid, including information about topology, electrical parameters, and existing load points.<br> 2. Load data files: comprise load demand time series for a year with hourly resolution and scenarios for the possible long-term development of peak load. These data describe an extended version of the reference system with information about possible new load points being added to the system in the future.<br> 3. Reliability data files: contain data necessary for carrying out reliability of supply analyses for the system.</p> <p>The data set is described in detail in the following data article:<br> I. B. Sperstad, O. B. Fosso, S. H. Jakobsen, A. O. Eggen, J. H. Evenstuen, and G. Kjølle, “Reference data set for a Norwegian medium voltage power distribution system,” Data in Brief, 109025, 2023, doi: 10.1016/j.dib.2023.109025.</p>
Bathymetry and Sediment thickness distribution of Lago dei Seracchi alpine lake, Rutor basin, Aosta Valley, Italy
<p>Maps of water depth and sediment accumulation in an Italian proglacial lake, done by Ground Penetrating Radar (GPR) in July 2021. Supporting Time domain reflectometry surveys and geotechnical analyses on the sediments are also provided. For details, see the readme file in the dataset folder.</p>
Belgian baseline distribution of invasive alien species of Union concern (Regulation (EU) 1143/2014)
<p><strong>Aims and scope</strong></p> <p>The European Alien Species Information Network team (EASIN, http://easin.jrc.ec.europa.eu) of the Joint Research Centre (JRC) requests the European member states to provide and verify the baseline distribution data of invasive alien species of Union Concern (Tsiamis et al. 2017) as provided by the EASIN mapping system (Katsanevakis et al. 2012). These are species with documented biodiversity impacts sensu the European Union Regulation on the prevention and management of the introduction and spread of Invasive Alien Species in Europe (IAS Regulation No 1143/2014) (European Union 2014). The purpose of this baseline is to set a representative geographic account of the distribution of these species at (i) country and (ii) 10km<sup>2</sup> grid level before the entry into force of the Regulation (and the listing of species through implementing regulations). This distribution provides the baseline for subsequent reporting by the member states as required by the IAS Regulation.</p> <p>The dataset provides a shapefile on the baseline distribution of the invasive species of EU concern in Belgium based on an aggregated dataset (<em>ias_belgium_t0_xxxx</em>). Data were compiled from various datasets holding invasive species observations such as data from research institutes and research projects (76%), citizen science observatories (23%) and a range of other sources (1%) such as governmental agencies, water managers, invasive species control companies, angling and hunting organizations etc. Data were normalized using a custom mapping of the original data files to Darwin Core (Wieczorek et al. 2012) where possible. Species names were mapped to the GBIF Backbone Taxonomy (GBIF 2016) using the species API (http://www.gbif.org/developer/species). Appropriate selection of records was performed based on predefined cut-off dates (see data range) and record content validation (see validation procedure). Data were then joined with GRID10k layer Belgium based on GRID10k cellcodes (ETRS_1989_LAEA).</p> <p><strong>File description</strong></p> <p>The dataset contains two types of data:</p> <ol> <li> <p>Shapefiles (<em>ias_belgium_t0_2016.zip, ias_belgium_t0_2018.zip, ias_belgium_t0_2020.zip and ias_belgium_t0_2023.zip</em>) providing the presence of the species of EU concern at 10km<sup>2</sup> (European Terrestrial Reference System projection - 1989 ETRS_1989_LAEA) level (resp. for 1st, 2nd, 3rd and 4th batch of species added to the Union List). The attributes table field “ACCEPTED” provides coded information on the distribution validation: correct squares (Y) represent data overlapping between the collated baseline data for Belgium and the EASIN maps. Incorrect data (N) can represent records mapped on wrong 10km2 squares, non-validated records or records that fall outside of the date range applied. New squares (New) represent previously unpublished data that were absent from EASIN. The work was supervised and validated by the Belgian national scientific council on invasive alien species, an official consultative structure coordinating scientific input and data aggregation between Belgian regions and institutions with regards to technical implementation of the Regulation No 1143/2014 on invasive alien species.</p> </li> <li> <p>A geojson version of the same shapefiles (<em>ias_belgium_t0_2016.geojson, ias_belgium_t0_2018.geojson, ias_belgium_t0_2020.geojson, ias_belgium_t0_2023.geojson</em>), in WGS84 projection.</p> </li> </ol> <p><strong>Date range</strong></p> <p>The baseline distribution reflects the current status and situation of the IAS of Union concern in Belgium at 10km<sup>2</sup> grid level. Historical records were not taken into consideration for the baseline. The choice of cut-off date was based on an analysis of the relative contribution of a year in defining the total distribution of the species at 1km<sup>2</sup> grid level (calculated as [the sum of unique UTM 1km<sup>2</sup> grid squares year-1/total number of unique UTM 1km<sup>2</sup> grid squares for that species]) based on the complete dataset. </p> <p>The dataset comprises observations of Union List invasive species <strong>from 2000 <em>until the entry into force </em>for every species</strong>, hence between January 2000 (2000-01-01) and February 2016 (2016-01-31) for the species of the first batch (<em>ias_belgium_t0_2016.zip</em>), between January 2000 (2000-01-01) and August 2017 (2017-08-31) for the species of the first update of the Union List (<em>ias_belgium_t0_2018.zip</em>), between January 2000 (2000-01-01) and August 2019 (2019-08-31) for the species of the second update of the Union List (<em>ias_belgium_t0_2020.zip</em>), between January 2000 (2000-01-01) and August 2022 (2022-08-2) for the species of the third update (<em>ias_belgium_t0_2023.zip</em>). For raccoon dog (<em>Nyctereutes procyonoides), </em>included in the second update (<em>ias_belgium_t0_2020.zip</em>) the date cut-off is 01/01/2000 to 31/01/2019. Note that <em>Pistia stratiotes</em>, <em>Xenopus laevis </em>and <em>Fundulus heteroclitus </em>enter into force only as from 2 August 2024, <em>Celastrus orbiculatus </em>on 2 August 2027 because of prolonged transitionary measures. However, these species are already included in the baseline now with a cut-off date set on August 2022. The data include both casual records as well as established populations and also comprise data from eradicated populations for the period 2000-2022.</p> <p><strong>Validation procedure</strong></p> <p>Record validation was performed to exclude dubious records, wrong identifications etc. This was done based on the IdentificationVerificationStatus field (to which validation information from original data were mapped) if available. In general, non-validated data were not considered for ias_belgium_t0_xxxx. Data were validated in the original datasets based on evidence (e.g. pictures), on the observer’s experience, or based on a set of predefined rules (e.g. automated validation based on geographic filtering). Data from research institutes were generally considered validated. A few casual records of EU list species that were clearly planted were discarded manually. When the original dataset did not mention any validation status, records were not considered validated and therefore not taken into account for ias_belgium_t0_xxxx, unless for Chinese mitten crab <em>Eriocheir sinensis</em>, ruddy duck <em>Oxyura jamaicensis</em>, raccoon <em>Procyon lotor</em>, Siberian ground squirrel <em>Tamias sibiricus</em>, sacred ibis <em>Threskiornis aethiopicus</em>, and red-eared slider <em>Trachemys spp</em>. For these species, we assumed all records were correct as they originate from dedicated sampling (<em>E. sinensis</em>) within research projects or represent species that are readily recognizable by people in the field. Likewise, for the second batch species, all records of Egyptian goose <em>Alopochen aegyptiaca, </em>Himalayan balsam <em>Impatiens glandulifera</em>, giant hogweed <em>Heracleum mantegazzianum </em>and muskrat <em>Ondatra zibethicus</em> (mostly derived from public eradication services) were considered validated and taken into account. For the third batch species, records of the widespread tree of heaven <em>Ailanthus altissima </em>and pumpkinseed <em>Lepomis gibbosus </em>were also considered validated. For species with less than 10 records (<em>Salvinia molesta</em>, <em>Acridotheres tristis</em>), every record was manually checked.</p> <p>A visual check was performed on the resulting distribution maps by representatives of the Belgian scientific council on IAS and the Belgian Comittee on IAS, two official bodies created in response to the EU Regulation within the framework of a cooperation agreement between the Belgian regions and the Federal Authority. Data in the distribution maps provided by EASIN but not present in ias_belgium_t0_xxxx were carefully checked and kept/rejected accordingly.</p> <p><strong>Data providers</strong></p> <p>The providers of the invasive species data for this exercise (individuals and their respective organizations) are listed in the "data providers" section of the dataset metadata. Much of the primary occurrence data that formed the basis for this aggregated dataset will be published as open data on the Global Biodiversity Information Facility (GBIF) within the framework of the <strong>Tracking Invasive Alien Species project (TrIAS, https://osf.io/7dpgr/, 2017-2020)</strong>.</p>
Distribution models for riparian landbirds and waterbirds in the Sacramento-San Joaquin Delta
<p><strong>SUMMARY</strong><br> Distribution models for 9 riparian landbird species and 6 groups of waterbird species in the Sacramento-San Joaquin River Delta of California. </p> <p><strong>DESCRIPTION</strong><br> These predictive models were developed to relate the probability of species or group presence as a function of the surrounding landscape, facilitating predictions of species presence or absence over the entire landscape. Each .RData object is structured as a list containing individual model objects of class `gbm` for each species or group.</p> <p>Models were developed using Boosted Regression Trees, implemented in R using the R packages `dismo` (Hijmans et al. 2021) and `gbm` (Greenwell et al. 2020). Models were developed from pre-existing bird survey data, including 2,547 surveys for riparian landbirds conducted at 716 unique locations throughout the Central Valley of California during the breeding season (May and June), 2011–2019, and 7,820 surveys for waterbirds conducted at 504 unique locations in the Delta during the fall (July 15–November 15) and winter (November 17–March 5) seasons, 2013–14 and 2014–15. Waterbird models were developed for each of the fall and winter seasons, with 46 species grouped into 6 distinct groups defined by similar habitat requirements, foraging style, and diet. </p> <p>These models were used to predict the distribution of each species and group across a baseline Delta landscape (representing land cover in 2018), and these predictions were used to identify Priority Bird Conservation Areas in the Delta. In addition, the models were used to predict distributions for alternative scenarios of future landscape change, and to evaluate the net change from the baseline distributions in the total area of suitable habitat. These models are required for evaluating the change in Biodiversity Support benefits using the R package "DeltaMultipleBenefits", which provides the code and work flow for repeating the initial scenario analyses or analyzing new scenarios.</p> <p>For additional details about the development and applications of these data, please see: </p> <ul> <li>Dybala K, Sesser K, Reiter M, Shuford WD, Golet GH, Hickey C, Gardali T. (<em>In review</em>) Priority Bird Conservation Areas in California’s Sacramento–San Joaquin Delta.</li> <li>Dybala KE, et al. (<em>In review</em>) Multiple-benefit Conservation in Practice: A Framework for Quantifying Multi-dimensional Impacts of Landscape Change in California’s Sacramento–San Joaquin Delta.</li> <li>Dybala KE (2023) <em>DeltaMultipleBenefits: Projecting the Multiple Benefits of Land Cover Change in the Sacramento-San Joaquin River Delta</em>. R package version 1.0.0. doi:10.5281/zenodo.7718620. https://pointblue.github.io/DeltaMultipleBenefits </li> </ul> <p><strong>Literature Cited:</strong></p> <ul> <li>Greenwell B, Boehmke B, Cunningham J, Developers G (2020). <em>gbm: Generalized Boosted Regression Models</em>. R package version 2.1.8. https://CRAN.R-project.org/package=gbm</li> <li>Hijmans RJ, Phillips S, Leathwick J, Elith J (2021). <em>dismo: Species Distribution Modeling</em>. R package version 1.3-5. https://CRAN.R-project.org/package=dismo</li> </ul> <p><strong>FUNDING STATEMENT</strong><br> These data were developed as part of the project "Trade-offs and Co-benefits of Landscape Change on Bird Communities and Ecosystem Services in the Sacramento–San Joaquin River Delta", funded by Proposition 1 Delta Water Quality and Ecosystem Restoration Program, Grant Agreement Number – Q1996022, administered by the California Department of Fish and Wildlife.</p> <p><strong>POINT OF CONTACT</strong><br> Kristen Dybala, Point Blue Conservation Science, kdybala@pointblue.org</p> <p><strong>SUGGESTED CITATION</strong><br> Dybala KE, Sesser KA, Reiter ME, Shuford WD, Golet GH, Hickey CM, Gardali T. 2023. Distribution models for riparian landbirds and waterbirds in the Sacramento-San Joaquin Delta. doi: 10.5281/zenodo.7531945</p> <p><strong>DATA DISTRIBUTION</strong><br> Zenodo. (https://doi.org/10.5281/zenodo.7531945)</p> <p><strong>PROGRESS</strong><br> Complete, but note that the accompanying manuscript has not yet undergone peer-review, and thus these data may require future revision.</p> <p><strong>UPDATE FREQUENCY</strong><br> Not Planned</p> <p><strong>DATE</strong><br> These models were developed 2019-2022, based on bird survey data collected 2011-2019.</p> <p><strong>FIELD DEFINITIONS</strong><br> N/A</p> <p><strong>ABBREVIATION DEFINITIONS</strong></p> <p>BRT_models_riparianlandbirds.RData:</p> <ul> <li><strong>NUWO:</strong> Nuttall's Woodpecker (<em>Picoides nuttallii</em>)</li> <li><strong>ATFL: </strong>Ash-throated Flycatcher (<em>Myiarchus cinerascens</em>)</li> <li><strong>BHGR: </strong>Black-headed Grosbeak (<em>Pheucticus melanocephalus</em>)</li> <li><strong>LAZB: </strong>Lazuli Bunting (<em>Passerina amoena</em>)</li> <li><strong>COYE:</strong> Common Yellowthroat (<em>Geothlypis trichas</em>)</li> <li><strong>YEWA: </strong>Yellow Warbler (<em>Setophaga petechia</em>)</li> <li><strong>SPTO: </strong>Spotted Towhee (<em>Pipilo maculatus</em>)</li> <li><strong>SOSP:</strong> Song Sparrow (<em>Melospiza melodia</em>)</li> <li><strong>YBCH: </strong>Yellow-breasted Chat (<em>Icteria virens</em>)</li> </ul> <p>BRT_models_waterbirds.RData:</p> <ul> <li><strong>geese:</strong> Geese <ul> <li>Greater White-fronted Goose (<em>Anser albifrons</em>)</li> <li>Snow Goose (<em>Anser caerulescens</em>)</li> <li>Ross's Goose (<em>Anser rossii</em>)</li> <li>Cackling Goose (<em>Branta hutchinsii</em>)</li> <li>Canada Goose (<em>Branta canadensis</em>)</li> </ul> </li> <li><strong>dblr: </strong>Dabbling ducks, including: <ul> <li>Wood Duck (<em>Aix sponsa</em>)</li> <li>Gadwall (<em>Mareca strepera</em>)</li> <li>American Wigeon (<em>Mareca americana</em>)</li> <li>Mallard (<em>Anas platyrhynchos</em>)</li> <li>Blue-winged Teal (<em>Spatula discors</em>)</li> <li>Cinnamon Teal (<em>Spatula cyanoptera</em>)</li> <li>Northern Shoveler (<em>Spatula clypeata</em>)</li> <li>Northern Pintail (<em>Anas acuta</em>)</li> <li>Green-winged Teal (<em>Anas carolinensis</em>)</li> </ul> </li> <li><strong>divduck: </strong>Diving ducks (<em>Note: this model was only developed for the winter season</em>) <ul> <li>Canvasback (<em>Aythya valisineria</em>)</li> <li>Ring-necked Duck (<em>Aythya collaris</em>)</li> <li>Lesser Scaup (<em>Aythya affinis</em>)</li> <li>Bufflehead (<em>Bucephala albeola</em>)</li> <li>Common Goldeneye (<em>Bucephala clangula</em>)</li> <li>Hooded Merganser (<em>Lophodytes cucullatus</em>)</li> <li>Common Merganser (<em>Mergus merganser</em>)</li> <li>Ruddy Duck (<em>Oxyura jamaicensis</em>)</li> </ul> </li> <li><strong>crane: </strong>Cranes <ul> <li>Greater Sandhill Crane (<em>Antigone canadensis tabida</em>)</li> <li>Lesser Sandhill Crane (<em>Antigone canadensis canadensis</em>)</li> </ul> </li> <li><strong>shore: </strong>Shorebirds <ul> <li>Western Sandpiper (<em>Calidris mauri</em>)</li> <li>Least Sandpiper (<em>Calidris minutilla</em>)</li> <li>Dunlin (<em>Calidris alpina</em>)</li> <li>Black-necked Stilt (<em>Himantopus mexicanus</em>)</li> <li>American Avocet (<em>Recurvirostra americana</em>)</li> <li>Greater Yellowlegs (<em>Tringa melanoleuca</em>)</li> <li>Lesser Yellowlegs (<em>Tringa flavipes</em>)</li> <li>Long-billed Dowitcher (<em>Limnodromus scolopaceus</em>)</li> <li>Short-billed Dowitcher (<em>Limnodromus griseus</em>)</li> <li>Wilson's Snipe (<em>Gallinago delicata</em>)</li> </ul> </li> <li><strong>cicon: </strong>Herons/Egrets (Ciconiiformes) <ul> <li>Great Blue Heron (<em>Ardea herodias</em>)</li> <li>Great Egret (<em>Ardea alba</em>)</li> <li>Snowy Egret (<em>Egretta thula</em>)</li> <li>Cattle Egret (<em>Bubulcus ibis</em>)</li> <li>Green Heron (<em>Butorides virescens</em>)</li> <li>Black-crowned Night-Heron (<em>Nycticorax nycticorax</em>)</li> </ul> </li> </ul> <p><strong>ACCESS & USE CONSTRAINTS</strong><br> CC-by-4.0 (https://creativecommons.org/licenses/by/4.0/)</p> <p><strong>KEYWORDS</strong></p> <ul> <li><strong>Themes: </strong>birds, landbirds, songbirds, waterbirds, waterfowl, shorebirds, distribution, habitat</li> <li><strong>Place: </strong>Sacramento-San Joaquin River Delta, Central Valley, California</li> </ul>
Distribution of functionally distinct native and non-indigenous species within marine urban habitats
<p>This data file (.xls) is composed of 5 sheets:</p> <ol> <li>The “Taxon labels”: Taxon code, full name, authority and status/type (Abiotic, Unassigned, Native, Cryptogenic, Non-Indigenous Species)</li> <li>The “Trait labels”: Trait modality and labels and correspondences.</li> <li>The “Taxon-by-Trait matrix”: Fuzzy coded scores for each trait modality and taxon</li> <li>The “Taxon-by-sample matrix”: Abundance data of retained taxa in samples</li> <li>The “Sample labels and description”: Site and experimental factors (Habitat, Age, Experimental Unit, Replicate, nested within site) corresponding to each sample.</li> </ol> <p>Sheets 4 and 5 are extracted from a published dataset, which cannot be shared at this stage of revision without revealing the name of several of the manuscript authors. This is done in respect with the journal guidelines about data storage.</p>
Distribution of functionally distinct native and non-indigenous species within marine urban habitats
<p>This data file (.xls) is composed of 5 sheets:</p> <ol> <li>The “Taxon labels”: Taxon code, full name, authority and status/type (Abiotic, Unassigned, Native, Cryptogenic, Non-Indigenous Species)</li> <li>The “Trait labels”: Trait modality and labels and correspondences.</li> <li>The “Taxon-by-Trait matrix”: Fuzzy coded scores for each trait modality and taxon</li> <li>The “Taxon-by-sample matrix”: Abundance data of retained taxa in samples</li> <li>The “Sample labels and description”: Site and experimental factors (Habitat, Age, Experimental Unit, Replicate, nested within site) corresponding to each sample.</li> </ol> <p>Sheets 4 and 5 are extracted from a published dataset, which cannot be shared at this stage of revision without revealing the name of several of the manuscript authors. This is done in respect with the journal guidelines about data storage.</p>
Data from: The spatial distribution and temporal trends of livestock damages caused by wolves in Europe
<p>The preprint of the corresponding manuscript can be found here: doi: https://doi.org/10.1101/2022.07.12.499715</p> <p>Wolf populations are recovering and expanding across Europe, causing conflicts with livestock owners. We here compiled incident-based livestock damage data caused by wolves across 21 European countries for the years 2018, 2019 and 2020.</p> <p>The file "<strong>wolf_damages_2018_2019_2020_complete_data_to_publish.csv</strong>" contains the following information per incident: country, target species, cause, number of animals killed/injured/missing, assessment level probability, reported date, number of days until inspection, location, incidentID, uniqueID, NUS1_ID, NUTS2_ID, NUTS3_ID, damage prevention measure, number of wolves attacking, latitude, longitude, comments, metadata constraints.</p> <p>The file "<strong>nuts3_regions_and_LC_where_wolves_are_present.csv</strong>" contains information of the percentage of area occupied by wolves per NUTS3 region for selected land cover variables.</p> <p>The file "<strong>prevention_measures.csv</strong>" contains information about the financial support of livestock damage prevention measure per country or NUTS region</p> <p>The file "<strong>wolf_presence_now_vs_50_years_ago_nuts3.csv</strong>" contains information on NUTS3 regions that had a documented wolf presence 50 years ago.</p> <p>The "<strong>scripts_to_publish.zip</strong>" folder contains the scripts that we used to conduct the analyses.</p>
Taxonomy, distribution and classification of ecosystem-types, integrating the recent IUCN function-based typology and local conceptualizations
<p>1. Introduction:</p> <p>This dataset is a work in progress. It compiles data gathered on ecosystem-types and their distribution based on a series of field studies led by the author, in Seychelles and West and Central Africa (Senterre 2014, Senterre & Wagner 2014, Senterre 2016, Senterre et al. 2017, 2019, 2020, 2021a, 2022). The aims of this dataset are:</p> <p>a. To share in an explicit and transparent way data on proposed taxonomies of ecosystems, i.e. conceptualizations of ecosystem-types, including explicit ecosystem names and management of synonymies.</p> <p>b. To develop ecosystem red listing based on transparent and falsifiable distribution raw data, combining distribution modeling (maps) and in situ observation of individual stand occurrences.</p> <p>c. To illustrate in detail how to deal with ecosystem data following the approach described in Senterre et al. (2021b) (i.e. "ecosystemology" approach).</p> <p>d. To integrate the above approach with the newly developed function-based typology of ecosystems (Keith et al. 2022), therefore contributing to bridging the persistent gap between the global and the local scales in ecosystem descriptions and classifications.</p> <p> </p> <p>2. Context and versions:</p> <p>This dataset was initially planned for publication on GBIF (Global Biodiversity Information Facility), as part of a project developed for the review of Key Biodiversity Areas in Seychelles: "Mainstreaming recent species and ecosystem distribution data into Key Biodiversity Areas assessments in Seychelles" (<a href="https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf">https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf</a>).</p> <p>In the first version of the GBIF dataset (<a href="https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf">https://www.gbif.org/dataset/f513fe98-b1c3-45ee-8e14-7f2a5b7890bf</a>), we proposed an analysis of the potential 'core' and 'extension' files available in GBIF for a publication of ecosystem-type names (and synonymies) and their corresponding occurrences recorded from field observations. This is an original analysis of taxonomic principles managed entirely at the scale of local observable objects, and their history of identifications or interpretations.</p> <p>Toward the end of the above-mentioned GBIF project, considering the limitations and gaps currently present in GBIF, it was decided to restrict the GBIF dataset to a simple 'metadata' entry and to publish the complete version of this dataset in Zenodo. This allows to include all tables needed, as well as all required fields without having to accommodate them within the limited GBIF structure (see metadata description on GBIF for more details). The fields of the tables published here are described in the GBIF metadata entry and in the ecosystemology paper (Senterre et al. 2021b).</p> <p> </p> <p>3. New development on typology aspects:</p> <p>In addition, considering that the new IUCN global typology of ecosystems is now published (Keith et al. 2022), we have reviewed in detail the possibility of integration of ecosystems conceptualized using our ecosystemology approach within the new IUCN typology. The result of this analysis is being considered for a publication, and this Zenodo dataset would then be published in full (i.e. including all typology aspects) as supplementary materials. In the meantime, I would be happy to discuss any of these aspects with whoever is interested.</p> <p> </p> <p>4. Access to ecosystem data for conservation actors:</p> <p>Finally, the actual data (published here) on ecosystem-types, their names, synonymies, classification, distribution, and red list status are compiled into a format that we designed to be useful to conservation actors in the form of interactive webpages (produced with R as shiny apps). This development is based on very limited resources, and the author is still quite new to R, so any help or feedback on ways to improve the scripts would be very much welcomed.</p> <p>The interactive page is available here (currently filtered to Seychelles' data only, although the dataset contains data beyond the Seychelles): https://shiny.bio.gov.sc/bioeco/</p> <p>The R scripts are available on Github: https://github.com/bsenterre/ecosystemology</p> <p> </p> <p>5. Tables contained in this dataset:</p> <p>a. Ecosystem taxonomy tables:</p> <p>ecoSpecies: Contains the list of all ecosystem-type names with their unique identifier.</p> <p>ecoOccurrences: Contains the list of individual stand occurrences, including ecosystem characters as standardized in Senterre et al. (2021b; i.e. virtual ecosystem specimen).</p> <p>ecoSpeciesProfiles: Contains basic metadata on ecosystem-types, such as their Red List evaluations.</p> <p>ecoIdentifications: Contains all the different interpretations/identifications (referring to the table ecoSpecies or to higher levels of classification, see below) made on the stands observed in the ecoOccurrences table.</p> <p> </p> <p>b. Ecosystem typology tables (TO BE ADDED LATER):</p> <p>IUCNL3: This is just a transcription, as is, of the IUCN global typology version 2.1.</p> <p>IUCNL3BIOCrossover: This table defines and comments correspondences between BIOL2 (the level 2 of the typology used by us) and the IUCN typology L3 (level 3).</p> <p>BIOL2: This is a variation based on the IUCN typology, here our level 2.</p> <p>BIOL3: This is a variation based on the IUCN typology, here our level 3.</p> <p>BIOL4: This is a variation based on the IUCN typology, here our level 4.</p> <p>ecoGenus: This is a general type of stand (thus excluding any regional ecosystem connotation), defined at a local scale and never combined with any geographic connotation (see ecosystemology paper: Senterre et al. 2021b).</p> <p>ecoFamily: This is a generalized version of the ecoGenus (i.e. still excluding any regional, sub-regional or geographic aspect).</p> <p>ecoOrder: This is a further generalized version of the ecoGenus (see also Senterre et al. 2020).</p> <p>lifeZone: This is a basic and incomplete list of life zones as defined following the Holdridge (1967) approach, with some additional elements proposed in Senterre et al. (2021b).</p> <p> </p> <p>6. Literature cited:</p> <p>Holdridge, L. R. 1967. Life zone ecology. Tropical Science Center, San Jose, Costa Rica.</p> <p>Keith, D. A., J. R. Ferrer-Paris, E. Nicholson, M. J. Bishop, B. A. Polidoro, E. Ramirez-Llodra, M. G. Tozer, J. L. Nel, R. Mac Nally, E. J. Gregr, K. E. Watermeyer, F. Essl, D. Faber-Langendoen, J. Franklin, C. E. R. Lehmann, A. Etter, D. J. Roux, J. S. Stark, J. A. Rowland, N. A. Brummitt, U. C. Fernandez-Arcaya, I. M. Suthers, S. K. Wiser, I. Donohue, L. J. Jackson, R. T. Pennington, T. M. Iliffe, V. Gerovasileiou, P. Giller, B. J. Robson, N. Pettorelli, A. Andrade, A. Lindgaard, T. Tahvanainen, A. Terauds, M. A. Chadwick, N. J. Murray, J. Moat, P. Pliscoff, I. Zager, and R. T. Kingsford. 2022. A function-based typology for Earth’s ecosystems. . Nature 610:513–518. doi:10.1038/s41586-022-05318-4.</p> <p>Senterre, B. 2014. Mapping habitat-types within the Hummingbird site at Dugbe (Liberia, West Africa). Consultancy Report, Missouri Botanical Garden. P. 56. https://doi.org/10.13140/RG.2.2.32628.48003.</p> <p>Senterre, B. 2016. Habitat-type ground-truthing and assessment of ecosystem conservation value in the Bel Air Alufer mining site (Guinea, West Africa), with recommendations for improving the draft map of land cover types. Consultancy Report, Missouri Botanical Garden, A study conducted for Alufer Mining Limited. P. 54.</p> <p>Senterre, B., E. Bidault, and T. Stévart. 2019. Identification et évaluation des écosystèmes menacés du Mont Nimba. Rapport de consultance, Missouri Botanical Garden (MBG), Africa and Madagascar Department. P. 106. https://doi.org/10.13140/RG.2.2.13242.93129.</p> <p>Senterre, B., E. Bidault, T. Stévart, and P. P. Lowry II. 2020. Assessment of Key Biodiversity Areas in the Lofa-Gola-Mano & Nimba complexes (West Africa) using ecosystem criteria. Final Report, Missouri Botanical Garden. P. 146. 10.13140/RG.2.2.17934.89924.</p> <p>Senterre, B., E. Bidault, T. Stévart, M. Wagner, and P. Lowry. 2017. Mapping habitat-types in south-east Kouilou (Republic of Congo). Consultancy Report, Missouri Botanical Garden (MBG), Africa and Madagascar Department, St. Louis, Missouri, USA. P. 163.</p> <p>Senterre, B., R. M. Bristol, G. Gendron, and E. Henriette. 2021a. Fine-tuning conservation priorities in Seychelles at the landscape scale, using global KBA guidelines with both species and ecosystem criteria. Consultancy Report, United Nations Development Programme, GOS/UNDP/GEF Programme Coordination Unit, Victoria, Seychelles.</p> <p>Senterre, B., P. P. Lowry II, E. Bidault, and T. Stévart. 2021b. Ecosystemology: a new approach toward a taxonomy of ecosystems. . Ecological Complexity 47:100945. doi:https://doi.org/10.1016/j.ecocom.2021.100945.</p> <p>Senterre, B., A.-H. Paradis, E. Bidault, T. Stévart, and P. P. Lowry II. 2022. Qualité et distribution des savanes montagnardes du Nimba. Rapport de consultance, Missouri Botanical Garden (MBG), Africa and Madagascar Department. P. 73. http://dx.doi.org/10.13140/RG.2.2.13433.34401.</p> <p>Senterre, B., and M. Wagner. 2014. Mapping Seychelles habitat-types on Mahé, Praslin, Silhouette, La Digue and Curieuse. Consultancy Report, Government of Seychelles, United Nations Development Programme, Victoria, Seychelles. P. 119. https://doi.org/10.13140/RG.2.1.4558.6009.</p>
Green Space Distribution m2 per capita in Valladolid city
<p>Urban green infrastructures are key part of the sustainable development in our cities. They can provide important Ecosystem Services in them, including provisioning, regulating, supporting and cultural services. The total surface of green areas needs to be relativized in terms of total area or per capita, in order to compare results with other cities or to observe the evolution within the same city.</p>
Landuse/Landcover predictors for invasive species distribution modelling in Europe.
<p><strong>Description</strong></p> <p>This data set contains a set of predictors characterizing land use/land cover derived from the CORINE dataset, anthropogenic pressure from the global terrestrial human footprint dataset, and the distance to the nearest waterbody, for continental Europe. All have been aligned with the 1 km<sup>2</sup> EEA Reference Grid. The climate variables based on historical (1976-2005) and future (2040-2070) scenarios are available from De Troch et al., 2020 also via Zenodo. These rasters represent the habitat and anthropogenic predictors needed in the Tracking Invasive Alien Species (TrIAS) workflow for invasive species distribution modelling (wiSDM).</p> <p><strong>Geographic coverage</strong></p> <p>Europe</p> <p><strong>Methods</strong></p> <p>Land use classes were extracted from the CORINE06 100 m GeoTiff downloaded from Copernicus. The percentage of each 1 km<sup>2</sup> EEA Reference Grid cell occupied by coniferous forest, deciduous forest, wetlands, grasslands and agriculture was calculated. Multiple land use sub-classes were aggregated for the following categories: agriculture, wetlands, grasslands (Table 1). These data layers have been processed in R to replace all NAs that are within the European landmass, with zeros to distinguish them from the ocean, which remain NA, as in the CORINE dataset. In this context, a zero reflects the absence of a given land cover attribute. </p> <p>The mean anthropogenic pressure per 1km<sup>2 </sup>EEA Reference Grid cell was extracted from the global terrestrial human footprint dataset (Venter et al, 2016). Distance to the nearest waterbody within each 1km<sup>2</sup> EEA Reference Grid cell was calculated using the 2016 Surface Water Bodies shapefile available from the EEA (https://www.eea.europa.eu/data-and-maps/data/wise-wfd-spatial/surface-water-body). </p> <p> </p> <table> <tbody> <tr> <td>Land Use Class</td> <td>CORINE LABEL</td> </tr> <tr> <td>Agriculture</td> <td>Non-irrigated arable land (211), Rice fields (213),Vineyards (221),Fruit trees and berry plantations (222),Olive groves (223),Pastures (231),Annual crops associated with permanent crops (241),Complex cultivation patterns (242),Land principally occupied by agriculture, with significant areas of natural vegetation (243)</td> </tr> <tr> <td> </td> </tr> <tr> <td> </td> </tr> <tr> <td>Coniferous forest</td> <td>Coniferous forest (312)</td> </tr> <tr> <td>Deciduous forest</td> <td>Broad-leaved forest (311)</td> </tr> <tr> <td>Grassland</td> <td>Natural grasslands (321), Moors and heathland, (322) Sclerophyllous vegetation (323)</td> </tr> <tr> <td>Wetland</td> <td>Inland marshes (411), Peat bogs (412)</td> </tr> </tbody> </table> <p>Table 1. How the the original land use/land cover types as labelled in CORINE were combined (or not).</p> <p><strong>Files</strong></p> <p>distance2water_EEA_1km.tif (distance to nearest waterbody)</p> <p>ESM1000m.tif (mean anthropogenic pressure)</p> <p>corine_perAgriculture.tif</p> <p>corine_perWetland.tif</p> <p>corine_pergrass.tif</p> <p>corine_perdeciduous.tif</p> <p>corine_perConiferous.tif</p> <p> </p>
THE StellaR PAth WP1: Sun-as-a-star plasma Emission Measure Distributions
<p>This folder contains a set of plasma Emission Measure Distributions (EMDs) vs. temperature, derived from observations of the solar corona with the Soft X-ray Telescope (SXT) on board the solar satellite Yohkoh, and the prescription to build EMDs for coronae of solar-type stars with different activity levels, including both quiescent and flaring components. For details read the Description PDF file.</p>
A Modular Quantum Compilation Framework for Distributed Quantum Computing
<p>This repository contains the data used for the plots in "<em>A Modular Quantum Compilation Framework for Distributed Quantum Computing</em>" by D. Ferrari, S. Carretta and M. Amoretti.</p> <p>Data is located in the <em>'data'</em> directory in <em>.csv</em> format, a python script to generate the plots can be found in the main directory. The script was tested with <strong>python3.10</strong> and needs <strong>matplotlib</strong>, <strong>pandas</strong> and <strong>seaborn</strong> packages. Plots are saved as <em>.pdf</em> files in the <em>'figures'</em> directory.</p>
Data associated with the article "Evolution and phylogenetic distribution of endo-α-mannosidase"
<p>Data associated with the article "Evolution and phylogenetic distribution of endo-α-mannosidase"</p> <p>Changelog:</p> <p>version 1.1</p> <ul> <li>added <em>Tunicaraptor</em> motif analysis alignment</li> </ul> <p>version 1.0</p> <ul> <li>Initial release</li> </ul> <p> </p> <p>Funding statement: National Science Centre of Poland is acknowledged for funding of the project 2020/36/C/NZ8/00081, "The role of glycosylation in the emergence of animal multicellularity", which enabled the creation of this research output.</p>
Federated Learning for Distributed Intrusion Detection Systems in Public Networks - Validation Dataset
<p>This dataset has been meticulously prepared and utilized as a validation set during the evaluation phase of "Meta IDS" to asses the performance of various machine learning models. It is now made available for interested users and researchers who seek a reliable and diverse dataset for training and testing their own custom models.</p> <p>The validation dataset comprises a comprehensive collection of labeled entries, that determines whether the packet type is "malicious" or "benign." It covers complex design patterns that are commonly encountered in real-world applications. The dataset is designed to be representative, encompassing edge and fog layers that are in contact with cloud layer, thereby enabling thorough testing and evaluation of different models. Each sample in the dataset is labeled with the corresponding ground truth, providing a reliable reference for model performance evaluation.</p> <p> </p> <p>To ensure convenient distribution and storage, the dataset has been broken down into three separate batches, each containing a portion of the dataset. This allows for convenient downloading and management of the dataset. The three batches are provided as individual compressed files.</p> <p> </p> <p>In order to extract the data, follow the following instructions:</p> <ul> <li>Download and install bzip2 (if not already installed) from the official website or your package manager.</li> <li>Place the compressed dataset file in a directory of your choice.</li> <li>Open a terminal or command prompt and navigate to the directory where the compressed dataset file is located.</li> <li>Execute the following command to uncompress the dataset: <ul> <li>bzip2 -d filename.bz2</li> </ul> </li> <li>Replace "filename.bz2" with the actual name of the compressed dataset file.</li> </ul> <p>Once uncompressed, you will have access to the dataset in its original format for further exploration, analysis, and model training etc. The total storage required for extraction is approximately 800 GB in total, with the first batch requiring approximately 302 GB, the second batch requiring approximately 203 GB, and the third batch requiring approximately 297 GB of data storage.</p> <p> </p> <p>The first batch contains 1,049,527,992 entries, where as the second batch contains 711,043,331 entries, and for the third and last batch we have 1,029,303,062 entries. The following table provides the feature names along with their explanation and example value once the dataset is extracted.</p> <p> </p> <table align="left"> <thead> <tr> <th scope="col">Feature</th> <th scope="col">Description</th> <th scope="col">Example Value</th> </tr> </thead> <tbody> <tr> <td>ip.src</td> <td>Source IP address in the packet</td> <td>a05d4ecc38da01406c9635ec694917e969622160e728495e3169f62822444e17</td> </tr> <tr> <td>ip.dst</td> <td>Destination IP address in the packet</td> <td>a52db0d87623d8a25d0db324d74f0900deb5ca4ec8ad9f346114db134e040ec5</td> </tr> <tr> <td>frame.time_epoch</td> <td>Epoch time of the frame</td> <td>1676165569.930869</td> </tr> <tr> <td>arp.hw.type</td> <td>Hardware type</td> <td>1</td> </tr> <tr> <td>arp.hw.size</td> <td>Hardware size</td> <td>6</td> </tr> <tr> <td>arp.proto.size</td> <td>Protocol size</td> <td>4</td> </tr> <tr> <td>arp.opcode</td> <td>Opcode</td> <td>2</td> </tr> <tr> <td>data.len</td> <td>Length</td> <td>2713</td> </tr> <tr> <td>eth.dst.lg</td> <td>Destination LG bit</td> <td>1</td> </tr> <tr> <td>eth.dst.ig</td> <td>Destination IG bit</td> <td>1</td> </tr> <tr> <td>eth.src.lg</td> <td>Source LG bit</td> <td>1</td> </tr> <tr> <td>eth.src.ig</td> <td>Source IG bit</td> <td>1</td> </tr> <tr> <td>frame.offset_shift</td> <td>Time shift for this packet</td> <td>0</td> </tr> <tr> <td>frame.len</td> <td>frame length on the wire</td> <td>1208</td> </tr> <tr> <td>frame.cap_len</td> <td>Frame length stored into the capture file</td> <td>215</td> </tr> <tr> <td>frame.marked</td> <td>Frame is marked</td> <td>0</td> </tr> <tr> <td>frame.ignored</td> <td>Frame is ignored</td> <td>0</td> </tr> <tr> <td>frame.encap_type</td> <td>Encapsulation type</td> <td>1</td> </tr> <tr> <td>gre</td> <td>Generic Routing Encapsulation</td> <td>'Generic Routing<br> Encapsulation (IP)’</td> </tr> <tr> <td>ip.version</td> <td>Version</td> <td>6</td> </tr> <tr> <td>ip.hdr_len</td> <td>Header length</td> <td>24</td> </tr> <tr> <td>ip.dsfield.dscp</td> <td>Differentiated Services<br> Codepoint</td> <td>56</td> </tr> <tr> <td>ip.dsfield.ecn</td> <td>Explicit Congestion<br> Notification</td> <td>2</td> </tr> <tr> <td>ip.len</td> <td>Total length</td> <td>614</td> </tr> <tr> <td>ip.flags.rb</td> <td>Reserved bit</td> <td>0</td> </tr> <tr> <td>ip.flags.df</td> <td>Don't fragment</td> <td>1</td> </tr> <tr> <td>ip.flags.mf</td> <td>More fragments</td> <td>0</td> </tr> <tr> <td>ip.frag_offset</td> <td>Fragment offset</td> <td>0</td> </tr> <tr> <td>ip.ttl</td> <td>Time to live</td> <td>31</td> </tr> <tr> <td>ip.proto</td> <td>Protocol</td> <td>47</td> </tr> <tr> <td>ip.checksum.status</td> <td>Header checksum status</td> <td>2</td> </tr> <tr> <td>tcp.srcport</td> <td>TCP source port</td> <td>53425</td> </tr> <tr> <td>tcp.flags</td> <td>Flags</td> <td>0x00000098</td> </tr> <tr> <td>tcp.flags.ns</td> <td>Nonce</td> <td>0</td> </tr> <tr> <td>tcp.flags.cwr</td> <td>Congestion Window Reduced<br> (CWR)</td> <td>1</td> </tr> <tr> <td>udp.srcport</td> <td>UDP source port</td> <td>64413</td> </tr> <tr> <td>udp.dstport</td> <td>UDP destination port</td> <td>54087</td> </tr> <tr> <td>udp.stream</td> <td>Stream index</td> <td>1345</td> </tr> <tr> <td>udp.length</td> <td>Length</td> <td>225</td> </tr> <tr> <td>udp.checksum.status</td> <td>Checksum status</td> <td>3</td> </tr> <tr> <td>packet_type</td> <td>Type of the packet which is either "benign" or "malicious"</td> <td>0</td> </tr> </tbody> </table> <p>Furthermore, in compliance with the GDPR and to ensure the privacy of individuals, all IP addresses present in the dataset have been anonymized through hashing. This anonymization process helps protect the identity of individuals while preserving the integrity and utility of the dataset for research and model development purposes.</p> <p> </p> <p>Please note that while the dataset provides valuable insights and a solid foundation for machine learning tasks, it is not a substitute for extensive real-world data collection. However, it serves as a valuable resource for researchers, practitioners, and enthusiasts in the machine learning community, offering a compliant and anonymized dataset for developing and validating custom models in a specific problem domain.</p> <p> </p> <p>By leveraging the validation dataset for machine learning model evaluation and custom model training, users can accelerate their research and development efforts, building upon the knowledge gained from my thesis while contributing to the advancement of the field.</p>
AntarcticBasins v1.04: Antarcitca's Sedimentary Basins Distribution and Classification
<p>GIS package for Antarctic Sedimentary Basins Distribution and Classification</p> <p>Supplement to: Aitken, A. R. A., Li, L., Kulessa, B., Schroeder, D., Jordan, T. A., Whittaker, J. M., et al. (2023). Antarctic sedimentary basins and their influence on ice-sheet dynamics. <em>Reviews of Geophysics</em>, 61, e2021RG000767. <a href="https://doi.org/10.1029/2021RG000767">https://doi.org/10.1029/2021RG000767</a></p>
A non-intrusive reduced order model for the characterisation of the spatial power distribution in large thermal reactors (dataset)
<p>This repository contains the software and datasets needed to reproduce the results presented in the article "<a href="https://doi.org/10.1016/j.anucene.2022.109674">A non-intrusive reduced order model for the characterisation of the spatial power distribution in large thermal reactors</a>", published in Annals of Nuclear Energy.</p>
Antarctic Sedimentary Basin Distribution and Classification
<p>This is the published version (<a href="https://zenodo.org/record/7984586">v1.04</a>) of the GIS package for Antarcitca's Sedimentary Basins Distribution and Classification. </p> <p>Supplement to: Aitken, A. R. A., Li, L., Kulessa, B., Schroeder, D., Jordan, T. A., Whittaker, J. M., et al. (2023). Antarctic sedimentary basins and their influence on ice-sheet dynamics. <em>Reviews of Geophysics</em>, 61, e2021RG000767. <a href="https://doi.org/10.1029/2021RG000767">https://doi.org/10.1029/2021RG000767</a></p> <p>With the release of the published version of the GIS package, future updates to the sedimentary basin mapping can be found at <a href="https://github.com/LL-Geo/AntarcticBasins">https://github.com/LL-Geo/AntarcticBasins</a>, and at <a href="https://doi.org/10.5281/zenodo.7955525">https://doi.org/10.5281/zenodo.7955525</a>.</p> <p>You can download individual GeoTIFF, Shapefile, and GeoJSON files. The complete DistroPackage contains all files, including styles in QGIS and ArcGIS projects.</p> <p> </p>
Afterslip distribution related to the Van Earthquake (Turkey) of 23 Oct. 2011
<p>Two distributions of slip (afterslip) related to the Van earthquake</p> <p>- after 4 days</p> <p>- after 17 days</p> <p>as reported in Figure 9 and 10 of the paper</p> <p>Deformation and Related Slip Due to the 2011 Van Earthquake (Turkey) Sequence Imaged by SAR Data and Numerical Modeling</p> <p><em>Elisa Trasatti, Cristiano Tolomei, Giuseppe Pezzo, Simone Atzori and Stefano Salvi</em></p> <p>Rem. Sens. 2016, 8, 532; doi:10.3390/rs8060532</p> <p> </p> <p>Coordinates, depth and slip in meters.</p> <p>The geographical coordinates are projected with UTM zone 38.</p> <p>The depth of the first row of patches is positive due to the elevation of the area.</p>
MCMC simulations from the posterior distributions of European migration flows
<p>This folder contains three RData files each containing MCMC simulations from the posterior distributions of a Bayesian hierarchal model used to estimate European migration flows, developed as part of the project Quantifying Migration Scenarios for Better Policy (QuantMig, <a href="https://eur03.safelinks.protection.outlook.com/?url=http%3A%2F%2Fwww.quantmig.eu%2F&data=05%7C01%7CP.W.Smith%40soton.ac.uk%7Cc166e7bff7f840a7983b08db94af1244%7C4a5378f929f44d3ebe89669d03ada9d8%7C0%7C0%7C638267252014422806%7CUnknown%7CTWFpbGZsb3d8eyJWIjoiMC4wLjAwMDAiLCJQIjoiV2luMzIiLCJBTiI6Ik1haWwiLCJXVCI6Mn0%3D%7C3000%7C%7C%7C&sdata=yMDYBngYz%2BnzvXmNj6MqseYwso0ICNbREE8TmqQ%2BRJ0%3D&reserved=0">www.quantmig.eu</a>); see Aristotelous, Smith and Bijak (2022) for details. The first set of estimates are disaggregated by origin, destination and time, a breakdown which we denote as ODT. The second and third sets of estimates are again disaggregated by origin, destination and time, but they are additionally disaggregated by other factors. The second set is further disaggregated by age and sex and the third by just birth region. We respectively denote these breakdowns as ODAST and ODBT. Also, included in this folder is an R script to calculate summaries of these posterior distributions (means, and lower and upper quartiles) and output them to csv files, which are also in the folder. For further details, see deliverable_6_4_v1_2.pdf also in the folder.</p>
Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESULTS DATASET (with Mega Journals)
<p>The dataset contains all the data produced running the research software for the study:"Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta".</p> <p>Disclaimer: these results are not considered to be representative, because we have fount that Mega Journals skewed significantly some of the data. The result datasets without Mega Journals are published <a href="https://zenodo.org/record/8249907">here</a>.</p> <p>Description of datasets:</p> <ul> <li><strong>SSH_Publications_in_OC_Meta_and_Open_Access_status.csv: </strong>containing information about OpenCitations Meta coverage of ERIH PLUS Journals as well as their Open Access availability. In this dataset, every row holds data for a Journal of ERIH PLUS also covered by OpenCitations Meta database. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>SSH_Publications_by_Discipline.csv:</strong> containing information about number of publications per discipline (in addition, number of journals per discipline are also included). The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>SSH_Publications_and_Journals_by_Country:</strong> containing information about number of publications and journals per country. The dataset has three columns, the first, labeled <strong>"Country",</strong> contains single countries of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>result_disciplines.json:</strong> the dictionary containing all disciplines as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>result_countries.json:</strong> the dictionary containing all countries as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>duplicate_omids.csv: </strong>a dataset containing the duplicated Journal entries in OpenCitations Meta, structured with two columns: "<strong>OC_omid"</strong>, the internal OC Meta identifier; "<strong>issn", </strong>the issn values associated to that identifier</li> <li><strong>eu_data.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>eu_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of european countries. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_eu.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>us_data.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>us_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of the United States. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_us.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> </ul> <p> </p> <p><strong>Abstract of the research: </strong></p> <p><strong>Purpose:</strong> this study aims to investigate the representation and distribution of Social Science and Humanities (SSH) journals within the OpenCitations Meta database, with a particular emphasis on their Open Access (OA) status, as well as their spread across different disciplines and countries. The underlying premise is that open infrastructures play a pivotal role in promoting transparency, reproducibility, and trust in scientific research.<br> <strong>Study Design and Methodology:</strong> the study is grounded on the premise that open infrastructures are crucial for ensuring transparency, reproducibility, and fostering trust in scientific research. The research methodology involved the use of secondary data sources, namely the OpenCitations Meta database, the ERIH PLUS bibliographic index, and the DOAJ index. A custom research software was developed in Python to facilitate the processing and analysis of the data.<br> <strong>Findings:</strong> the results reveal that 78.1% of SSH journals listed in the European Reference Index for the Humanities (ERIH-PLUS) are included in the OpenCitations Meta database. The discipline of Psychology has the highest number of publications. The United States and the United Kingdom are the leading contributors in terms of the number of publications. However, the study also uncovers that only 38% of the SSH journals in the OpenCitations Meta database are OA.<br> <strong>Originality:</strong> this research adds to the existing body of knowledge by providing insights into the representation of SSH in open bibliographic databases and the role of open access in this domain. The study highlights the necessity for advocating OA practices within SSH and the significance of open data for bibliometric studies. It further encourages additional research into the impact of OA on various facets of citation patterns and the factors leading to disparity across disciplinary representation.</p> <p><strong>Related resources:</strong></p> <p>Ghasempouri S., Ghiotto M., & Giacomini S. (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESEARCH ARTICLE. <a href="https://doi.org/10.5281/zenodo.8263908">https://doi.org/10.5281/zenodo.8263908</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S., (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - DATA MANAGEMENT PLAN (Version 4). Zenodo. <a href="https://doi.org/10.5281/zenodo.8174644">https://doi.org/10.5281/zenodo.8174644</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S. (2023e). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - PROTOCOL. V.5. (<a href="https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5">https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5</a>)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.