Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,561
datasets available to search
ShareScore release 0.9.0
Dataset results
1,561 results for “institution”
Jupyter Usage in Institutions with Coordinates
<p>A dataset with the coordinates of several Institutions which are using Jupyter along with some metadata</p>
Research institutions clustering based on the intensity of academic collaboration
<p>The clustering of research institutions has been conducted using the Louvain modularity algorithm. The Louvain modularity is a state-of-the-art method of identifying communities (clusters) in large networks. Modularity is a value between -1 and 1 that measures the density of edges inside communities to edges outside communities. Optimizing this value results in the best possible grouping of the nodes of a given network.</p> <p>In our exercise, the Louvain methods were applied to identify clusters of institutions within ACM and SSRN networks. In the network, nodes are constituted of institutions, and edges are represented by the intensity of research collaboration measured by number of papers co-authored by authors affiliated with the institutions.</p> <p>As an example, including the paper: <em>Fast unfolding of communities in large networks</em>, written by V. D. Blondel (Universite Catholique de Louvain), J-L. Guillaume (Universite Pierre et Marie Curie), R. Lambiotte (Imperial College London) and Etienne Lefebvre (Universite Catholique de Louvain) would impact the number of edges in our analysis in the following way:</p> <p>“Universite catholique de Louvain” ⇔ “Imperial College London” =+1</p> <p>“Universite catholique de Louvain” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>“Imperial College London” ⇔ “Universite Pierre et Marie Curie” =+1</p> <p>In our largest network we analyse 5362 institution nodes with 147 482 edges. The number of identified clusters highly depends on the resolution parameter. Resolution is a parameter for the Louvain community detection algorithm that affects the size of the recovered clusters. Smaller resolutions recover smaller, and therefore a larger number of clusters, and conversely, larger values recover clusters containing more data points. In all clusterizations, we have used a default resolution (1.0) tuned in the popular Gephi software for network analysis. Resolutions equal to one result in a moderate number of clusters, characterised by satisfactory statistical distribution. </p> <p><strong>Source:</strong></p> <p>- Association for Computing Machinery (ACM)</p> <p>Characteristics of the ACM Data Set following geographical classification</p> <p>Number of institutions: 5477</p> <p>Number of papers: 674684</p> <p>Number of countries: 122</p> <p>Years: 2011-2018</p> <p>As ACM contains publications across various areas of computer science, a more in-depth analysis requires the classification of papers into fields of interests. During the analysis, we looked at 3 wide areas:</p> <ul> <li> <p>Artificial intelligence and machine learning</p> </li> <li> <p>Technology (hardware, emerging technologies, infrastructure)</p> </li> <li> <p>Social issues</p> </li> </ul> <p>The 3 categories were set following expert analysis of the 1000 most frequent keywords in the dataset. If a term from the following list appeared among the paper’s keywords, the paper was assigned to that group, allowing a paper to assign to more than one group.</p> <p><strong>Files:</strong></p> <p>mod_ai.csv (based on keywords related to artificial intelligence)</p> <p>mod_tech.csv (based on keywords related to technologies)</p> <p>mod_soc.csv (based on keywords related to social issues)</p> <p>mod_all.csv (based on all papers)</p> <p> </p>
Observed and WRF-simulated air temperature and wind speed at the Czech Hydrometeorological Institute weather stations Lučina, Lysá hora and Olomouc
<p>The dataset contains two csv files with observed 2-m air temperature and 10-m wind speed data at Lučina, Lysá hora and Olomouc meteorological stations in the Czech Republic and analogical time series produced by the Weather Research and Forecasting (WRF) model. The dataset covers a period of 27 October 2010, 01:00 UTC to 01 November 2010, 00:00 UTC. WRF output is given for three model configurations:</p> <p>1) QNSE boundary layer scheme</p> <p>2) 3DTKE boundary layer scheme with Revised MM5 surface layer scheme</p> <p>3) 3DTKE boundary layer scheme with MYNN surface layer scheme</p>
answered questionnaire to Bachelor Thesis "Wie sinnvoll ist die Ergänzung des Resource Discovery Systems an der Bibliothek des Max-Planck-Instituts für evolutionäre Anthropologie durch einen zusätzlichen, externen Index?"
<p>The dataset contains the answers that were given in the online questionnaire that was conducted as part of the Bachelor Thesis "Wie sinnvoll ist die Ergänzung des Resource Discovery Systems an der Bibliothek des Max-Planck-Instituts für evolutionäre Anthropologie durch einen zusätzlichen, externen Index?"</p> <p>The questionnaire and further information can be found in the Bachelor Thesis, which is linked uner Related Works.</p>
Data and code for "Carbon neutrality should not be the end goal: Lessons for institutional climate action from U.S. higher education"
<p>Code and data for the paper "Carbon neutrality should not be the end goal: Lessons for institutional climate action from U.S. higher education"</p> <p>File descriptions:</p> <p>'HEI_analysis_OneEarth.Rmd' is the code with improved annotation and colorblind-friendly figures.</p> <p>All other data files are provided as excel and csv for convenience.</p> <p>'working_master_data' contains data from the Second Nature reporting platform on emissions by category for each institution analyzed in the paper (measured in metric tons). All adjustments necessary to fill in the data gaps in this file are documented at the beginning of 'HEI_analysis'.</p> <p>'offsets' contains data on the type(s) of offsets purchased by each school in their carbon neutral year (measured in metric tons). This data was assembled from a variety of sources which are documented at the beginning of 'HEI_analysis'.</p> <p>'carbon_neutral_years' contains yearly counts of higher education neutrality goals that were reported to Second Nature as of November 2020.</p>
A census of research software in 171 academic institutional repositories.
<p> A dataset of metadata for 171 UK academic institutional repositories, including a census of research software contained.</p> <table> <tbody> <tr> <td><strong>URL</strong></td> <td>The OAI url</td> </tr> <tr> <td><strong>id</strong></td> <td>CORE Identifier</td> </tr> <tr> <td><strong>openDoarId</strong></td> <td>Open DOAR identifier</td> </tr> <tr> <td><strong>name</strong></td> <td>Name of repository</td> </tr> <tr> <td><strong>Russell_member</strong></td> <td>If the university is a member of the Russell Group of research intensive universities</td> </tr> <tr> <td><strong>RSE_group</strong></td> <td>If an RSE group is present (based on Soc of RSE data)</td> </tr> <tr> <td><strong>email</strong></td> <td>Redacted</td> </tr> <tr> <td><strong>uri</strong></td> <td>Not used</td> </tr> <tr> <td><strong>uni_sld</strong></td> <td>Second level domain (the part of the url between . And .ac.uk</td> </tr> <tr> <td><strong>homepageUrl</strong></td> <td>University website</td> </tr> <tr> <td><strong>source</strong></td> <td>Not used</td> </tr> <tr> <td><strong>ris_software</strong></td> <td>the Research Information System software used</td> </tr> <tr> <td><strong>ris_software_enum</strong></td> <td>Resolve ris_software into similar types (e.g. Eprints 3, EPrints3.3.16 both equal eprints)</td> </tr> <tr> <td><strong>metadataFormat</strong></td> <td>the protocol used for metadata</td> </tr> <tr> <td><strong>createdDate</strong></td> <td>Repository creation date</td> </tr> <tr> <td><strong>location</strong></td> <td>location of university</td> </tr> <tr> <td><strong>logo</strong></td> <td>University logo (resolves in error)</td> </tr> <tr> <td><strong>type</strong></td> <td>Only = Repository for this dataset. Can be = journal etc.</td> </tr> <tr> <td><strong>stats</strong></td> <td>Not used</td> </tr> <tr> <td><strong>contains_software_set</strong></td> <td>Whether the OAI-PMH software set is present in the repository.</td> </tr> <tr> <td><strong>Num_sw_records</strong></td> <td>The response of the OAI-PMH query for software (erroneous as discussed in paper)</td> </tr> <tr> <td><strong>Error</strong></td> <td>The category of error returned by the experiment’s OAI-PMH queries (see paper)</td> </tr> <tr> <td><strong>Manual_Num_sw_records</strong></td> <td>The true amount of software contained in the repository as found by a manual exhaustive search of each university website</td> </tr> <tr> <td><strong>Category</strong></td> <td>Whether the repository (a) contains software; (b) can contain software, but doesn’t yet; (c) has no separate type of research output called software or similar</td> </tr> </tbody> </table>
Metrics and peer review agreement at the institutional level - Data
<p>This data is released to accompany the paper:</p> <p>Traag, VA, Malgarini, M and Sarlo, S (2020) Metrics and peer review agreement at the institutional level.</p>
Global Biodiversity Information Facility (GBIF): an exhaustive list of gbif record ids, dataset keys, and their associated Occurrence IDs, Institution Code, Collection Codes and Catalog Numbers. hash://sha256/ea88f03a7bfd1ba853fdbea3203d54ab81ac3cdc8e8da7c96bbbba9c4b05d933 hash://md5/c49fe34785354847b37ea4509261e130
<p>The Global Biodiversity Information Facility (GBIF) indexes thousands of biodiversity datasets from Natural History Collections, citizen science initiatives (e.g., iNaturalist, eBird), and other sources. As part of the index process, GBIF associates at least two identifiers with indexed records: a record id (aka gbifID) and a dataset id (aka dataset key). These ids are central to do lookup, reference data, and package interpreted data products.</p> <p>This publication contains an exhaustive list of GBIF IDs and ids associated by their data providers as derived from:</p> <p>GBIF.org (01 March 2023) GBIF Occurrence Download https://doi.org/10.15468/dl.pk3trq</p> <p>The resource (size: ~260GB) provided by GBIF had content id hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 and was used to generate the resource included in this publication using</p> <pre><code class="language-bash">preston cat 'zip:hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97!/0015281-230224095556074.csv'\ | cut -f 1,2,3,37,38,39\ | gzip\ > gbifid.tsv.gz </code></pre> <p>with the content id of gbifid.tsv.gz (size: ~35GB) being hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8 .</p> <p>the first 10 lines of gbifid.tsv.gz as extracted via</p> <pre><code>preston cat --remote https://zenodo.org/record/7789866/files,https://linker.bio hash://sha256/a339e32e10edaad585f61f2ded06cbb23e0618c65a6360db18d7d729054940a8\ | gunzip\ | head</code></pre> <p>are:</p> <pre><code>gbifID datasetKey occurrenceID institutionCode collectionCode catalogNumber 2997162320 c71c8000-9fc7-422c-804a-ce6abe751771 3399442 CEPEC CEPEC CEPEC00109669 2997162309 c71c8000-9fc7-422c-804a-ce6abe751771 2733085 CEPEC CEPEC CEPEC00000818 2997162317 c71c8000-9fc7-422c-804a-ce6abe751771 2733086 CEPEC CEPEC CEPEC00000888 2997162313 c71c8000-9fc7-422c-804a-ce6abe751771 3399443 CEPEC CEPEC CEPEC00109744 2997162306 c71c8000-9fc7-422c-804a-ce6abe751771 2733087 CEPEC CEPEC CEPEC00000889 2997162316 c71c8000-9fc7-422c-804a-ce6abe751771 3399440 CEPEC CEPEC CEPEC00109605 2997162324 c71c8000-9fc7-422c-804a-ce6abe751771 2733088 CEPEC CEPEC CEPEC00000890 2997162308 c71c8000-9fc7-422c-804a-ce6abe751771 3399441 CEPEC CEPEC CEPEC00109615 2997162303 c71c8000-9fc7-422c-804a-ce6abe751771 2733089 CEPEC CEPEC CEPEC00000891</code></pre> <p>Note that at time of writing, the html resource associated with the occurrence id 2997162320, and data set key c71c8000-9fc7-422c-804a-ce6abe751771 (extracted from of the first data row example above) are available via:</p> <p>https://gbif.org/occurrence/2997162320</p> <p>and</p> <p>https://gbif.org/dataset/c71c8000-9fc7-422c-804a-ce6abe751771</p> <p>respectively.</p> <p>This resource was initially created to help integrate with Bionomia (https://bionomia.net) to help associate people identifiers provided by bionomia to their original records via their GBIF ids. Bionomia re-uses GBIF records ids as a way to define links between records and the people (e.g., curators, collectors, identifiers) that worked on them. </p> <p>In other words, this resource provides a versioned translation table from the GBIF data universe (as defined by GBIF record ids, and dataset keys) to the data collections that exist (and evolve) independent of it. </p> <p>Note that the resource identified by hash://sha256/c8bac8acb28c8524c53589b3a40e322dbbbdadf5689fef2e20266fbf6ddf6b97 was not included in this publication it was too big (260GB) to fit. You may be able to retrieve the resource from its original location at https://api.gbif.org/v1/occurrence/download/request/0015281-230224095556074.zip .</p>
Two years of CO2 CO and CH4 from The Cyprus Institute at Nicosia, Cyprus
<p>Two years of carbon dioxide (CO2), carbon monoxide (CO) and methane (CH4) concentration measurements, were performed for the first time in the city of Nicosia, Cyprus from 11/02/2020 to 07/09/2022.</p> <p>The dataset is generated from three Picarro G2401s. Property of LSCE (187) and CYI (1172 &1173). They were consecutively installed at the Cyprus Institute, on top of the NTL building, in Nicosia residential area. The dataset was processed by LSCE at Gif-sur-Yvette in France and calibrated against a World Meteorological Organization (WMO) reference scale. </p> <p>The Eastern Mediterranean and Middle East (EMME) region, with its population of more than 400 million, is identified as one of the primary climate “hot spots” worldwide, experiencing adverse impacts ranging from extreme weather events to poor air quality. Projections show that these phenomena are expected to further exacerbate in the coming decades. At central position, lies Cyprus, an island country that receives long-range transported pollution from various anthropogenic and natural sources.</p> <p> </p>
Summary of Input from Stakeholders and other institutions involved in dynamic force applications
<p>Survey data used to create Deliverable 1 of ComTraForce project "Roadmap detailing the future requirements for improved force transfer standards and associated calibration methods for force testing machines taking into account realistic uncertainties".</p>
A dataset of metadata for UK academic institutional repositories, including a census of research software contained.
<p>A dataset of metadata for UK academic institutional repositories, including a census of research software contained.</p> <table> <tbody> <tr> <td><strong>URL</strong></td> <td>The OAI url</td> </tr> <tr> <td><strong>id</strong></td> <td>CORE Identifier</td> </tr> <tr> <td><strong>openDoarId</strong></td> <td>Open DOAR identifier</td> </tr> <tr> <td><strong>name</strong></td> <td>Name of repository</td> </tr> <tr> <td><strong>Russell_member</strong></td> <td>If the university is a member of the Russell Group of research intensive universities</td> </tr> <tr> <td><strong>RSE_group</strong></td> <td>If an RSE group is present (based on Soc of RSE data)</td> </tr> <tr> <td><strong>email</strong></td> <td>Redacted</td> </tr> <tr> <td><strong>uri</strong></td> <td>Not used</td> </tr> <tr> <td><strong>uni_sld</strong></td> <td>Second level domain (the part of the url between . And .ac.uk</td> </tr> <tr> <td><strong>homepageUrl</strong></td> <td>University website</td> </tr> <tr> <td><strong>source</strong></td> <td>Not used</td> </tr> <tr> <td><strong>ris_software</strong></td> <td>the Research Information System software used</td> </tr> <tr> <td><strong>ris_software_enum</strong></td> <td>Resolve ris_software into similar types (e.g. Eprints 3, EPrints3.3.16 both equal eprints)</td> </tr> <tr> <td><strong>metadataFormat</strong></td> <td>the protocol used for metadata</td> </tr> <tr> <td><strong>createdDate</strong></td> <td>Repository creation date</td> </tr> <tr> <td><strong>location</strong></td> <td>location of university</td> </tr> <tr> <td><strong>logo</strong></td> <td>University logo (resolves in error)</td> </tr> <tr> <td><strong>type</strong></td> <td>Only = Repository for this dataset. Can be = journal etc.</td> </tr> <tr> <td><strong>stats</strong></td> <td>Not used</td> </tr> <tr> <td><strong>contains_software_set</strong></td> <td>Whether the OAI-PMH software set is present in the repository.</td> </tr> <tr> <td><strong>Num_sw_records</strong></td> <td>The response of the OAI-PMH query for software (erroneous as discussed in paper)</td> </tr> <tr> <td><strong>Error</strong></td> <td>The category of error returned by the experiment’s OAI-PMH queries (see paper)</td> </tr> <tr> <td><strong>Manual_Num_sw_records</strong></td> <td>The true amount of software contained in the repository as found by a manual exhaustive search of each university website</td> </tr> <tr> <td><strong>Category</strong></td> <td>Whether the repository (a) contains software; (b) can contain software, but doesn’t yet; (c) has no separate type of research output called software or similar</td> </tr> </tbody> </table> <p> </p>
Higher Education Institutions in Poland Dataset
<p><strong>Higher Education Institutions in Poland Dataset</strong></p> <p>This repository contains a dataset of higher education institutions in Poland. The dataset comprises 131 public higher education institutions and 216 private higher education institutions in Poland. The data was collected on 24/11/2022. <br> This dataset was compiled in response to a cybersecurity investigation of Poland's higher education institutions' websites [1]. The data is being made publicly available to promote open science principles [2].</p> <p><strong>Data</strong></p> <p>The data includes the following fields for each institution:</p> <ul> <li>Id: A unique identifier assigned to each institution.</li> <li>Region: The federal state in which the institution is located.</li> <li>Name: The original name of the institution in Polish.</li> <li>Name_EN: The international name of the institution in English.</li> <li>Category: Indicates whether the institution is public or private.</li> <li>Url: The website of the institution.</li> </ul> <p><strong>Methodology</strong></p> <p>The dataset was compiled using data from two primary sources:</p> <ul> <li>Public Higher Education Institutions: Data was sourced from the official website of the Ministry of Education and Science of Poland [3].</li> <li>Private Higher Education Institutions: Data was obtained from the RAD-on system, which is part of the Integrated Information Network on Science and Higher Education [4].</li> </ul> <p>For the international names in English, the following methodology was employed:</p> <p>Both Polish and English names were retained for each institution. This decision was based on the fact that some universities do not have their English versions available in official sources.</p> <p>English names were primarily sourced from:</p> <ul> <li>The Polish National Agency for Academic Exchange's official document [5].</li> <li>The website Studies in English [6].</li> <li>Official websites of the respective Higher Education Institutions.</li> </ul> <p>In instances where English names were not readily available from the aforementioned sources, the GPT-3.5 model was employed to propose suitable names. These proposed names are distinctly marked in blue within the dataset file (hei_poland_en.xls).</p> <p><strong>Usage</strong></p> <p>This data is available under the Creative Commons Zero (CC0) license and can be used for academic research purposes. We encourage the sharing of knowledge and the advancement of research in this field by adhering to open science principles [2].</p> <p>If you use this data in your research, please cite the source and include a link to this repository. To properly attribute this data, please use the following DOI:<br> <strong>10.5281/zenodo.8333573</strong></p> <p><strong>Contribution</strong></p> <p>If you have any updates or corrections to the data, please feel free to open a pull request or contact us directly. Let's work together to keep this data accurate and up-to-date.</p> <p><strong>Acknowledgment</strong></p> <p>We would like to express our gratitude to the Ministry of Education and Science of Poland and the RAD-on system for providing the information used in this dataset.</p> <p>We would like to acknowledge the support of the Norte Portugal Regional Operational Programme (NORTE 2020), under the PORTUGAL 2020 Partnership Agreement, through the European Regional Development Fund (ERDF), within the project "Cybers SeC IP" (NORTE-01-0145-FEDER-000044). This study was also developed as part of the Master in Cybersecurity Program at the Polytechnic University of Viana do Castelo, Portugal.</p> <p><strong>References</strong></p> <ol> <li>Pending.</li> <li>S. Bezjak, A. Clyburne-Sherin, P. Conzett, P. Fernandes, E. Görögh, K. Helbig, B. Kramer, I. Labastida, K. Niemeyer, F. Psomopoulos, T. Ross-Hellauer, R. Schneider, J. Tennant, E. Verbakel, H. Brinken, and L. Heller, Open Science Training Handbook. Zenodo, Apr. 2018. [Online]. Available: [<a href="https://doi.org/10.5281/zenodo.1212496">https://doi.org/10.5281/zenodo.1212496</a>]</li> <li>Ministry of Education and Science of Poland. "Wykaz uczelni publicznych nadzorowanych przez Ministra właściwego ds. szkolnictwa wyższego - publiczne uczelnie akademickie." Nov 2022. [Online]. Available: <a href="https://www.gov.pl/web/edukacja-i-nauka/wykaz-uczelni-publicznych-nadzorowanych-przez-ministra-wlasciwego-ds-szkolnictwa-wyzszego-publiczne-uczelnie-akademickie">https://www.gov.pl/web/edukacja-i-nauka/wykaz-uczelni-publicznych-nadzorowanych-przez-ministra-wlasciwego-ds-szkolnictwa-wyzszego-publiczne-uczelnie-akademickie</a></li> <li>RAD-on System. "Dane instytucji systemu szkolnictwa wyższego i nauki." Nov 2022. [Online]. Available: <a href="https://radon.nauka.gov.pl/dane/instytucje-systemu-szkolnictwa-wyzszego-i-nauki">https://radon.nauka.gov.pl/dane/instytucje-systemu-szkolnictwa-wyzszego-i-nauki</a></li> <li>Polish National Agency for Academic Exchange. "List of the university-type HEIs." 2023. [Online]. Available: <a href="https://nawa.gov.pl/images/Aktualnosci/2023/Att.-2.-List-of-the-university-type-HEIs.pdf">https://nawa.gov.pl/images/Aktualnosci/2023/Att.-2.-List-of-the-university-type-HEIs.pdf</a></li> <li>Studies in English. [Online]. Available: <a href="http://www.studies-in-english.pl/">www.studies-in-english.pl</a></li> </ol>
Final chlorophyll and temperature measurements at 10m depth, offshore of Dana Point, California as part of an Ocean Institute time series, 2006 - 2025.
Time series of chlorophyll and temperature at 10m offshore of Dana Point, California. Measurements made as part of education/outreach programs involving students and teachers in oceanographic sampling and analytic sensitivity to time series. Cruises are conducted twice, monthly (or other) where sampling is performed by students assisted by technicians and supervisors.
Institutional Dimensions of Restoring Everglades Water Quality - Social Capital Analysis (FCE), Florida Everglades Agricultural Area from September 2014 to July 2015
These data were compiled through the Institutional Dimensions of Restoring Everglades Water Quality research project. One of the manuscripts generated by this project focused on the social capital dynamics in the Everglades Agricultural Area. These data represent different social capital aspects reflected by the responses of interview subjects. The purpose of analyzing social capital was to explore why and how farmers cooperated given that the state law, the Everglades Forever Act, which required the adoption of best management practices, relied on shared compliance for farmers to improve water quality. The study sought to undrestand how different aspects of social capital (broadly pro-social norms of reciprocity and trust) either encouraged or discouraged farmers to adopt BMPs effectively.
Annual summaries of daily climatological observations from the National Weather Service weather station at the UGA Marine Institute on Sapelo Island, Georgia for 1958 to 2004
Daily summaries of climatological observations from the National Weather Service weather station on Sapelo Island, Georgia, were obtained from the NOAA National Climatic Data Center (http://www.ncdc.noaa.gov/) covering the period 1958 through 2004. Records for incomplete years with more than one month of missing observations were deleted (i.e. 1964, 1969-1971), then missing values of daily minimum, maximum and mean temperature were estimated by cubic spline interpolation to fill in data gaps of five or fewer consecutive days. Annual summary statistics were then calculated for daily minimum, maximum and mean air temperature and total precipitation.
A set of six databases used in a study of the biogeography of Greater Caribbean reef fishes entitled: Comparing biodiversity databases: Greater Caribbean reef-fishes as a case study Iliana Chollett1, D. Ross Robertson2 1 Sea Cottage, Louisburgh, Co. Mayo, Ireland 2 Smithsonian Tropical Research Institute, Balboa, Panamá
<p><strong>A set of six databases used in a study of the biogeography of Greater Caribbean reef fishes entitled:</strong></p> <p><strong><em> </em></strong></p> <p><strong><em>Comparing biodiversity databases: Greater Caribbean reef-fishes as a case study</em></strong></p> <p> </p> <p> Iliana Chollett, D. Ross Robertson</p> <p><strong> </strong></p> <p><strong> </strong></p> <p><strong>Database Authors: D Ross Robertson and Ernesto Peña, Smithsonian Tropical Research Institute, Panamá</strong></p> <p><strong> </strong></p> <p>This set of six databases contains georeferenced location records from six sources as described below.These six sources provided georeferenced records of occurrence of fishes found in the Greater Caribbean study area (6-33<sup>0</sup> N, 57-100<sup>0</sup> W). Each occurrence record consists of a species name and associated latitude and longitude. Databases included in the comparisons made here are from five major online aggregators. Since their content overlaps to some extent, and OBIS, iDigBio and FishNet collaborate with GBIF, their data might be expected to produce similar biogeographic patterns. STRI includes a curated compendium of data from those five aggregators, enriched with data from many additional sources.</p> <p> </p> <p>Only reef-associated fish species were included in the present analysis. These mostly represent demersal species known to occur on hard bottoms (coral, rock and oyster substrata), but also include species living on rubble, sand and vegetated bottoms within and around the immediate fringes of reefs, and pelagic species regularly found on reefs. All exotic and non-resident species and species other than reef-associated fishes were excluded from all databases prior to comparisons. Non-residents were defined as otherwise widespread species only rarely seen in the study area. Shore-fishes, including what are generally regarded as reef fishes, include those found in the waters of continental and insular shelves, i.e. between 0-200m. Reef-fish assemblages dominated by shallow-water taxa extend down to that depth in the study area (Baldwin <em>et al.</em> 2018). We used the shelf edge as a breakpoint and excluded records in areas deeper than 200m, identifying those areas using the General Bathymetric Chart of the Oceans (Kapoor, 1981; GEBCO Compilation Group, 2019).</p> <p> </p> <p>Before the analyses, for all databases, duplicate records were deleted. Subsequently, records in the Pacific or on land were deleted. We used the Global Self-consistent, Hierarchical, High-resolution Geography Database (Wessel & Smith, 1996) to identify these areas. The spatial distribution of species-records in each database is shown in Figure 1 of the publication.</p> <p><strong> </strong></p> <p><strong>Global Biodiversity Information Facility </strong>(GBIF, https://www.gbif.org/): GBIF is an international network and research infrastructure aimed at providing open access to data about all types of life on earth. GBIF works through participant nodes using common standards and open-source tools that enable them to share information. Data from among the 49,000+ datasets hosted by GBIF that were used here range from those on museum specimens collected since the 18th century, to published scientific checklists, to curated local checklists produced by trained science sources such as the Atlantic and Gulf Rapid Assessment Program (https://www.agrra.org/),to geotagged smartphone photos (that act as vouchers allowing verification) shared by amateur and scientific naturalists through iNaturalist (https://www.inaturalist.org/), to unvouchered, unverified and unverifiable observation records from untrained divers such as those contributing to DiveBoard (http://www.diveboard.com). GBIF data are standardized in Darwin Core format. GBIF data were obtained from a polygon of the region of study and subject to taxonomic review and selection after downloading. GBIF data were obtained from a polygon of the study area and subject to taxonomic review after downloading (accessed through the GBIF portal, https://www.gbif.org/, on or about 2019-05-19).</p> <p> </p> <p><strong>Ocean Biogeographic Information System</strong> (OBIS, <a href="https://obis.org/">https://obis.org/</a>): OBIS is a global open-access data and information clearing-house on marine biodiversity (OBIS, 2019) that was adopted as a project of the Intergovernmental Oceanographic Data and Information Exchange of the Intergovernmental Commission of UNESCO . Its range of sources is similar to that of GBIF. OBIS hosts data from organizations or programs that join it as one of 13 “nodes”, and harvest the data from the IPT (Integrated Publishing Toolkit), where providers publish their data. The IPT is developed and maintained by the GBIF, and OBIS is a major contributor of marine data to GBIF. Data are standardized in Darwin Core format. OBIS data were obtained for the region of study by downloading data on each family, then retaining only data inside the study area, which were then subject to taxonomic review and selection (accessed through the OBIS portal, https://obis.org/, on or about 2019-05-19).</p> <p> </p> <p><strong>Integrated Digitized Biocollections</strong> (iDigBio, https://portal.idigbio.org/portal/search): iDigBio is sponsored by the a US National Science Foundation and run by the University of Florida that provides digital data from public, non-federal, US collections. Data are standardized in a Darwin Core format, and provided “as is”. IDigBio joined the GBIF network in 2017. IDigBio records were downloaded from a polygon of the region of study and subject to taxonomic review and selection (accessed through the iDigBio portal, https://portal.idigbio.org/portal/search, on or about 2019-05-19).</p> <p> </p> <p><strong>FishNet2 </strong>(http://www.fishnet2.net/): FishNet2 is a collaborative effort that aggregates data on fish collections around the world to share and distribute data on specimen holdings from ~75 museums, universities and other institutions. FishNet2 distributes data in Darwin Core, and data are provided “as is”. FishNet2 is part of the network VerNet, which has contributed to GBIF since 2013 and became part of IDigBio in 2016. While FishNet2 has made substantial efforts to georeference location-record data it hosts, many hosted records still lack georeferencing. FishNet2 data were obtained from a polygon of the study area and subject to taxonomic review after downloading (accessed through the Fishnet2 Portal, www.fishnet2.org, 2019-05-19).</p> <p> </p> <p><strong>FishBase</strong> (<a href="http://www.fishbase.org/">http://www.fishbase.org</a>): FishBase is a global biodiversity information system supervise by a consortium of nine non-USA international institutions, and hosts data on fin fishes and elasmobranchs (Froese & Pauly, 2009). Information presented in FishBase is extracted from the scientific literature, reports and museum or aggregator (GBIF) databases, and standardized by a team of specialists. Data from Fishbase were downloaded for the following ecosystems: Caribbean Sea, Gulf of Mexico, Southeast U.S. Continental Shelf, Atlantic Ocean, Sargasso Sea and Bermuda, and subject to taxonomic review and selection after downloading (2019-05-19).</p> <p> </p> <p><strong>Smithsonian Tropical Research Institute</strong> (STRI; <a href="https://biogeodb.stri.si.edu/caribbean/en/pages">https://biogeodb.stri.si.edu/caribbean/en/pages</a>): The STRI database was compiled by DRR and Ernesto Peña at STRI’s Naos Marine Laboratory, and represents about 15 years accumulation of curated data (see below) from the following sources: data downloaded at roughly two year intervals from the five aggregators; data from online databases of various museums that supply aggregators (data directly downloaded from a museum sometimes differs from that available in an aggregator from the same museum), including the Swedish Museum of Natural History, the American Museum of Natural History, the Natural History Museum of Denmark, the Gulf Coast Research Laboratory, the Colombian Museum of Natural Marine History, the United States National Museum, and the United States Geological Survey; data from national aggregators of Colombia (Sistema de Información Sobre Biodiversidad de Colombia (https://sibcolombia.net/), and Sistema de Información Ambiental Marina de Colombia, https://siam.invemar.org.co/), Mexico (La Comisión Nacional para el Conocimiento y Uso de la Biodiversidad, CONABIO; http://www.conabio.gob.mx/informacion/gis/), and Costa Rica (Museo de Zoologia de la Universidad de Costa Rica, http://museo.biologia.ucr.ac.cr/); verified (by DRR) underwater photographs of fishes taken at known locations; peer reviewed publications containing location information (species descriptions; taxonomic revisions of species, genera and families; regional and local checklists); fisheries reports; digital tagging data for species such as elasmobranchs; diving surveys and collections of local faunas by DRR (e.g. Robertson et al. 2019). In addition selected data from two sources that collect species lists at sites scattered throughout the Greater Caribbean are incorporated: from the Atlantic and Gulf Rapid Reef Assessment program (AGRRA, https://www.agrra.org/: Kramer & Lang, 2003) and from trained citizen scientists who contribute data on fishes to the Reef Environmental Education Foundation’s database (REEF: Pattengill-Semmens & Semmens, 2003). The bibliographic module (https://biogeodb.stri.si.edu/caribbean/en/library) of Robertson & VanTassel (2019) contains ~1700 publications linked to species names, among them the publications from which location data were extracted.</p> <p> </p> <p>Data from the aggregators is presented “as is” and the aggregators themselves do not do data curation. Duplicates (and occasionally triplicates and quaduplicates) of the same museum record often are included from multiple sources (e.g. the original museum source, derivative checklists, an aggregator), sometimes with slightly different georeferenced coordinates. Data available in one year may subsequently disappear from an aggregator, and different data may be available for the same species under different names (e.g. the old and new names when a species is reassigned to another genus). Errors, sometimes large errors (Robertson, 2008), are common in aggregator data, from museums as well as other sources, and longstanding errors can seem to take on a perpetual existence. For example the damselfish <em>Abudefduf saxatilis </em>is a common and widespread inhabitant of tropical reefs on both sides of the Atlantic, but does not naturally occur outside that ocean. Despite the fact that its taxonomic status and range were resolved ~30 y ago (e.g. see Allen, 1991) museum data presented by the all five aggregators that contributed to the multi-source database used in this study currently (December 10, 2019) show large numbers of records of this species throughout the entire tropical Indo-Pacific, as well as across its native range in the Atlantic. Since many of the databases accumulating on aggregators are derivative (lists derived from records and from other derivative lists) it will become increasingly difficult to eliminate such errors as corrections to data in primary sources do not automatically propagate through the chain of usage by different databases. Due to increasing limitations on resources for taxonomic work, museums themselves have difficulty dealing with errors in specimen identity and location, and old specimens become unidentifiable, specimens never get returned when loaned out, or simply vanish, and entire collections can get destroyed by hurricanes or fires, or get dumped when museums close or experience a major change in mission. Georeferenced location data on fish distributions in the neotropics (and presumably most other areas) hosted by aggregators, particularly GBIF and OBIS, which take data from a broad range of source types, might best be described as messy, and the significant potential for errors in location records and an inability to verify records always needs to be taken into account when incorporating data from aggregators, primary museum sources, and analog sources.</p> <p> </p> <p>Data considered for inclusion in the STRI database were screened as follows to exclude questionable records. Data from two databases hosted by OBIS and GBIF were excluded entirely due to lack of reliability: BioGoMx (https://www.gulfbase.org/project/biodiversity-gulf-mexico-biogomx-database) and Diveboard (http://www.diveboard.com). The only REEF data used were from “expert” REEF recorders on readily identifiable species that are unlikely to be confused with similar species (e.g. data for some genera of sparids, gerreids, labrisomids and gobies that include various sympatric species with very similar appearances, were not used). After data from aggregators and museum sources were combined into a single database duplicate records were filtered out by rounding all records to three decimal places and eliminating duplicates, a process that inevitably deleted some valid records as well as duplicates. The sizes of the databases and abundance of such duplicates precluded individual manual exclusion. Finally, all location data for each species were revised by DRR by examining the distribution of its georeferenced coordinates overlayed on a digital map of the current known distribution range of that species (for such range information see Carpenter & De Angelis, 2002; Ebert <em>et al.</em>, 2013; Last <em>et al.</em>, 2016; Robertson & Van Tassell, 2019; IUCN Redlist species accounts for most species considered here: https://www.iucnredlist.org/search). Such revision took into account any recent modifications to taxonomy and distributions due to new data and new publications, or as a result of discussions between DRR and experts in the taxonomy of particular species or genera. Source information of many individual questionable records provided by aggregators with the hosted data was inspected to try and assess their validity. Records thought likely to be erroneous were deleted. Those included inexplicable records lacking adequate documentation located well outside the known distribution range, and records in unlikely habitats (e.g. on land for marine species; in deep water for shallow-water species). This revision process reduced the number of records by about 30%.</p> <p> </p> <p>Data from the five individual aggregator databases that are used in the comparisons described here were all downloaded from their online portals during May, 2019. However, data from those five aggregators that were incorporated in the STRI database were downloaded in March 2017, with data from other sources described above added to the STRI database intermittently between then and May 2019, when the entire dataset was curated as described above. Hence the five individual aggregator databases analyzed in this study undoubtedly contain data not included in the version of the STRI database used in the present analyses.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <p> </p> <p>Data acquisition and construction of the STRI database was supported by funds from STRI, the Smithsonian Marine Science Network, the Smithsonian Publications Fund, the Smithsonian’s Deep Reef Observation Project, the National Geographic Society, the IUCN Red List program, the Harte Research Institute, and CONABIO. We thank REEF and AGRRA for supplying species-location records, various people for taxonomic and location-record information used to construct that database (principal among them C Baldwin, S Brandl, K Conway B Frable, T Menut, T Munroe, R Robins, L Tornabene, J Van Tassell and B Victor), and hundreds of citizen-scientist submarine photographers whose images (see <a href="https://biogeodb.stri.si.edu/caribbean/en/contributors/citizen_scientists">https://biogeodb.stri.si.edu/caribbean/en/contributors/citizen_scientists</a>) acted as vouchers for location records.</p> <p> </p> <p><strong>References</strong></p> <p><strong> </strong></p> <p>Allen, G.R. (1991) <em>Damselfishes of the World</em>. Mergus, Melle, 271 p.</p> <p>Baldwin, C.C., Tornabene, L. & Robertson, D.R. (2018) Below the mesophotic. <em>Scientific Reports</em>, 8, 4920.</p> <p>Carpenter, K.E. (Ed) (2002) <em>The living marine resources of the Western Central Atlantic.</em> Vols 1-3, FAO, Rome, 2127 p.</p> <p>Ebert, D.A., Fowler, S., Compagno, L. (2013) <em>Sharks of the World: a fully illustrated guide</em>. Wild Nature Press, Plymouth. 528 p.</p> <p>GEBCO Compilation Group (2019) GEBCO 2019 Grid (doi:10.5285/836f016a-33be-6ddc-e053-6c86abc0788e).</p> <p>Kapoor, D.C. (1981) General bathymetric chart of the oceans (GEBCO). <em>Marine Geodesy</em>, 5, 73–80.</p> <p>Kramer, P.R. & Lang, J.C. (2003) Appendix one: The Atlantic and Gulf Rapid Reef Assessment (AGRRA) Protocols: Former Version 2. 2. <em>Atoll Research Bulletin</em>, 496, 611–624.</p> <p>Last, P. R., White, W.A., de Carvalho, M.R., Séret, B., Stehmann, F.W., & Naylor, J.P. (2016). <em>Rays of the World</em>. CSIRO, Clayton. 790 p.</p> <p>Pattengill-Semmens, C.V. & Semmens, B.X. (2003) <em>Conservation and management applications of the reef volunteer fish monitoring program</em>. <em>Coastal Monitoring through Partnerships: Proceedings of the Fifth Symposium on the Environmental Monitoring and Assessment Program (EMAP) Pensacola Beach, FL, U.S.A., April 24–27, 2001</em> (ed. by B.D. Melzian), V. Engle), M. McAlister), S. Sandhu), and L.K. Eads), pp. 43–50. Springer Netherlands, Dordrecht.</p> <p>Robertson, D. R. (2008) Global biogeographic databases on marine fishes: caveat emptor. <em>Diversity and Distributions, 14<strong>,</strong> 891-892</em></p> <p>Robertson, D.R,, Dominguez-Dominguez, O., Lopez Arollo, Y.M., Moreno Mendoza. R., Simoes, N. (2019) Reef-associated fishes from the offshore reefs of western Campeche Bank, Mexico, with a discussion of mangroves and seagrass beds as nursery habitats. <em>Zookeys </em>843: 71-115. <a href="https://doi.org/10.3897/zookeys.843.33873">https://doi.org/10.3897/zookeys.843.33873</a></p> <p>Robertson, D.R & Van Tassell, J. (2019) Shorefishes of the Greater Caribbean: online information system. Version 2.0. <em>Smithsonian Tropical Research Institute, Balboa, Panamá</em>. <a href="https://biogeodb.stri.si.edu/caribbean/en/pages">https://biogeodb.stri.si.edu/caribbean/en/pages</a>.</p> <p>Wessel, P. & Smith, W.H.F. (1996) A global, self-consistent, hierarchical, high-resolution shoreline database. <em>Journal of Geophysical Research: Solid Earth</em>, 101, 8741–8743.</p>
Institutional arrangements regarding Minimum Wage Setting in 195 countries
<p>Most countries in the world have country-level policies concerning their minimum wage-fixing machinery. These policies vary widely, and therefore it becomes important to have adequate classifications of these policies. This paper reviews databases that classify country-level policies for determining minimum wages. Several databases - we found twelve - classify countries according to their minimum wage-fixing mechanisms and the coverage of these mechanisms. The mechanisms indicate whether the minimum wages are set by Law, by Collective Bargaining or any policy in between, the coverage indicates whether the minimum wages cover the entire dependent labour force or only one or more sections within the labour force. The twelve databases vary with respect to the years covered, the countries covered and the characteristics coded. We restricted our analysis to the years 2011 to 2015. The number of countries covered in these databases range from 29 to 189, with 195 countries in total. The merged database reveals that countries are not classified similarly across databases. Between 75% and 93% of the countries apply a statutory minimum wage-fixing mechanism across years and databases. Less than one in ten countries relies solely on minimum wage setting by collective bargaining. In the EU28 plus Norway this percentage is relatively high, but in countries outside Europe it is far below 10%. Two ILO conventions refer to minimum wage-fixing mechanisms. Across years and databases roughly three in five countries that apply a statutory minimum wage-fixing mechanism have signed the oldest Convention (C26), whereas roughly one in three has done so with the most recent Convention (C131). Obviously, many more countries could have signed the Conventions. Only a few countries have signed the Conventions but do not have a statutory minimum wage-fixing mechanism. Among others a few EU28 countries rely solely on collective bargaining for minimum wage setting, and consider that as a national wide fixing mechanism. If countries apply a statutory wage-fixing mechanism, does the minimum wage then cover the entire dependent labour force? Globally, more than half of the countries with a statutory minimum wage apply differentiated minimum wages. Most frequently reported breakdowns are by industry or occupation. Countries with multiple minimum wage rates mimic collective bargaining, particularly when they break down the rates by industry or occupation. The aim of this paper is to generate a Minimum Wage Policies Database (MWPDB) from the merged dataset. Using a set of rules for generating data from the source databases, we indicate for almost half of the 195 countries the presence or absence of a statutory minimum wage for all five years from 2011 to 2015. For 16 countries no valid data is available for any year. Particularly for Europe and South America, MWPDB has satisfactory number of observations, whereas the opposite holds for the small islands in Oceania. The MWPDB results show that approximately nine in ten countries do apply a minimum wage policy, and that this is slightly increasing between 2011 to 2015. </p>
Sudanese Jibbah in UK Institutions
<p><em>22/09/20: dataset updated to include 5 records from the Pitt Rivers Museum</em></p> <p>This dataset was created by Elvira Thomas during the summer of 2020 as part of her<a href="https://www.sussex.ac.uk/study/undergraduate/undergraduate-research/junior-research-associates"> Junior Research Associate Project funded by the University of Sussex</a>. The project is attached to the “<a href="http://makingafricanconnections.org/">Making African Connections</a>” (MAC) Project of which Sussex University and the <a href="https://www.re-museum.co.uk/">Royal Engineers Museum</a> are partners.</p> <p>The project, which was titled “Plundered artefacts and the Colonial Gaze: exhibiting Sudan at the Royal Engineers Museum after Covid-19”, investigates museum display in UK institutions of artefacts which were plundered by colonial British forces at the end of the 19<sup>th</sup> Century. It takes an object-focus on Mahdist <em>jibbah</em>, tunics worn by Ansar (followers of al Mahdi) in conflicts during the Mahdist war in Sudan at the end of the 19<sup>th</sup> C. Many of these <em>jibbah</em> were plundered after the Battle of Omdurman (1898), taken back to Britain and dispersed into private collections and cultural institutions across the UK. Five of these Mahdist <em>jibbah</em> are now held at the Royal Engineers Museum. In order to explore how these important cultural artefacts of Sudanese heritage have been historically framed in triumphalist narratives of empire in UK institutions, the project combined a case study of the museum display at the Royal Engineers Museum with a wider investigation of Mahdist <em>jibbah</em> held in institutions across the UK.</p> <p>This dataset was created as part of the wider investigation of <em>jibbah</em> held in UK institutions. Due to the outbreak of the Covid-19 pandemic which restricted close contact with the <em>jibbah</em> held at REM, this part of the project became a crucial element of research. There were three main goals for this investigation:</p> <ol> <li> <p>to produce an estimate of the amount of Sudanese <em>jibbah</em> existing in UK cultural institutions</p> </li> <li> <p>to locate, record and publish a list of these artefacts to make them knowable to ‘source’ communities or anyone searching for Sudanese heritage in UK institutions</p> </li> <li> <p>to analyse digital catalogue entries (accessibility, data prioritisation and linguistic analysis) of the <em>jibbah</em> held in UK institutions and consider to what extent legacies of colonialism are encountered in the digital display of plundered artefacts held in the UK</p> </li> </ol>
Silene seeds from the laboratories of the Institute of Biophysics (Academy of Sciences) of Brno (Czech Republic)
<p>Seeds of <em>Silene </em>for the analysis of morphology (Martín Gómez et al.) obtained from the laboratories of the Academy of Sciences of Brno (Czech Republic). Photos contains 40 seeds of:</p> <p><em>Silene acutifolia; S. colpophylla; S. conica; S. diclinis; S. dioica; S. gallica; S. italica; S. latifolia; S. noctiflora; S. nutans; S. otites; S. pendula; S. saxifraga; S. schafta; S. tatarica; S. viscosa; S. vulgaris; S. wolgensis; S. zawadzkii</em></p>
BDRC Institutional Affiliation Network Data
<p>These two files correspond to the edge list and the node list used to build the network of institutional affiliations (individuals' positions in institutions) in the <em>Biographical Dictionary of Republican China</em> (BDRC). </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.