Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,025

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

6,025 results for “Science of science”

Learn how ShareScore rates datasets ↗
zenodo44/100

Open Source and Open Science Sustainability Year-long Study - Pseudonymised

<p>Research software is often abandoned or shut down, for one reason or another. While some reasons may be straightforward, e.g. a sole maintainer has moved on, or grant funding has ceased - some projects are able to withstand these barriers and may remain active and maintained despite adversity.</p> <p>This study monitors open source projects over the period of a year, measuring common performance indicators, to see if any indicators are common to projects that remain sustainable and active.</p> <p>This study uses mixed methods:</p> <ol> <li>Initial survey gathers info about the project age, leadership, and GitHub (or other source control) URLs. Participants are asked to add <a href="https://sustainable-open-science-and-software.github.io/readme_notice">a short notice to their readme</a>.</li> <li>After the initial survey, we gathered information about the GitHub projects such as number of contributors, number of PRs, time taken to close/merge these PRs, and issues closed. Some of this info is gathered using scripts,&nbsp; and other parts are gathered manually. An example of a manual metric is the Code of Conduct - while we can programmatically check for the <em>existence</em> of CodeOfConduct.md, we can&rsquo;t easily check for enforcement contacts without manual checks.</li> <li>6 months and 12 months after the initial survey, we send follow up surveys, and in month 12 we re-run the GitHub metrics to compare to month 0.</li> </ol> <p>&nbsp;</p> <p>For more study info see: <a href="https://sustainable-open-science-and-software.github.io/">https://sustainable-open-science-and-software.github.io/</a></p> <p>&nbsp;</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

Making Social Science Research Transparent [Webinar recording]

<p>High-quality data have the potential to be reused in many ways. Archiving and publishing your data properly is at the core of making your data FAIR and will enable both your future self as well as others to get the most out of your data. Recently, more and more scientific journals are implementing open data policies, leading to researchers&#39; dilemmas about where, when and how to publish the data. Consequently, the way that social science research is conducted and disseminated is gradually changing. A crucial element of that change is research transparency. Introduction to the topic took place in the first part of the event.</p> <p>In the second part, panellists presented in-depth the processes, policies and tools implemented for facilitating transparent research in the social sciences. They discussed the processes that need to be in place for an open research cycle, the role of data archives and repositories in sharing research data and materials, tools for reproducing research findings in practice, collaborations between archives and social science journals, and implementing Transparency and Openness Promotion Guidelines in different social science disciplines.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=B-phrIMETGk"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Data management in the social sciences in Macedonia [Webinar recording]

<p>This webinar aimed to introduce social science researchers to the basic principles of data management, including the creation of a Data management plan, which is an important tool for planning the research project.</p> <p>The webinar consisted of three parts. The first part introduced researchers with the basic principles of data management including the benefits of adopting Data management plans (DMPs). The DMP follows the research projects&rsquo; life cycle, starting with the initial phases of Planning and Organization and documentation of research data. This part also included a presentation of best practices for creation of appropriate structure of folders and data files, as well as instructions for their naming, documentation and organization.</p> <p>The second part of the webinar focused on the following three phases of the project life cycle: Data processing, Preservation and Protection. Contemporary social science presumes the respect of high level ethical standards during the handling of research data, in accordance with legal rules and best practices in this area.</p> <p>The last part of the webinar was dedicated to the phases of Publication - familiarizing the researchers with the possibilities of data preservation and publishing; and Data discovery - discussing the ways and means to acquire social science data, including the secondary use of data produced by other researchers.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=n05WTs58CMY"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Research Data Management and data protection in the Social Sciences [Workshop recording]

<p>This online workshop organized by The Austrian Social Science Data Archive (AUSSDA) focused on the Research Data Management basics, Data Management Plans and common data protection issues in the Social Sciences.</p> <p>The first part of the workshop was dedicated to RDM basics and Data Management Plans (DMPs). In many projects, DMPs are mandatory deliverables that need to be submitted at the beginning of a project and are updated throughout the project life cycle. During the workshop, it was explained which aspect funders expect to be part of DMPs in Social Sciences and how researchers can benefit from (writing) these documents.</p> <p>In the second part of the workshop, data protection issues that are common in Social Sciences were addressed and how they can be handled. In particular, differences in the curation of quantitative and qualitative data need in order to comply with data protection regulations in general and AUSSDA deposit guidelines in particular. Presentation on how AUSSDA scans quantitative data for potential data protection violations using STATA and gives participants the opportunity to test the code on their own data and devices.</p> <p>The video is available on<a href="https://www.youtube.com/watch?v=DhiL9J-Iwqg"> the&nbsp;CESSDA Training&nbsp;YouTube channel</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo44/100

Delphi Study: Exploring the Implications of Large Language Models on the Science System

<p><strong>Sample description:</strong> Our target audience consisted of researchers working in the fields of science, technology, and society with a specific interest in Large Language Models (LLMs).</p> <p><strong>Collection method: </strong>Participants were recruited through the professional and personal networks of the authors, as well as the Alexander von Humboldt Institute (HIIG), using a combination of generic emails via LimeSurvey and personal contacts.</p> <p><strong>Description. </strong>The aim of this study was to explore the impact of large language models, specifically ChatGPT, on scholarly practice and academic writing, targeting researchers and experts in the fields of artificial intelligence, science, and technology who publish their research and scientific work. The two-stage Delphi survey sought to identify and assess the potential opportunities and challenges associated with the use of ChatGPT in academic work and scientific writing, with a specific focus on research impact rather than university teaching. Phase 1 yielded 72 responses, while Phase 2 had 52 responses.</p> <p>To conduct our analysis, we developed two distinct codebooks (see Files ChatGPT Delphi Codebook Phase 1.csv and ChatGPT Delphi Codebook Phase 2.csv) for the Delphi study. The first codebook was created by examining approximately half of the responses, extracting relevant information, and generating codes through inductive reasoning. We then categorized and developed subcodes based on these initial codes, assigning them to each participant&#39;s answers using deductive reasoning. For example, when addressing the potential applications of ChatGPT and other language models (LLMs), we identified six subcategories with precise definitions and illustrative examples. The analysis in Phase 1 led to the formulation of ranking questions for Phase 2, focusing on determining the most frequently utilized applications of ChatGPT and other LLMs based on the established codes.</p> <p>During Phase 2, we introduced two additional open-ended questions to explore the impact of ChatGPT and LLMs on the scientific system and society, aiming to envision future scenarios. The analysis of these questions in the second codebook followed a similar approach to Phase 1, including inductive reasoning for code generation and deductive reasoning for assigning codes to the answers. We observed overlapping codes with the Phase 1 codebook and assigned them to the second codebook. Additionally, we noted a shift in the connotation of certain answers from neutral in Phase 1 to being perceived as either positive or negative consequences of ChatGPT and other LLMs. This observation prompted the bifurcation of specific codes to capture the nuanced perspectives. For instance, applications such as reducing administrative tasks initially seen as valuable aids for researchers were sometimes viewed as potential causes for job replacement, implying negative outcomes.</p> <p>For detailed information on the analytical approach employed, including references to these methodologies, please refer to the methodology chapter in the official publication.</p> <p><strong>Content</strong></p> <ol> <li> <p>Questionaire-ChatGPT-Delphi-Phase1-Limesurvey-Export.pdf &ndash; This file file is an exported version of the Phase 1 questionnaire from Limesurvey. It includes the description, socio demographic questions, content questions, and a request for participant naming.</p> </li> <li> <p>Questionaire-ChatGPT-Delphi-Phase2-Limesurvey-Export.pdf &ndash; This file file is an exported version of the Phase 2 questionnaire from Limesurvey. It includes the description, socio demographic questions, content questions, and a request for participant naming.</p> </li> <li> <p>ChatGPT Delphi - Results Phase 1.pdf &ndash; This file contains the responses and corresponding questions from Phase 1 of the Delphi study. The responses provided by the participants are in the form of open-ended answers. As part of this publication, we have ensured the anonymity of the participants.</p> </li> <li> <p>ChatGPT Delphi - Results Phase 2. pdf &ndash; This file contains the responses and corresponding questions from Phase 2 of the Delphi study.&nbsp; It encompasses the ranking answers provided by the participants, as well as two open-ended answers. To maintain anonymity consistently, all participants have been anonymized again in this publication of our results.</p> </li> <li> <p>ChatGPT Delphi Codebook Phase 1.pdf &ndash; This file contains the Phase 1 codebook, which presents the primary codes, their respective subcodes, detailed definitions, and noteworthy examples.</p> </li> <li> <p>ChatGPT Delphi Codebook Phase 2.pdf &ndash; This file contains the Phase 1 codebook, which provides a comprehensive overview of the primary codes within the given scenario. It includes their corresponding subcodes, detailed definitions, and notable examples to enhance understanding and interpretation.</p> </li> </ol>

opencc-by-4.0Jun 2023View details →
zenodo44/100

CS3MESH4EOSC Final Event Science Mesh - Unlocking Open Science and Collaborative Research Landscape

<p>The recap video of CS3MESH4EOSC final event. The CS3MESH4EOSC final event, took place on 22 June 2023, at the EGI Conference in Poznan (Poland), showcased how the Science Mesh is contributing to an easier and more robust open science across Europe, thanks to novel approaches for data sharing and synchronisation.</p> <p><strong>The first half of the event</strong> will count with live demonstrations, where each data service from the Science Mesh will be presented from a user-perspective point of view. Event attendees will get practical information on how they can join the Science Mesh as a researcher, a software developer or a service provider. The event will also bring together representatives of Science Mesh use cases, who will explain how Science Mesh is making a difference in their lives, thanks to easier data sharing and synchronisation. A panel discussion with representatives of different sectors, from research to industry and education, will discuss the most urgent trends &amp; priorities for cross-border science collaboration between different sciences.</p> <p><strong>The second half of the event</strong> will be focused on the technical novelties within the Science technical foundation. A series of demonstrations will be presented, followed by a panel discussion, with representatives of CS3MESH4EOSC members that are part of the EOSC Task Forces, on how the Science Mesh contributes to EOSC&rsquo;s success and the overall EOSC Strategic Research and Innovation Agenda (SRIA).</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Different facets of the same niche: integrating citizen science and scientific survey data to predict biological invasion risk under multiple global change drivers

<p>Raw data (occurrences and&nbsp;environmental predictors) used in&nbsp;the manuscript &quot;Different facets of the same niche: integrating citizen&nbsp;science&nbsp;and&nbsp;scientific survey&nbsp;data&nbsp;to&nbsp;predict&nbsp;biological&nbsp;invasion risk under&nbsp;multiple&nbsp;global change&nbsp;drivers&quot;</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESULTS DATASET (with Mega Journals)

<p>The dataset contains all the data produced running the research software for the study:&quot;Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta&quot;.</p> <p>Disclaimer: these results are not considered to be representative, because we have fount that Mega Journals skewed significantly some of the data. The result datasets without Mega Journals are published <a href="https://zenodo.org/record/8249907">here</a>.</p> <p>Description of datasets:</p> <ul> <li><strong>SSH_Publications_in_OC_Meta_and_Open_Access_status.csv:&nbsp;</strong>containing information about OpenCitations Meta coverage of ERIH PLUS Journals as well as their Open Access availability. In this dataset, every row holds data for a Journal of ERIH PLUS also covered by OpenCitations Meta database. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> <li><strong>SSH_Publications_by_Discipline.csv:</strong>&nbsp;containing information about number of publications per&nbsp;discipline&nbsp;(in addition, number of journals&nbsp;per discipline are also included). The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>SSH_Publications_and_Journals_by_Country:</strong>&nbsp;containing information about number of publications and journals per&nbsp;country.&nbsp;The dataset has three columns, the first, labeled <strong>&quot;Country&quot;,</strong>&nbsp;contains single countries of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>result_disciplines.json:</strong> the dictionary containing all disciplines as key and a list of&nbsp;related ERIH PLUS venue identifiers as value.</li> <li><strong>result_countries.json:</strong>&nbsp;the dictionary containing all countries as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>duplicate_omids.csv: </strong>a dataset containing the duplicated Journal entries in OpenCitations Meta, structured with two columns: &quot;<strong>OC_omid&quot;</strong>, the internal OC Meta identifier; &quot;<strong>issn&quot;,&nbsp;</strong>the issn values associated to that identifier</li> <li><strong>eu_data.csv: </strong>contains the data specific for&nbsp;European countries&#39; SSH Journals&nbsp;covered in OCMeta. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;Original_Title&quot;</strong>,<strong> &quot;Country_of_Publication&quot;</strong>,<strong>&quot;ERIH_PLUS_Disciplines&quot;</strong>, <strong>&quot;disc_count&quot;</strong>, the number of disciplines per Journal.</li> <li><strong>eu_disciplines_count.csv:&nbsp;</strong>containing information about number of publications per&nbsp;discipline and number of journals&nbsp;per discipline of european countries. The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_eu.csv:&nbsp;</strong>contains the data specific for&nbsp;European countries&#39; SSH Journals&nbsp;covered in OCMeta.&nbsp;It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> <li><strong>us_data.csv:&nbsp;</strong>contains the data specific for the&nbsp;United States&#39; SSH Journals&nbsp;covered in OCMeta. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;Original_Title&quot;</strong>,<strong> &quot;Country_of_Publication&quot;</strong>,<strong>&quot;ERIH_PLUS_Disciplines&quot;</strong>, <strong>&quot;disc_count&quot;</strong>, the number of disciplines per Journal.</li> <li><strong>us_disciplines_count.csv:&nbsp;</strong>containing information about number of publications per&nbsp;discipline and number of journals&nbsp;per discipline of the United States. The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_us.csv:&nbsp;</strong>contains the data specific for the United States&#39; SSH Journals&nbsp;covered in OCMeta.&nbsp;It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> </ul> <p>&nbsp;</p> <p><strong>Abstract of the research:&nbsp;</strong></p> <p><strong>Purpose:</strong>&nbsp;this study aims to investigate the representation and distribution of Social Science and Humanities (SSH) journals within the OpenCitations Meta database, with a particular emphasis on their Open Access (OA) status, as well as their spread across different disciplines and countries. The underlying premise is that open infrastructures play a pivotal role in promoting transparency, reproducibility, and trust in scientific research.<br> <strong>Study Design and Methodology:</strong>&nbsp;the study is grounded on the premise that open infrastructures are crucial for ensuring transparency, reproducibility, and fostering trust in scientific research. The research methodology involved the use of secondary data sources, namely the OpenCitations Meta database, the ERIH PLUS bibliographic index, and the DOAJ index. A custom research software was developed in Python to facilitate the processing and analysis of the data.<br> <strong>Findings:</strong>&nbsp;the results reveal that 78.1% of SSH journals listed in the European Reference Index for the Humanities (ERIH-PLUS) are included in the OpenCitations Meta database. The discipline of Psychology has the highest number of publications. The United States and the United Kingdom are the leading contributors in terms of the number of publications. However, the study also uncovers that only 38% of the SSH journals in the OpenCitations Meta database are OA.<br> <strong>Originality:</strong>&nbsp;this research adds to the existing body of knowledge by providing insights into the representation of SSH in open bibliographic databases and the role of open access in this domain. The study highlights the necessity for advocating OA practices within SSH and the significance of open data for bibliometric studies. It further encourages additional research into the impact of OA on various facets of citation patterns and the factors leading to disparity across disciplinary representation.</p> <p><strong>Related resources:</strong></p> <p>Ghasempouri S., Ghiotto M., &amp; Giacomini S. (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESEARCH ARTICLE.&nbsp;<a href="https://doi.org/10.5281/zenodo.8263908">https://doi.org/10.5281/zenodo.8263908</a></p> <p>Ghasempouri, S.,&nbsp;Ghiotto, M., Giacomini, S., (2023).&nbsp; Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - DATA MANAGEMENT PLAN (Version 4). Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.8174644">https://doi.org/10.5281/zenodo.8174644</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S. (2023e). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - PROTOCOL. V.5. (<a href="https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5">https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5</a>)</p>

opencc-byMay 2023View details →
zenodo44/100

Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESULTS DATASET (without Mega Journals)

<p>The dataset contains all the data produced running the research software for the study <em>Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta</em>, a research carried out in the contest of the Open Science course 22/23 at the University of Bologna.</p> <p>Mega Journals have been excluded form the datasets, since we found they were significantly skewing the results, the only datasets not interested by this exclusion are&nbsp;<strong>SSH_Publications_in_OC_Meta_and_Open_Access_status </strong>and<strong>&nbsp;duplicate_omids.</strong>&nbsp;The result datasets with Mega Journals included are published <a href="https://doi.org/10.5281/zenodo.8250858">here</a><br> The Journals excluded from the results are: PLOS ONE (issn:1932-6203), PNAS (issn:1091-6490), Science (issn:1095-9203), Nature(issn:0028-0836).</p> <p>Description of datasets:</p> <ul> <li><strong>SSH_Publications_in_OC_Meta_and_Open_Access_status.csv:&nbsp;</strong>containing information about OpenCitations Meta coverage of ERIH PLUS Journals as well as their Open Access availability. In this dataset, every row holds data for a Journal of ERIH PLUS also covered by OpenCitations Meta database. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> <li><strong>SSH_Publications_by_Discipline.csv:</strong>&nbsp;containing information about number of publications per&nbsp;discipline&nbsp;(in addition, number of journals&nbsp;per discipline are also included). The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>SSH_Publications_and_Journals_by_Country:</strong>&nbsp;containing information about number of publications and journals per&nbsp;country.&nbsp;The dataset has three columns, the first, labeled <strong>&quot;Country&quot;,</strong>&nbsp;contains single countries of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>result_disciplines.json:</strong> the dictionary containing all disciplines as key and a list of&nbsp;related ERIH PLUS venue identifiers as value.</li> <li><strong>result_countries.json:</strong>&nbsp;the dictionary containing all countries as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>duplicate_omids.csv: </strong>a dataset containing the duplicated Journal entries in OpenCitations Meta, structured with two columns: &quot;<strong>OC_omid&quot;</strong>, the internal OC Meta identifier; &quot;<strong>issn&quot;,&nbsp;</strong>the issn values associated to that identifier</li> <li><strong>eu_data.csv: </strong>contains the data specific for&nbsp;European countries&#39; SSH Journals&nbsp;covered in OCMeta. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;Original_Title&quot;</strong>,<strong> &quot;Country_of_Publication&quot;</strong>,<strong>&quot;ERIH_PLUS_Disciplines&quot;</strong>, <strong>&quot;disc_count&quot;</strong>, the number of disciplines per Journal.</li> <li><strong>eu_disciplines_count.csv:&nbsp;</strong>containing information about number of publications per&nbsp;discipline and number of journals&nbsp;per discipline of european countries. The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_eu.csv:&nbsp;</strong>contains the data specific for&nbsp;European countries&#39; SSH Journals&nbsp;covered in OCMeta.&nbsp;It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> <li><strong>us_data.csv:&nbsp;</strong>contains the data specific for the&nbsp;United States&#39; SSH Journals&nbsp;covered in OCMeta. It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;Original_Title&quot;</strong>,<strong> &quot;Country_of_Publication&quot;</strong>,<strong>&quot;ERIH_PLUS_Disciplines&quot;</strong>, <strong>&quot;disc_count&quot;</strong>, the number of disciplines per Journal.</li> <li><strong>us_disciplines_count.csv:&nbsp;</strong>containing information about number of publications per&nbsp;discipline and number of journals&nbsp;per discipline of the United States. The dataset has three columns, the first, labeled <strong>&quot;Discipline&quot;,</strong>&nbsp;contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>&quot;Journal_count&quot;&nbsp;</strong>and <strong>&quot;Publication_count&quot;,&nbsp;</strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_us.csv:&nbsp;</strong>contains the data specific for the United States&#39; SSH Journals&nbsp;covered in OCMeta.&nbsp;It is structured with the following columns:&nbsp; &quot;<strong>EP_id&quot;, </strong>the internal ERIH PLUS identifier; <strong>&quot;Publications_in_venue&quot;, </strong>the<strong>&nbsp;</strong>numbers of Publications counted in each venue; <strong>&quot;</strong><strong>OC_omid&quot;, </strong>the internal OpenCitations Meta identifier for the venue;<strong>&nbsp;&quot;issn&quot;,</strong> numbers of publications in each venue;<strong>&nbsp;&quot;Open Access&quot;,</strong> a value to represent if the journal is OA or not, either &quot;True&quot; or &quot;Unknown&quot;.</li> </ul> <p>&nbsp;</p> <p><strong>Abstract of the research:&nbsp;</strong></p> <p><strong>Purpose:</strong>&nbsp;this study aims to investigate the representation and distribution of Social Science and Humanities (SSH) journals within the OpenCitations Meta database, with a particular emphasis on their Open Access (OA) status, as well as their spread across different disciplines and countries. The underlying premise is that open infrastructures play a pivotal role in promoting transparency, reproducibility, and trust in scientific research.<br> <strong>Study Design and Methodology:</strong>&nbsp;the study is grounded on the premise that open infrastructures are crucial for ensuring transparency, reproducibility, and fostering trust in scientific research. The research methodology involved the use of secondary data sources, namely the OpenCitations Meta database, the ERIH PLUS bibliographic index, and the DOAJ index. A custom research software was developed in Python to facilitate the processing and analysis of the data.<br> <strong>Findings:</strong>&nbsp;the results reveal that 78.1% of SSH journals listed in the European Reference Index for the Humanities (ERIH-PLUS) are included in the OpenCitations Meta database. The discipline of Psychology has the highest number of publications. The United States and the United Kingdom are the leading contributors in terms of the number of publications. However, the study also uncovers that only 38% of the SSH journals in the OpenCitations Meta database are OA.<br> <strong>Originality:</strong>&nbsp;this research adds to the existing body of knowledge by providing insights into the representation of SSH in open bibliographic databases and the role of open access in this domain. The study highlights the necessity for advocating OA practices within SSH and the significance of open data for bibliometric studies. It further encourages additional research into the impact of OA on various facets of citation patterns and the factors leading to disparity across disciplinary representation.</p> <p><strong>Related resources:</strong></p> <p>Ghasempouri S., Ghiotto M., &amp; Giacomini S. (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESEARCH ARTICLE.&nbsp;<a href="https://doi.org/10.5281/zenodo.8263908">https://doi.org/10.5281/zenodo.8263908</a></p> <p>Ghasempouri, S.,&nbsp;Ghiotto, M., Giacomini, S., (2023).&nbsp; Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - DATA MANAGEMENT PLAN (Version 4). Zenodo.&nbsp;<a href="https://doi.org/10.5281/zenodo.8174644">https://doi.org/10.5281/zenodo.8174644</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S. (2023e). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - PROTOCOL. V.5. (<a href="https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5">https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5</a>)</p>

opencc-byMay 2023View details →
zenodo44/100

A Shortlist of Diamond Open Access Journals for the Faculty of Science at Utrecht University

<p><strong>Context</strong></p> <p>The following shortlist of diamond open-access journals was compiled to increase awareness of alternative scholarly publication models among the six departments of the <a href="https://www.uu.nl/en/organisation/faculty-of-science">Faculty of Science at Utrecht University</a>. The list is relevant to the six disciplines at the Faculty of Science: Biology, Chemistry, Mathematics, Information and Computing Sciences, Physics, and Pharmaceutical Sciences. For this purpose, a &quot;diamond journal&quot; is defined as a journal indexed in the <a href="https://www.doaj.org/">Directory of Open Access Journals (DOAJ)</a> that does not charge an article processing charge (APC).</p> <p>&nbsp;</p> <p><strong>Contents and Results</strong></p> <p>The Excel file titled &ldquo;Diamond_journals_faculty_of_science_UU&rdquo; contains the list of selected diamond journals based on the following criteria: they allow submissions in English, have a plagiarism screening policy, possess an electronic ISSN number, and accept submissions in Biology, Chemistry, Mathematics, Information and Computing Sciences, Physics, and Pharmaceutical Sciences. In this shortlist, 355 journals meet the criteria. Out of these 355 journals, only 29 have received a DOAJ seal, 150 journals are indexed in <a href="https://www.scopus.com/">Scopus</a>, and 94 journals are indexed in <a href="https://mjl.clarivate.com/home">Web of Science</a>.</p> <p>A detailed description of the methods employed to obtain this shortlist can be found in the Word file titled &quot;Methods_and_Results&quot;.</p> <p>The raw CSV data has been included under the name &quot;Raw_DOAJ_journal_metadata_2023_07_25&quot;.</p> <p>&nbsp;</p> <p><strong>Limitations</strong></p> <p>The compilers of this shortlist are aware that some current diamond journals could change their status to non-diamond by charging article processing fees at a later stage. Since the journal record is not always updated by the publishers, we strongly recommend the users double-check the latest open access status directly on the journal&#39;s homepage (journal URLs are provided in the Excel file). The same applies for Scopus and WOS indexations.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Supplementary file 1 from: Moliner Cachazo L, Makati K, Chadwick MA, Catford JA, Price BW, Mackay AW, Guiry MD, Murray-Hudson M, Murray-Hudson F (2023) A review of the freshwater diversity in the Okavango Delta and Lake Ngami (Botswana): taxonomic composition, ecology, comparison with similar systems and conservation status. Aquatic Sciences

<p>Dataset&nbsp;with 2,204&nbsp;freshwater species from the Okavango Delta and Lake Ngami (Botswana), with additional 355&nbsp;species found in other areas of Botswana that are likely to be present in the study region. The dataset&nbsp;covers the following groups: amphibians, birds, fishes, macroinvertebrates, macrophytes, mammals, reptiles, phytoplankton, and zooplankton. The following information is given for each species: status in the Okavango Delta and Lake Ngami (present/potentially present);&nbsp;conservation status globally,&nbsp;Phylum,&nbsp;Class,&nbsp;Order,&nbsp;Family, Genus, species name, cited synonyms, common name, habitat, presence in high water, presence in low water, ecology, distribution in continental Africa, confirmed locations in the Okavango Delta, site coordinates, references, notes.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Data and figures for "Atlas of Science Collaboration, 1971–2020"

<p><strong>Abstract</strong></p><p>The evolving landscape of interinstitutional collaborative research across 15 natural science disciplines is explored using the open data sourced from OpenAlex.&nbsp;This extensive exploration spans the years from 1971 to 2020, facilitating a thorough investigation of leading scientific output producers and their collaborative relationships based on coauthorships.&nbsp;The findings are visually presented on world maps and other diagrams, offering a clear and insightful portrayal of notable variations in both national and international collaboration patterns across various fields and time periods.&nbsp;These visual representations serve as valuable resources for science policymakers, diplomats and institutional researchers, providing them with a comprehensive overview of global collaboration and aiding their intuitive grasp of the evolving nature of these partnerships over time.</p><p>&nbsp;</p><p><strong>Intended Readership</strong></p><ul><li>The booklet, entitled<i>&nbsp;'</i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>'</i>, aims to offer a broad overview of international and interinstitutional research collaboration, shedding light on its present status and evolution on a global scale. While it might not delve into intricate scholarly or academic data analysis, it remains a valuable resource for those seeking a general understanding of the collaborative relationships that have been established between research institutions in the world of science.</li><li>The intended readership including science and technology (S&amp;T) policymakers and diplomats, government research and development (R&amp;D) agencies, international organisations, S&amp;T think tanks, as well as institutional research divisions of universities or R&amp;D institutions.</li></ul><p>&nbsp;</p><p><strong>Data Source</strong></p><ul><li>The<i> </i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>&nbsp;</i>is based on data retrieved from&nbsp;<a href="https://docs.openalex.org/">OpenAlex</a>, a free and open (the CC0 license) catalogue of the world's scholarly papers, researchers, journals and institutions. Launched in January 2022, OpenAlex replaced&nbsp;<a href="https://www.microsoft.com/en-us/research/project/microsoft-academic-graph/">Microsoft Academic Graph (MAG)</a>, which retired at the beginning of 2022.</li><li>OpenAlex collects information on scientific publications, including journal articles, non-journal articles, preprints, conference papers, books and datasets—hereafter collectively referred to as 'works'—from various platforms such as&nbsp;<a href="https://www.crossref.org/">Crossref</a>,&nbsp;<a href="https://orcid.org/">ORCID</a>,&nbsp;<a href="https://ror.org/">ROR</a>,&nbsp;<a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>, preprint servers like&nbsp;<a href="https://arxiv.org/">arXiv</a>, and institutional or disciplinary repositories like&nbsp;<a href="https://zenodo.org/">Zenodo</a>. For comparison with other scholarly data sources such as&nbsp;<a href="https://www.scopus.com/">Scopus</a>,&nbsp;<a href="https://clarivate.com/products/scientific-and-academic-research/research-discovery-and-workflow-solutions/webofscience-platform/">Web of Science</a>&nbsp;and&nbsp;<a href="https://www.dimensions.ai/">Dimensions</a>, please refer to&nbsp;<a href="https://openalex.org/about#comparison">OpenAlex's website</a>.</li><li>OpenAlex offers extensive coverage of meta-information across a diverse spectrum of works, encompassing not only journal publications but also non-journal works, non-English works and contributions from the Global South. This attribute proves beneficial by providing a more precise augmentation of the extent of R&amp;D activities, along with their associated scholarly outputs. This is especially crucial in fields where journals are not the predominant channel for disseminating research outcomes. Furthermore, OpenAlex effectively captures outputs in the preprint format, which might persist for varying durations, spanning from months to years or even indefinitely, without necessarily transitioning into journal publications.</li><li>The present edition (August 2023) of&nbsp;the<i> </i><a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a><i>&nbsp;</i>was compiled using data obtained via the&nbsp;<a href="https://docs.openalex.org/how-to-use-the-api/api-overview">OpenAlex API</a>&nbsp;during the period from the 12th to the 15th of August 2023. It is essential to note that OpenAlex is an ongoing project, continuously updating its data and improving its system. Consequently, the visualisations in this booklet may not provide the most comprehensive view or accurate data. Expect more accurate results when acquiring data in the future as OpenAlex undergoes further upgrades. Revised editions of&nbsp;the<i> Atlas of Science Collaboration&nbsp;</i>may be made available on&nbsp;<a href="https://zenodo.org/">Zenodo</a>&nbsp;or other open platforms beyond this release.</li></ul><p>&nbsp;</p><p><strong>R&amp;D Disciplines</strong></p><ul><li>In this current edition, the primary focus centres around the level-1 'concepts' listed in the following table&nbsp;sourced from the OpenAlex classification, as previously explored in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a>. Each level-1 concept is accompanied by 'related concepts', which can offer a finer or broader delineation compared to the level-1 concept. Using this characteristic, an enhanced notion of R&amp;D discipline is constructed by including all associated subconcepts of level 2 or higher for each of the 15 level-1 concepts. For instance, our defined discipline of 'Artificial Intelligence' includes OpenAlex's level-2 concepts of '<a href="https://explore.openalex.org/concepts/C50644808">Artificial Neural Network</a>' and '<a href="https://explore.openalex.org/concepts/C108583219">Deep Learning</a>', but not the level-0 concepts of '<a href="https://explore.openalex.org/concepts/C41008148">Computer Science</a>' or '<a href="https://explore.openalex.org/concepts/C33923547">Mathematics</a>'.</li></ul><p>&nbsp;</p><p>&nbsp; OpenAlex Concept / Identifier / Discipline Code&nbsp;</p><ol><li>Artificial intelligence&nbsp;/ <a href="https://explore.openalex.org/concepts/C154945302">C154945302</a> / "ai"</li><li>Quantum mechanics&nbsp;/ <a href="https://explore.openalex.org/concepts/C62520636">C62520636</a> / "quantum"</li><li>Biotechnology&nbsp;/ <a href="https://explore.openalex.org/concepts/C150903083">C150903083</a> / "bio"</li><li>Nanotechnology&nbsp;/ <a href="https://explore.openalex.org/concepts/C171250308">C171250308</a> / "nano"</li><li>Agricultural engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C88463610">C88463610</a> / "agri"</li><li>Particle physics&nbsp;/ <a href="https://explore.openalex.org/concepts/C109214941">C109214941</a> / "particle"</li><li>Aerospace engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C146978453">C146978453</a> / "aerospace"</li><li>Nuclear engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C116915560">C116915560</a> / "nuclear"</li><li>Marine engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/c199104240">C199104240</a> / "marine"</li><li>Neuroscience&nbsp;/ <a href="https://explore.openalex.org/concepts/c169760540">C169760540</a> / "neuro"</li><li>Condensed matter physics&nbsp;/ <a href="https://explore.openalex.org/concepts/C26873012">C26873012</a> / "condensed"</li><li>Environmental engineering&nbsp;/ <a href="https://explore.openalex.org/concepts/C87717796">C87717796</a> / "envi"</li><li>Earth science&nbsp;/ <a href="https://explore.openalex.org/concepts/c1965285">C1965285</a> / "earth"</li><li>Astronomy&nbsp;/ <a href="https://explore.openalex.org/concepts/c1276947">C1276947</a> / "astro"</li><li>Pure mathematics&nbsp;/ <a href="https://explore.openalex.org/concepts/C202444582">C202444582</a> / "math"</li></ol><p>&nbsp;</p><p><strong>Analysis and Visualisation</strong></p><ul><li>First,&nbsp;the<i> World Map of Science Collaboration</i> ('<strong>wmap_bilat</strong>' folder)&nbsp;divides the period from 1971 to 2020 into four intervals: 1971–1990, 1991–2000, 2001–2010 and 2011–2020. For each period and discipline, bubbles represent the top 199 research institutions in terms of work production. Additionally, for the top 50 research institutions, their locations are connected on the world map using great circle curves (the shortest route between them) to illustrate bilateral coauthorship relationships. Coauthorship relationships with fewer than five coauthored papers are not displayed. The background world map utilises the&nbsp;world&nbsp;data from the&nbsp;<a href="https://cran.r-project.org/package=maps">maps</a>&nbsp;package&nbsp;in R. The connection visualisation between two research institutions leverages the&nbsp;gcIntermediate()&nbsp;function from the&nbsp;<a href="https://cran.r-project.org/package=geosphere">geosphere</a>&nbsp;package&nbsp;in R. The sizes of the bubbles are proportional to the volume of work and can be compared across the different period panels.</li><li>Second,&nbsp;the<i> Top 30 Productive Institutions on the World Map&nbsp;</i>('<strong>wmap_topinst</strong>' folder)&nbsp;displays the leading 30 institutions in terms of work production on the World Map for each discipline and the three respective periods: 1991–2000, 2001–2010 and 2011–2020. The background world map employs the&nbsp;world&nbsp;data from the&nbsp;<a href="https://cran.r-project.org/package=maps">maps</a>&nbsp;package in R along with the&nbsp;<a href="https://cran.r-project.org/package=ggplot2">ggplot2</a>&nbsp;package16&nbsp;in R. The sizes of the bubbles are proportional to the volume of work, standardised within each period panel, and cannot be compared across panels.</li><li>Third,&nbsp;the<i> Interregional Collaboration Matrix Diagram&nbsp;</i>('<strong>halfmat</strong>' folder)&nbsp;exhibits a half-matrix diagram at the country level for each discipline and the three respective periods: 1991–2000, 2001–2010 and 2011–2020. It counts the number of bilateral coauthorship relationships represented on the World Map. Each bubble's size (area) displayed in the matrix cell is proportional to the number of bilateral coauthorship relationships.&nbsp;This edition particularly focuses on five pivotal parties: the US, China, EU27, the UK and Japan.&nbsp;These parties were specifically selected due to their substantial contributions to work production across all scientific fields from 1971 to 2020.&nbsp;These choices also align with the nations acclaimed as the 'Big 5' science nations&nbsp;(the US, China, Germany, the UK and Japan) in the <a href="https://www.nature.com/articles/d41586-022-00569-7"><i>Nature Index</i></a>.&nbsp;Please note that the Matrix Diagram&nbsp;only takes into account the top 50 institutions in terms of work production for each period and discipline.&nbsp;Therefore, if a cell shows zero (as small dots), it does not necessarily imply the absence of coauthorship relationships for the corresponding bilateral pair.</li><li>Forth,&nbsp;the<i> Interinstitutional Collaboration Dendrogram&nbsp;</i>('<strong>cdend</strong>' folder)&nbsp;elucidates the development and evolution of interinstitutional research collaboration clusters spanning the last five decades. This is accomplished through hierarchical clustering analysis of institutions, considering the top 50 institutions in terms of work production across the four periods: 1971–1990, 1991–2000, 2001–2010 and 2011–2020.<ul><li>The method used for hierarchical clustering analysis is the same as developed in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a>. The distance between institutions X and Y is defined as the number of works with nationalities from both X and Y divided by the total number of works with nationalities from at least one of X and Y, subtracted from 1. Hierarchical clustering analysis was performed on the distance matrix using the&nbsp;hclust&nbsp;function implemented in R with the&nbsp;ward.D2&nbsp;option (i.e. the original Ward's method) specified.</li><li>The method of dendrogram visualisation is primarily derived from an example detailed on the&nbsp;<a href="https://cran.r-project.org/web/packages/dendextend/vignettes/dendextend.html">dendextend&nbsp;website</a>. Circular dendrograms were created using the&nbsp;<a href="https://cran.r-project.org/package=dendextend">dendextend</a>&nbsp;and&nbsp;<a href="https://cran.r-project.org/package=circlize">circlize</a>&nbsp;packages in R. As one moves inward from the outer edge of the circle towards its centre, institutions or clusters of institutions that are in closer proximity to each other merge earlier.</li><li>To indicate the country where the institutions are located, the country names are included at the beginning of the terms of research institutions, using the two-letter&nbsp;<a href="https://en.wikipedia.org/wiki/ISO_3166-1_alpha-2">ISO3166-1alpha-2</a>&nbsp;code.&nbsp;The accompanied circularised bar graphs represent the number of works for the institutions&nbsp;involved.&nbsp;<a href="https://ror.org/">ROR</a>s are used as the canonical identifiers of the research institutions. Readers of this booklet in PDF format can click on the ROR-based URL ('https://ror.org/...') in the diagrams to view the corresponding ROR webpage from their browser.</li></ul></li><li>Additionally, for each discipline and the respective periods of 1971–1990, 1991–2000, 2001–2010 and 2011–2020, the top 100 institutions in terms of work production are displayed in tabular format ('<strong>table</strong>' folder), showing their respective country codes and production volumes. If multiple research institutions have equal production volumes during each period, they are organised alphabetically by country codes and then by organisation names. Even if distinct rankings are shown, they lack significance and are treated as ties.</li></ul><p>&nbsp;</p><p><strong>Important Notes</strong></p><ul><li>It is worth reiterating that the data from OpenAlex used to compile&nbsp;the <a href="https://arxiv.org/abs/2308.16810"><i>Atlas of Science Collaboration</i></a>, even when incorporating bibliometric data related to past works, lacks consistent finality. As of the data acquisition for this version (August 2023), OpenAlex encompassed information regarding approximately 240 million works, with an additional influx of about 50,000 new data entries related to works being added daily.&nbsp;Furthermore, for a substantial portion of these works, information regarding the corresponding institution to which the authors belong remains unknown. As a result, should the same analyses as those embedded within this booklet be replicated in the future, although the qualitative extent of change remains uncertain, it is undeniable that quantitatively distinct data will be acquired. Nonetheless, for individuals seeking an understanding of the global scope and evolution of international and interinstitutional collaborative research, the potential availability of this booklet or an enhanced, continuously updated evidence base holds inherent value.</li><li>Further, it is worth reiterating that the term 'works' encompasses a wide variety of scholarly publications. The analyses conducted in the compilation of this booklet do not take into consideration whether these works are peer-reviewed articles or not, nor do they encompass considerations of their prominence, impact or quality. It is emphasised that the primary intent behind the visualisations in this booklet is to quantitatively capture the momentum of scholarly knowledge production outputs from diverse research institutions, and to identify how productive institutions collaborate internationally and interinstitutionally. Caution must be exercised, with acknowledgment that relying solely on the quantity of scholarly output produced by institutions falls short in encompassing discussions about their research potential, contributions to academia, or their relative superiority or inferiority. Further, it is recommended to consider the limitations discussed in <a href="https://doi.org/10.48550/arXiv.2211.04429">Okamura (2023)</a> when using this booklet.</li></ul><p>&nbsp;</p><p><strong>Miscellaneous</strong></p><ul><li>It is important to note that some research institutions may encounter difficulties in accurately assessing the actual production volume at the institutional level within each analysis period due to challenges related to name disambiguation and the influence of historical organisational changes in bibliometric databases.</li><li>For the Interinstitutional Collaboration Dendrograms and the rankings of the top 100 productive institutions, entities like universities and R&amp;D institutions are primarily identified using the nomenclature employed in OpenAlex. However, certain portions have been presented through abbreviations or acronyms, both for illustrative purposes and to effectively accommodate limited space. For instance, 'University of' is abbreviated as 'U.', 'Institution' and 'Institute' as 'Inst', 'National Laboratory' as 'NL', and 'Science' and 'Technology' as 'Sci' and 'Tech', correspondingly, among others. Should readers possess more fitting suggestions for abbreviations specific to particular organisations, or any other ideas aimed at enhancing the content of this booklet, we would greatly appreciate their input.</li></ul><p>&nbsp;</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Science, Technology & Society Eurobarometers 1993 - 2021: Trend Data Collection

<p>These files contain structured collections of Eurobarometer (EB) survey data from 1993 to 2021, focusing on European citizen&rsquo;s views on science and technology (S&amp;T). The primary aim of these data collections is to facilitate research on trends over time regarding people&rsquo;s knowledge, perception and attitudes towards S&amp;T. The European Union has collected extensive survey data from the general public over the past 50 years through its official polling instrument, the Eurobarometer. Its general goal is monitoring the state of public opinion on diverse subjects and issues throughout Europe, one of which is S&amp;T. The data collection files, provided in the folder &lsquo;EB_data_csv&rsquo;, include data from seven different Eurobarometer surveys. These surveys were selected because their raw datasets were openly accessible through Open EU Datasets. To ensure sufficient data points for plotting specific trends over time, we included only survey questions that appeared in at least three different EB surveys.&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Materiali del sito Open Science - Università degli Studi di Palermo Sistema Bibliotecario e Archivio storico di Ateneo (SBA) - Servizi di supporto alla ricerca

<p>Contenuti, immagini e poster del sito Open Science, sottosezione del portale delle biblioteche del portale di Ateneo, Universit&agrave; degli Studi di Palermo</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Review of documents about Open Science

<p>In a review exercise of a sample of 46 documents about Open Science governance, from institutions and regulatory actors, we identified opportunities to strengthen Open Science practices related to integrity, traceability, and preservation.&nbsp;</p><p>Variables to be considered: Document name, file type, institution – country, PDF/A, have DOI, have other PID, recognized by Zotero, recognized by Mendeley, metadata in PDF properties.</p><p>&nbsp;</p>

opencc-by-4.0Oct 2023View details →
edi44/100

Data from Citizen science data reveal regional heterogeneity in phenological response to climate in the large milkweed bug, Oncopeltus fasciatus

These data include annotations for life stage, mating behavior, and plant part occupancy of large milkweed bug observations in North America as well as information about climate and environment.

openCC0Feb 2023View details →
edi44/100

Lake ice surveys, 1874-2022, Adirondack Long-Term Ecological Monitoring Program Project No. 8 by Adirondack Ecological Center of the State University of New York College of Environmental Science and Forestry, Newcomb, New York. Environmental Data Initiative.

The objective of this dataset is to document ice-in and ice-out dates on several lakes on the State University of New York College of Environmental Science and Forestry's Huntington Wildlife Forest (HWF). Lakes include: Arbutus, Catlin, Deer, Military, Rich, Wolf and Lodo Pond; some records exist for Long Pond and other water bodies but they are not included here except in some comment fields.

openCC (other)Dec 2022View details →
edi44/100

Return on Investment Metrics for Data Repositories in Earth and Environmental Sciences

Despite a growing recognition of the importance of data to the economy and to science, investment in repositories to manage and disseminate that data in easily accessible and understandable ways is scarce. Keeping repository services active and up-to-date for a long time period is difficult due to this funding situation. As a result, repositories must continually provide proof of their value, their Return on Investment (ROI) to their sponsors; yet doing so has always been difficult, problematic and not always successful. In this work, an analysis of approaches for assessing the ROI of several scientific data repositories has identified various techniques that repositories use to report on the impact and value of their data products and services. A survey of selected repositories rated the set of metrics identified and rated each by its importance as well as the ease with which the metric could be measured. The discussion is broken down into considerations for calculating costs, perceived value of repositories and suggested metrics that would allow a repository to calculate an ROI. The authors, representatives of environmental data repositories, concluded that easily obtainable data use metrics, such as data downloads, etc., have limited value while more informative analyses would require additional resources.

openCC (other)Feb 2019View details →
edi44/100

Alaska 2004 Burns: Densities of tree seedlings after fire; measured in 2006, 2008, 2011, and 2017 by Joint Fire Science Program

This dataset contains counts and densities of tree seedlings that established naturally in the years after the 2004 burns in interior Alaska. Records are from a network of 90 sites established in 2005 along the Steese, Taylor, and Dalton highways as part of a Joint Fire Science Project. Counts are reported here for sample years in 2006, 2008, 2011, and 2017.

openOpenMay 2018View details →
OpenNeuro40/100

FSL open science dev dataset

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record