Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
16
datasets available to search
ShareScore release 0.7.1
Dataset results
16 results for “OpenCitations”
OpenCitations Meta RDF dataset of agent roles metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>agent roles</strong> of bibliographic resources<strong> </strong>(<a href="http://purl.org/spar/pro/RoleInTime" target="_blank" rel="noopener">http://purl.org/spar/pro/RoleInTime</a>). These agents can be authors, editors, or publishers. It contains all the metadata and its provenance information, structured specifically around agent roles, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /ar/06250/10000/1000/1000.zip, while information about provenance in /ar/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of page numbers metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>page numbers</strong> of bibliographic resources, known as <strong>manifestations </strong>(<a href="http://purl.org/spar/fabio/Manifestation" target="_new">http://purl.org/spar/fabio/Manifestation</a>). It contains all the bibliographic metadata and its provenance information, structured specifically around manifestations (page numbers), in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of identifiers metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>identifiers </strong>(<a href="http://purl.org/spar/datacite/Identifier" target="_blank" rel="noopener">http://purl.org/spar/datacite/Identifier</a>) of bibliographic resources. It contains all the metadata and its provenance information, structured specifically around identifiers, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /id/06250/10000/1000/1000.zip, while information about provenance in /id/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of bibliographic resources metadata and its provenance information
<div> <p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>bibliographic resources </strong>(<a href="http://purl.org/spar/fabio/Expression" target="_blank" rel="noopener">http:///purl.org/spar/fabio/Expression</a>). It contains all the metadata and its provenance information, structured specifically around bibliographic resources, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p> <p> </p> </div>
OpenCitations Meta CSV dataset of all bibliographic metadata
<p>This dataset contains all the bibliographic metadata (in CSV format) included in OpenCitations Meta. In particular, each line of the CSV file defines a bibliographic resource, and includes the following information:</p> <ul> <li><strong>[field "id"]</strong> the IDs for the document described within the line;</li> <li><strong>[field "title"]</strong> the document's title;</li> <li><strong>[field "author"]</strong> the authors of the document;</li> <li><strong>[field "pub_date"]</strong> the date of publication;</li> <li><strong>[field "venue"]</strong> information about the venue, i.e. the bibliographical resource to which the document belongs;</li> <li><strong>[field "volume"]</strong> the volume sequence identifier (e.g. a number) to which the entity belongs;</li> <li><strong>[field "issue"]</strong> the issuesequence identifier (e.g. a number) to which the entity belongs;</li> <li><strong>[field "page"]</strong> the page range of the resource described in the row;</li> <li><strong>[field "type"]</strong> the type of resource described in the row;</li> <li><strong>[field "publisher"]</strong> the entity responsible for making the resource available;</li> <li><strong>[field "editor"] </strong>the editors of the document.</li> </ul> <p>This version of the dataset contains:</p> <ul> <li>114,621,237 bibliographic entities</li> <li>298,847,794 authors and 2,465,711 editors (counted by their roles, without disambiguating individual</li> <li>711,711 publication venues</li> <li>241,783 publishers</li> </ul> <p>The zipped dataset weighs 11 GB, while, when extracted, it weighs 46 GB on an ext4 filesystem.</p> <p>Additional information about OpenCitations Meta at <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
Bibliographic dataset based on Scientometrics, containing provenance information compliant with the OpenCitations Data Model and non disambigued authors
<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known. The data was extracted via Crossref. It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works. Journals and bibliographic resources always appear unambiguously, without duplicates. On the contrary, the authors have not been disambigued. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p>
Bibliographic dataset based on Scientometrics, including provenance information compliant with the OpenCitations Data Model
<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known. The data was extracted via Crossref. It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works. Journals, bibliographic resources, and authors always appear unambiguously, without duplicates. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p> <p>The dataset is distributed as two journal files, one for the data and one for the provenance, readable via the triplestore Blazegraph. There are 4,960,087 data triples and 19,348,027 provenance triples, which corresponds to 1,134,545 entities and 2,696,689 snapshots. Therefore, on average, each entity has two snapshots. Among the data, there are 231,217 agent roles, 221,602 responsible agents, 206,003 bibliographic resources, 142,472 citations, 141,555 bibliographical references, 108,112 identifiers, and 83,584 resource embodiments.</p> <p>The code to generate and modify such collections is available at <a href="https://doi.org/10.5281/zenodo.5579754">https://doi.org/10.5281/zenodo.5579754</a>. </p>
Coverage of DOAJ journals' citations through OpenCitations - Result DataSet
<p>The dataset contains: </p> <ul> <li><strong>by_journal.json</strong>: a file containing all information extracted by Open Citations about DOAJ journals divide by year and journal name. Inside the file, the metadata about the journal are: <ul> <li>ISSN</li> <li>EISSN</li> <li>number of articles overall in the journal</li> <li>subject(s) </li> <li>number of citations received</li> <li>number of citations done</li> <li>ratio between citations done and received</li> <li>number of citations received from DOAJ journals</li> <li>number of citations done to DOAJ journals</li> <li>ratio between citations done to and received from DOAJ journals.</li> </ul> </li> </ul> <ul> <li><strong>normal.json</strong>: a file containing all information extracted from Open Citations about DOAJ journals divided only by year. Inside the file, the data by year are: <ul> <li>number of citations received.</li> <li>number of citations done.</li> <li>ratio between citations done and received.</li> <li>number of self-citations made by DOAJ inside Open Citations.</li> <li>ratio between the self-citation and the total citations received and done by DOAJ.</li> </ul> </li> </ul> <ul> <li><strong>errors.json</strong>: a file containing the count of all errors obtained from computations. Inside the file: <ul> <li>errors about records that don't have any specified date (null dates).</li> <li>errors about records that have impossible dates (wrong dates).</li> <li>errors about articles that don't have any specified Dois.</li> <li>errors about Open Citations records that don't have any Dois in the citing or cited fields.</li> </ul> </li> <li><strong>DOAJ_metrics.json</strong>: a file containing metrics about DOAJ and Open Citations, obtained by computations. Inside the file are these fields: <ul> <li>number of journals with dois.</li> <li>number of articles which have been processed during computations.</li> <li>number of used Dois. All dois (with no repetition) which are used for the adding journal operation.</li> <li>number of repeated Dois. All dois which are repeated inside the same or in another journal.</li> <li>number of accepted Dois. All articles (with repetition) which have both a defined journal and a defined doi.</li> </ul> </li> </ul> <p> </p>
Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESULTS DATASET (with Mega Journals)
<p>The dataset contains all the data produced running the research software for the study:"Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta".</p> <p>Disclaimer: these results are not considered to be representative, because we have fount that Mega Journals skewed significantly some of the data. The result datasets without Mega Journals are published <a href="https://zenodo.org/record/8249907">here</a>.</p> <p>Description of datasets:</p> <ul> <li><strong>SSH_Publications_in_OC_Meta_and_Open_Access_status.csv: </strong>containing information about OpenCitations Meta coverage of ERIH PLUS Journals as well as their Open Access availability. In this dataset, every row holds data for a Journal of ERIH PLUS also covered by OpenCitations Meta database. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>SSH_Publications_by_Discipline.csv:</strong> containing information about number of publications per discipline (in addition, number of journals per discipline are also included). The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>SSH_Publications_and_Journals_by_Country:</strong> containing information about number of publications and journals per country. The dataset has three columns, the first, labeled <strong>"Country",</strong> contains single countries of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>result_disciplines.json:</strong> the dictionary containing all disciplines as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>result_countries.json:</strong> the dictionary containing all countries as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>duplicate_omids.csv: </strong>a dataset containing the duplicated Journal entries in OpenCitations Meta, structured with two columns: "<strong>OC_omid"</strong>, the internal OC Meta identifier; "<strong>issn", </strong>the issn values associated to that identifier</li> <li><strong>eu_data.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>eu_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of european countries. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_eu.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>us_data.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>us_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of the United States. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_us.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> </ul> <p> </p> <p><strong>Abstract of the research: </strong></p> <p><strong>Purpose:</strong> this study aims to investigate the representation and distribution of Social Science and Humanities (SSH) journals within the OpenCitations Meta database, with a particular emphasis on their Open Access (OA) status, as well as their spread across different disciplines and countries. The underlying premise is that open infrastructures play a pivotal role in promoting transparency, reproducibility, and trust in scientific research.<br> <strong>Study Design and Methodology:</strong> the study is grounded on the premise that open infrastructures are crucial for ensuring transparency, reproducibility, and fostering trust in scientific research. The research methodology involved the use of secondary data sources, namely the OpenCitations Meta database, the ERIH PLUS bibliographic index, and the DOAJ index. A custom research software was developed in Python to facilitate the processing and analysis of the data.<br> <strong>Findings:</strong> the results reveal that 78.1% of SSH journals listed in the European Reference Index for the Humanities (ERIH-PLUS) are included in the OpenCitations Meta database. The discipline of Psychology has the highest number of publications. The United States and the United Kingdom are the leading contributors in terms of the number of publications. However, the study also uncovers that only 38% of the SSH journals in the OpenCitations Meta database are OA.<br> <strong>Originality:</strong> this research adds to the existing body of knowledge by providing insights into the representation of SSH in open bibliographic databases and the role of open access in this domain. The study highlights the necessity for advocating OA practices within SSH and the significance of open data for bibliometric studies. It further encourages additional research into the impact of OA on various facets of citation patterns and the factors leading to disparity across disciplinary representation.</p> <p><strong>Related resources:</strong></p> <p>Ghasempouri S., Ghiotto M., & Giacomini S. (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESEARCH ARTICLE. <a href="https://doi.org/10.5281/zenodo.8263908">https://doi.org/10.5281/zenodo.8263908</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S., (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - DATA MANAGEMENT PLAN (Version 4). Zenodo. <a href="https://doi.org/10.5281/zenodo.8174644">https://doi.org/10.5281/zenodo.8174644</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S. (2023e). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - PROTOCOL. V.5. (<a href="https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5">https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5</a>)</p>
Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESULTS DATASET (without Mega Journals)
<p>The dataset contains all the data produced running the research software for the study <em>Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta</em>, a research carried out in the contest of the Open Science course 22/23 at the University of Bologna.</p> <p>Mega Journals have been excluded form the datasets, since we found they were significantly skewing the results, the only datasets not interested by this exclusion are <strong>SSH_Publications_in_OC_Meta_and_Open_Access_status </strong>and<strong> duplicate_omids.</strong> The result datasets with Mega Journals included are published <a href="https://doi.org/10.5281/zenodo.8250858">here</a><br> The Journals excluded from the results are: PLOS ONE (issn:1932-6203), PNAS (issn:1091-6490), Science (issn:1095-9203), Nature(issn:0028-0836).</p> <p>Description of datasets:</p> <ul> <li><strong>SSH_Publications_in_OC_Meta_and_Open_Access_status.csv: </strong>containing information about OpenCitations Meta coverage of ERIH PLUS Journals as well as their Open Access availability. In this dataset, every row holds data for a Journal of ERIH PLUS also covered by OpenCitations Meta database. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>SSH_Publications_by_Discipline.csv:</strong> containing information about number of publications per discipline (in addition, number of journals per discipline are also included). The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>SSH_Publications_and_Journals_by_Country:</strong> containing information about number of publications and journals per country. The dataset has three columns, the first, labeled <strong>"Country",</strong> contains single countries of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>result_disciplines.json:</strong> the dictionary containing all disciplines as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>result_countries.json:</strong> the dictionary containing all countries as key and a list of related ERIH PLUS venue identifiers as value.</li> <li><strong>duplicate_omids.csv: </strong>a dataset containing the duplicated Journal entries in OpenCitations Meta, structured with two columns: "<strong>OC_omid"</strong>, the internal OC Meta identifier; "<strong>issn", </strong>the issn values associated to that identifier</li> <li><strong>eu_data.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>eu_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of european countries. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_eu.csv: </strong>contains the data specific for European countries' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> <li><strong>us_data.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"Original_Title"</strong>,<strong> "Country_of_Publication"</strong>,<strong>"ERIH_PLUS_Disciplines"</strong>, <strong>"disc_count"</strong>, the number of disciplines per Journal.</li> <li><strong>us_disciplines_count.csv: </strong>containing information about number of publications per discipline and number of journals per discipline of the United States. The dataset has three columns, the first, labeled <strong>"Discipline",</strong> contains single disciplines of the ERIH classificaton, the second and the third, labeled <strong>"Journal_count" </strong>and <strong>"Publication_count", </strong>respectively, the number of Journals and the number of Publications counted for each discipline.</li> <li><strong>meta_coverage_us.csv: </strong>contains the data specific for the United States' SSH Journals covered in OCMeta. It is structured with the following columns: "<strong>EP_id", </strong>the internal ERIH PLUS identifier; <strong>"Publications_in_venue", </strong>the<strong> </strong>numbers of Publications counted in each venue; <strong>"</strong><strong>OC_omid", </strong>the internal OpenCitations Meta identifier for the venue;<strong> "issn",</strong> numbers of publications in each venue;<strong> "Open Access",</strong> a value to represent if the journal is OA or not, either "True" or "Unknown".</li> </ul> <p> </p> <p><strong>Abstract of the research: </strong></p> <p><strong>Purpose:</strong> this study aims to investigate the representation and distribution of Social Science and Humanities (SSH) journals within the OpenCitations Meta database, with a particular emphasis on their Open Access (OA) status, as well as their spread across different disciplines and countries. The underlying premise is that open infrastructures play a pivotal role in promoting transparency, reproducibility, and trust in scientific research.<br> <strong>Study Design and Methodology:</strong> the study is grounded on the premise that open infrastructures are crucial for ensuring transparency, reproducibility, and fostering trust in scientific research. The research methodology involved the use of secondary data sources, namely the OpenCitations Meta database, the ERIH PLUS bibliographic index, and the DOAJ index. A custom research software was developed in Python to facilitate the processing and analysis of the data.<br> <strong>Findings:</strong> the results reveal that 78.1% of SSH journals listed in the European Reference Index for the Humanities (ERIH-PLUS) are included in the OpenCitations Meta database. The discipline of Psychology has the highest number of publications. The United States and the United Kingdom are the leading contributors in terms of the number of publications. However, the study also uncovers that only 38% of the SSH journals in the OpenCitations Meta database are OA.<br> <strong>Originality:</strong> this research adds to the existing body of knowledge by providing insights into the representation of SSH in open bibliographic databases and the role of open access in this domain. The study highlights the necessity for advocating OA practices within SSH and the significance of open data for bibliometric studies. It further encourages additional research into the impact of OA on various facets of citation patterns and the factors leading to disparity across disciplinary representation.</p> <p><strong>Related resources:</strong></p> <p>Ghasempouri S., Ghiotto M., & Giacomini S. (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - RESEARCH ARTICLE. <a href="https://doi.org/10.5281/zenodo.8263908">https://doi.org/10.5281/zenodo.8263908</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S., (2023). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - DATA MANAGEMENT PLAN (Version 4). Zenodo. <a href="https://doi.org/10.5281/zenodo.8174644">https://doi.org/10.5281/zenodo.8174644</a></p> <p>Ghasempouri, S., Ghiotto, M., Giacomini, S. (2023e). Open Science for Social Sciences and Humanities: Open Access availability and distribution across disciplines and Countries in OpenCitations Meta - PROTOCOL. V.5. (<a href="https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5">https://dx.doi.org/10.17504/protocols.io.5jyl8jo1rg2w/v5</a>)</p>
OC-782K: Knowledge Graph of "Scientometrics" modelled according to the OpenCitations Data Model
<p>This dataset is a knowledge graph extracted from a <a href="https://static.aminer.cn/misc/na-data-kdd18.zip">t</a>riplestore covering information about the journal <em>Scientometrics</em> and modelled according to the OpenCitations Data Model. The original triplestore is available <a href="https://doi.org/10.5281/zenodo.5151264">here</a>. This KG was extracted for a research project on knowledge graph embeddings (KGEs) for author disambiguation. Structural triples of the knowledge graph are split into training, testing and validation for applying representation learning methods. Textual literals and numeric literals were stored separately in order to implement multimodal approaches for KGEs (see <a href="https://arxiv.org/abs/1802.00934">arXiv:1802.00934</a>). For the same reason, textual literals and numeric literals are already stored into sentence embeddings and a numeric matrix respectively in the files <em>textual_literals.npy </em>and <em>numeric_literals.npy</em>. The file <em>and_eval</em><em>.json </em>contains the evaluation dataset used for evaluating our AND architecture. For the script used to gather this dataset see the GitHub repository: <a href="https://github.com/sntcristian/and-kge/tree/main/aminer">https://github.com/sntcristian/and-kge/tree/main/open-citations</a>.</p>
PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)
<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>
Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals - DATA PRODUCED
<p>This zipped folders contain all the data produced for the research "Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals": the results datasets (dataset_map_disciplines, dataset_no_SSH, dataset_SSH, erih_meta_with_disciplines and erih_meta_without_disciplines).</p> <ul> <li> <p><strong>dataset_map_disciplines.zip </strong>contains CSV files with four columns ("id", "citing", "cited", "disciplines") giving information about publications stored in OpenCitations META (version 3 released on February 2023) and part of SSH journals, according to ERIH PLUS (version downloaded on 2023-04-27), specifying the disciplines associated to them and a boolean value stating if they cite or are cited, according to the OpenCitations COCI dataset (version 19 released on January 2023).</p> </li> <li> <p><strong>dataset_no_SSH.zip </strong>and <strong>dataset_SSH.zip</strong> contain CSV files with the same structure. Each dataset has four columns: "citing", "is_citing_SSH", "cited", and "is_cited_SSH". ”Citing” and “cited” columns are filled with DOIs of publications stored in OpenCitations META that according to OpenCitations COCI are involved in a citation. The "is_citing_SSH" and "is_cited_SSH" columns contain boolean values: "True" if the corresponding publication is associated with a SSH (Social Sciences and Humanities) discipline, according to ERIH PLUS, and "False" otherwise. The two datasets are built starting from the two different subsets obtained as a result of the union between OpenCitations META and ERIH PLUS: dataset_SSH comes from erih_meta_with_disciplines and dataset_no_SSH from <strong>erih_meta_without_disciplines. </strong>dataset_no_SSH comes from <strong>erih_meta_with_disciplines.zip</strong> and erih_meta_without_disciplines.zip, as explained before, contain CSV files originating from ERIH PLUS and META. erih_meta_without_disciplines has just one column “id” and contains the DOIs of all the publications in META that do not have any discipline associated, that is, have not been published on a SSH journal, while erih_meta_with_disciplines derives from all the publications in META that have at least one linked discipline and has two columns: “id” and “erih_disciplines”, containing a string with all the disciplines linked to that publication like "History, Interdisciplinary research in the Humanities, Interdisciplinary research in the Social Sciences, Sociology".</p> </li> </ul> <p>Software: https://doi.org/10.5281/zenodo.8326023</p> <p>Data preprocessed: https://doi.org/10.5281/zenodo.7973159</p> <p>Article: https://zenodo.org/record/8326044</p> <p>DMP: https://zenodo.org/record/8324973</p> <p>Protocol: https://doi.org/10.17504/protocols.io.n92ldpeenl5b/v5</p>
Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals - DATA PREPROCESSED
<p>This zipped folders contain all the data preprocessed for the research "Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals": the cleaned datasets (coci_preprocessed, meta_preprocessed, erih_preprocessed and erih_meta).</p> <ul> <li> <p><strong>coci_preprocessed.zip</strong>: this archive contains CSVs with two columns “citing” and “cited”, giving information about publications involved in citations according to the OpenCitations COCI dataset (version 19 released on January 2023), and that are entirely contained in OpenCitations META (version 3 released on February 2023). This means that the citations which have either the citing or the cited entity (or both) not contained in META are excluded from coci_preprocessed dataset.</p> </li> <li> <p><strong>meta_preprocessed.zip</strong>: all the original columns of OpenCitations META are maintained in this dataset, so the CSVs have the columns: “id”, “title”, “author”, “issue”, “volume”, “venue”, “page”, “pub_date”, “type”, “publisher” and “editor”. The only difference with the original dataset is that meta_preprocessed in the columns “id” and “venue” has respectively just the DOIs and the ISSNs, without all the other identifiers specified for each entity in META.</p> </li> <li> <p><strong>erih_preprocessed.zip</strong>: it contains a CSV file with two columns "venue_id" and "ERIH_disciplines". "venue_id" is the union of the original columns "Online ISSN" and "Print ISSN" of ERIH_PLUS (version downloaded on 2023-04-27).</p> </li> <li> <p><strong>erih_meta.zip</strong>: it contains CSV files obtained from the union of meta_preprocessed and erih_preprocessed, they have all the columns of meta_preprocessed plus a new column “erih_disciplines” containing all the disciplines linked to a venue (identified by an ISSN).</p> </li> </ul> <p> </p> <p>Software: https://doi.org/10.5281/zenodo.8326023</p> <p>Data produced: https://doi.org/10.5281/zenodo.7974816</p> <p>Article: https://zenodo.org/record/8326044</p> <p>DMP: https://zenodo.org/record/8324973</p> <p>Protocol: https://doi.org/10.17504/protocols.io.n92ldpeenl5b/v5</p>
OpenCitations 2018-2020 requests: SPARQL endpoints vs REST APIs
<p><strong>The number of requests received by the OpenCitations (</strong><a href="http://opencitations.net/">http://opencitations.net/</a><strong>) SPARQL endpoints vs. the calls to the OpenCitations REST APIs between January 2018 and March 2020.</strong></p> <p> </p>
TEST [OpenCitations crowdsourcing: deposits of the week before 2022-09-17]
<p>OpenCitations collects citation data and related metadata from the community through issues on the GitHub repository <a href="https://github.com/opencitations/crowdsourcing">https://github.com/opencitations/crowdsourcing</a>. In order to preserve long-term provenance information, such data is uploaded to Zenodo every week. This is is a test upload containing the data of deposit issues published in the week before 2022-09-17.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.