Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
15
datasets available to search
ShareScore release 0.9.0
Dataset results
15 results for “Crossref”
Types, open citations, closed citations, publishers, and participation reports of Crossref entities
<p>This publication contains several datasets that have been used in the paper "Crowdsourcing open citations with CROCI – An analysis of the current status of open citations, and a proposal" submitted to the <a href="https://www.issi2019.org/">17th International Conference on Scientometrics and Bibliometrics (ISSI 2019)</a>, available at <a href="https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/">https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/</a>.</p> <p>Additional information about the analyses described in the paper, including the code and the data we have used to compute all the figures, is available as a Jupyter notebook at <a href="https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb">https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb</a>. The datasets contain the following information.</p> <p><strong>non_open.zip:</strong> it is a zipped (~5 GB unzipped) CSV file containing the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, dated October 2018. All the entity types retrieved from Crossref were aligned to one of following five categories: journal, book, proceedings, dataset, other. The open CC0 citation data we used came from the CSV dump of <a href="https://doi.org/10.6084/m9.figshare.6741422.v3">most recent release of COCI dated 12 November 2018</a>. The number of closed citations was calculated by subtracting the number of open citations to each entity available within COCI from the value “is-referenced-by-count” available in the Crossref metadata for that particular cited entity, which reports all the DOI-to-DOI citation links that point to the cited entity from within the whole Crossref database (including those present in the Crossref ‘closed’ dataset).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>doi:</em> the DOI of the publication in Crossref;</li> <li><em>type:</em> the type of the publication as indicated in Crossref;</li> <li><em>cited_by:</em> the number of open citations received by the publication according to COCI;</li> <li><em>non_open:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>croci_types.csv:</strong> it is a CSV file that contains the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, as collected in the previous CSV file, alligned in five classes depening on the entity types retrieved from Crossref: <em>journal</em> (Crossref types: journal-article, journal-issue, journal-volume, journal), <em>book</em> (Crossref types: book, book-chapter, book-section, monograph, book track, book-part, book-set, reference-book, dissertation, book series, edited book), <em>proceedings</em> (Crossref types: proceedings-article, proceedings, proceedings-series), <em>dataset</em> (Crossref types: dataset), <em>other</em> (Crossref types: other, report, peer review, reference-entry, component, report-series, standard, posted-content, standard-series).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>type:</em> the type publication between "journal", "book", "proceedings", "dataset", "other";</li> <li><em>label:</em> the label assigned to the type for visualisation purposes;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publication type according to COCI;</li> <li><em>crossref_close_cit:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>publishers_cits.csv:</strong> it is a CSV file that contains the top twenty publishers that received the greatest number of open citations. The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher</em>: the name of the publisher;</li> <li><em>doi_prefix</em>: the list of DOI prefixes used assigned by the publisher;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publications of the publisher according to COCI;</li> <li><em>crossref_close_cit</em>: the number of closed citations received by the publications of the publishers according to Crossref + COCI;</li> <li><em>total_cit</em>: the total number of citations received by the publications of the publisher (= <em>coci_open_cit</em> + <em>crossref_close_cit</em>).</li> </ul> <p><strong>20publishers_cr.csv: </strong>it is a CSV file that contains the numbers of the contributions to open citations made by the twenty publishers introduced in the previous CSV file as of 24 January 2018, according to the data available through the Crossref API. The counts listed in this file refers to the number of publications for which each publisher has submitted metadata to Crossref that include the publication’s reference list. The categories 'closed', 'limited' and 'open' refer to publications for which the reference lists are not visible to anyone outside the Crossref Cited-by membership, are visible only to them and to Crossref Metadata Plus members, or are visible to all, respectively. In addition, the file also record the total number of publications for which the publisher has submitted metadata to Crossref, whether or not those metadata include the reference lists of those publications.</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher: </em>the name of the publisher;</li> <li><em>open: </em>the number of publications in Crossref with an 'open' visibility for their reference lists;</li> <li><em>limited: </em>the number of publications in Crossref with an 'limited' visibility for their reference lists;</li> <li><em>closed: </em>the number of publications in Crossref with an 'closed' visibility for their reference lists;</li> <li><em>overall_deposited:</em> the overall number of publications for which the publisher has submitted metadata to Crossref.</li> </ul>
Crossref metadata of COCI bibliographic resources, as of November 2018 and LCC categories of the ISBN entities in the dataset
<p>The <em>all.zip</em> CSV file (zipped) contains citation counts obtained from the November 2018 dump of COCI (https://doi.org/10.6084/m9.figshare.6741422.v3) and some metadata (title, DOI, number of authors, ISBN, ISBN of the container, type of the bibliographic resource) of the related citing and cited entities obtained by using the Crossref dump downloaded in October 2018 – which is the same dump used to create the COCI data.</p> <p>In addition, it contains all the Library of Congress Classification (LCC) categories associated with each ISBN in the previous dataset (file <em>isbn_cat_lcc.csv</em>), according to the data retrieved using the services at <a href="http://classify.oclc.org/classify2/api_docs/index.html">http://classify.oclc.org/classify2/api_docs/index.html</a>. Two ancillary mapping files have been also added: one (<em>ddc_to_lcc_mapping.csv</em>) for converting a Dewey Decimal Classification (DDC) categories into LCC categories, in the case the service mentioned above returned only DDC categories for some ISBN; the other (<em>lcc_to_wos_mapping.csv</em>) to map each LCC category into the related <a href="https://images.webofknowledge.com/images/help/WOS/hp_research_areas_easca.html">Web of Science research area</a>.</p>
Dataset of the deleted DOIs extracted from the difference set between Crossref DOIs as of March 2017 and January 2021
<p><strong>Abstract</strong></p> <p>Digital Object Identifiers (DOIs) are regarded as persistent; however, they are sometimes deleted. Deleted DOIs are an important issue not only for persistent access to scholarly content but also for bibliometrics, because they may cause problems in correctly identifying scholarly articles. However, little is known about how much of deleted DOIs and what causes them. We identified deleted DOIs by comparing the datasets of all Crossref DOIs on two different dates, investigated the number of deleted DOIs in the scholarly content along with the corresponding document types, and analyzed the factors that cause deleted DOIs. Using the proposed method, 708,282 deleted DOIs were identified. The majority corresponded to individual scholarly articles such as journal articles, proceedings articles, and book chapters. There were cases of many DOIs assigned to the same content, e.g., retracted journal articles and abstracts of international conferences. We show the publishers and academic societies which are the most common in deleted DOIs. In addition, the top cases of single scholarly content with a large number of deleted DOIs were revealed. The findings of this study are useful for citation analysis and altmetrics, as well as for avoiding deleted DOIs.</p> <p> </p> <p><strong>Data Records</strong></p> <p>The data format of the dataset is JSON lines, where each line is a single record. In this dataset, we identified the deleted DOIs by from the difference set between Crossref DOIs as of March 2017 and January 2021. We note that the file "00_Non-Crossref_DOIs.jsonl.gz" is not deleted DOIs but other files are deleted DOIs. Please refer the conference paper shown in the references for details. Sample of the record is the following.</p> <ul> <li>doi -- DOI name (String), e.g., "10.xxxx/xxxx"</li> <li>whichRA -- Registration agency name or error message for the DOI name according to the “<a href="https://www.doi.org/factsheets/DOIProxy.html#whichra">whichRA?</a>.” (String). "Airiti," "Crossref," "DOI does not exist," "DataCite," "KISTI," "Public," or "mEDRA."</li> <li>redirects -- Redirected URIs for the DOI name obtained by curl command (Array of String), e.g., ["https://doi.org/10.1001/archinte.166.4.387","http://archinte.jamanetwork.com/article.aspx?doi=10.1001/archinte.166.4.387"]</li> <li>redirect_to_other_doi -- The other DOI when the DOI link redirects to. (Array of String), e.g., "[10.1001/archinte.166.4.387]"</li> <li>timestamp -- Date the data was retrieved (Datetime), "2022-01-30T02:10:45Z"</li> <li>label -- The group where the DOI belongs to. Alias DOIs," "DOIs with Deleted Description in Metadata," "DOIs without Redirects," "Defunct DOIs," "Non-Crossref DOIs," "Non-existing DOIs," or "Other DOIs."</li> </ul> <p>As for the file "04_DOIs_with_Deleted_Description_on_Metadata.jsonl.gz," additional records are available as follows.</p> <ul> <li>alias_doi -- Alias DOI name, the same as the value of "doi." (String), e.g., "10.1007/bf00400428."</li> <li>primary_doi -- Primary DOI name for the alias DOI name, the same as the first value of "redirect_to_other_doi". (String), e.g., "10.1007/bf00400429"</li> <li>container_title_of_alias_doi -- the container title for the alias DOI according to the Crossref REST API. e.g., "CrossRef Listing of Deleted DOIs."</li> <li>title_of_alias_doi -- the title for the alias DOI according to the Crossref REST API. e.g., "CrossRef Listing of Deleted DOIs."</li> <li>container_title_of_primary_doi -- the container title for the primary DOI according to the Crossref REST API. e.g., "CrossRef Listing of Deleted DOIs."</li> <li>title_of_alias_doi -- the title for the primary DOI according to the Crossref REST API. e.g., "CrossRef Listing of Deleted DOIs."</li> </ul> <p><strong>References</strong></p> <ul> <li>Kikkawa, J., Takaku, M. & Yoshikane, F. "Analysis of the deletions of DOIs: What factors undermine their persistence and to what extent?", Proceedings of the 26th International Conference on Theory and Practice of Digital Libraries (<a href="http://tpdl2022.dei.unipd.it/"><em>TPDL 2022</em></a>), (to appear), 2022.</li> </ul> <p><strong>FUNDING</strong></p> <ul> <li>JSPS KAKENHI Grant Number <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-21K21303">JP21K21303</a>, <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-22K18147">JP22K18147</a>, <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-20K12543">JP20K12543</a>, and <a href="https://kaken.nii.ac.jp/en/grant/KAKENHI-PROJECT-21K12592">JP21K12592</a></li> </ul>
Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic
<p>This data set contains supplementary material for the paper 'Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic' by Martijn Visser, Nees Jan van Eck, and Ludo Waltman. The data set provides the statistics presented in the figures in the paper.</p>
Crossref metadata statistics
<p>This dataset contains the data underlying the statistics reported in the paper '<a href="https://doi.org/10.31222/osf.io/smxe5_v2">Crossref as a source of open bibliographic metadata</a>' by Nees Jan van Eck and Ludo Waltman. The data provides insight into the availability of different metadata elements (i.e., references, abstracts, ORCIDs, author affiliations, funding information, and license information) for journal articles, book chapters, conference papers, and preprints in Crossref. The data is based on the Crossref XML Metadata Plus Snapshot from January 2025.</p> <p>The file 'Crossref_metadata_statistics.ods' contains for each combination of a publisher, a publication type, and a publication year the total number of records and the number of records for which specific metadata elements are available.</p> <p>The file 'Crossref_metadata_figures.ods' contains for each figure in the paper the statistics presented in the figure.</p>
Crossref open citations
<p>This text file contains citations between documents classified as journal article, book content, conference paper, or preprint in Crossref. Each line in the file contains the DOI of a citing document followed by the DOI of a cited document. There are 147,161,531 documents classified as journal article, book content, conference paper, or preprint in Crossref. There are 1,723,702,594 citations between these documents in Crossref. The data has been extracted from Crossref’s XML Metadata Plus Snapshot downloaded on February 10, 2025.</p>
Crossref relationships between preprints and journal articles
<p>This dataset contains preprint-journal article relationships deposited by publishers with Crossref and/or discovered by an automated preprint matching strategy. It includes preprints and journal articles deposited until the end of February 2025. The following fields are included:</p> <ul> <li>preprint DOI (string)</li> <li>journal article DOI (string)</li> <li>whether the publisher of the journal article deposited this relationship (boolean)</li> <li>whether the publisher of the preprint deposited this relationship (boolean)</li> <li>the confidence score returned by the strategy (float, empty if the strategy did not discover this relationship)</li> </ul> <p>The code of the preprint matching strategy is available <a href="https://gitlab.com/crossref/labs/marple/-/tree/main/strategies_available/preprint_sbmv">here</a>.</p>
Extraction of Crossref datadump for benchmarking purpose
<p>Extracted form a datadump provided <a href="https://www.crossref.org/blog/2022-public-data-file-of-more-than-134-million-metadata-records-now-available/">here</a>. The file contains several JSON files limiting to benchmark tools that have to process huge dataset in JSON format.</p>
Crossref Funder Registry - Mapping to top-level funding organizations
<p>The <a href="https://www.crossref.org/services/funder-registry/">Crossref Funder Registry</a> is an open registry of names and identifiers of research funding organizations. The registry has a hierarchical structure. Each organization may have relations with parent and child organizations.<br> <br> We provide a mapping of organizations in the Crossref Funder Registry to the corresponding top-level organizations. The mapping is based on version 1.34 of the Crossref Funder Registry, made available in an RDF file in the <a href="https://gitlab.com/crossref/open_funder_registry/">Crossref Funder Registry GitLab repository</a>.<br> <br> To create the mapping to top-level organizations, we first extracted the hierarchical structure of the Crossref Funder Registry based on all ‘narrower’ and ‘broader’ assignments. Using this hierarchical structure, we then identified for each of the 27,681 active organizations in the registry the corresponding top-level organizations. For most organizations, we identified a single top-level organization. However, organizations may have multiple parent organizations, and for some organizations we therefore identified multiple top-level organizations.<br> <br> Example: ‘H2020 Societal Challenges’ has ‘Horizon 2020 Framework Programme’ as its only parent organization, ‘Horizon 2020 Framework Programme’ has ‘European Commission’ as its only parent organization, and ‘European Commission’ does not have a parent organization. The top-level organization corresponding to ‘H2020 Societal Challenges’ therefore is ‘European Commission’.<br> <br> We ignored the following parent-child relations in order to avoid having cycles in the network of parent-child relations:</p> <ul> <li>Child organization: Ministry of Science and Technology of the People's Republic of China (CHN) [DOI: 10.13039/501100002855]<br> Parent organization: National Science and Technology Major Project (CHN) [DOI: 10.13039/501100018537]</li> <li>Child organization: University of Georgia (USA) [DOI: 10.13039/100007699]<br> Parent organization: College of William and Mary (USA) [DOI: 10.13039/100008277]</li> </ul> <p>We ignored the following parent-child relations because they seem to be mistakes:</p> <ul> <li>Child organization: College of William and Mary (USA) [DOI: 10.13039/100008277]<br> Parent organization: University of Georgia (USA) [DOI: 10.13039/100007699]</li> <li>Child organization: Fakulti Pertanian, Universiti Putra Malaysia (MYS) [DOI: 10.13039/501100006018]<br> Parent organization: Department of Child Health, University of Manchester (GBR) [DOI: 10.13039/100007546]</li> <li>Child organization: Pusat Penyelidikan dan Inovasi, Universiti Malaysia Sabah (MYS) [DOI: 10.13039/501100006243]<br> Parent organization: Workplace Health, Safety and Compensation Commission (CAN) [DOI: 10.13039/501100000183]</li> <li>Child organization: Workplace Health, Safety and Compensation Commission (CAN) [DOI: 10.13039/501100000183]<br> Parent organization: Shahid Beheshti University of Medical Sciences (IRN) [DOI: 10.13039/501100005851]</li> </ul>
More open abstracts? Comparing abstract coverage in Crossref and OpenAlex [dataset]
<p>Aggregated data underlying the blogpost:<strong><br><br>More open abstracts? Comparing abstract coverage in Crossref and OpenAlex<br></strong><a href="https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/">https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/</a><strong><br></strong><br>The dataset contains the following files:</p> <ul> <li><em>abstracts_crossref_openalex_202410.csv</em></li> <li><em>abstracts_crossref_openalex_data_dictionary.txt</em></li> </ul> <p>The csv file contains data on abstract coverage for Crossref DOIs in Crossref and OpenAlex, aggregated by publisher, for the top 1000 publishers in terms of number of retrieved dois. Scope is limited to publications with Crossref type 'journal-articles' and publication years 2022-2024. Variables are described in the data dictionary included in this record.<br><br>This analysis was performed using <a href="https://openknowledge.community/" rel="nofollow">Curtin Open Knowledge Initiative (COKI)</a> infrastructure, which is documented on GitHub: <a href="https://github.com/The-Academic-Observatory">https://github.com/The-Academic-Observatory</a>. Here, a number of open data sources (including Crossref, OpenAlex and OpenAIRE) are ingested into a Google Big Query environment, which can then be queried via SQL.<br><br>The following data sources were used:</p> <ul> <li> <p>Crossref (Metadata Plus snaphot 2024-10-31, Crossref member route API 2024-11-20)</p> </li> <li> <p>OpenAlex (data snapshot 2024-10-30)</p> </li> </ul> <p><br>The code used to generate the dataset is available on GitHub: <a href="https://github.com/bmkramer/more_open_abstracts">https://github.com/bmkramer/more_open_abstracts</a></p> <p> </p>
Crossref datadump extraction for benchmarking purpose of Jsonpath
<p> Several aggregation of Crossref datadump to amplify the size for benchmarking purpose.</p>
Tidy coverage data for all 9943 CrossRef members
<p>Included in this deposit are the scripts and data collected on the coverage of metadata by the CrossRef members. Collection date is March 27, 2018. The data is available in long, <a href="http://vita.had.co.nz/papers/tidy-data.pdf">tidy</a>, format. Each row depicts the coverage of one specific aspect, indicating which member and how many DOIs that member deposited. This allows for easy parsing in for example ggplot2 or dplyr.</p>
Crossref as a source of scientometric data for social & human sciences [dataset]
<p>This is a dataset used in and produced by research described in the paper titled "Crossref as a source of scientometric data for social & human sciences".</p>
Extraction of Crossref datadump for benchmarking purpose (small version)
<p> Several aggregation of Crossref datadump to amplify the size for benchmarking purpose.</p>
Extraction of Crossref datadump for benchmarking purpose (minify version)
<p> A chunk of Crossref datadump minified (removing whitespace)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.