Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

28

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

28 results for “Dataset citation”

Learn how ShareScore rates datasets ↗
zenodo52/100

Synthetic Dataset of Citation Strings in 12 Styles

<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland.&nbsp;</p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br>&nbsp; &nbsp; title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}},&nbsp;<br>&nbsp; &nbsp; author = {Iana Atanassova and Marc Bertin},<br>&nbsp; &nbsp; year = {2024},<br>&nbsp; &nbsp; booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br>&nbsp; &nbsp; address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Dataset related to the manuscript: "An open-source integrated framework for the automation of citation collection and screening in systematic reviews"

<p>Dataset related to the manuscript: &ldquo;An open-source integrated framework for the automation of citation collection and screening in systematic reviews&rdquo;, to be used together with the code stored at&nbsp;https://github.com/AD-Papers-Material/BART_SystReviewClassifier to reproduce the results.</p> <p>There are three datasets:<br> - The Record data collected from the online scientific databases;<br> - The session journal which describes the search session, i.e., how many records were collected and from which source, for each query/session pairs.<br> - The session data which is the outcome of the classification and review tasks;</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

[review paper] Sustaining the 'Frozen Footprints' of Scholarly Communication through Open Citations_Dataset

<p>This dataset belongs to the review titled&nbsp;<em>Sustaining the &lsquo;Frozen Footprints&rsquo; of Scholarly Communication through Open Citations</em>. The review explores the developments in the open citations movement, the OpenCitations infrastructure, and the Initiative for Open Citations (I4OC), providing a comprehensive overview of key milestones and initiatives.</p> <p>The dataset includes bibliographic and citation data for 174 scholarly outputs and 149 blogposts analyzed in the review. These outputs were drawn from a range of sources, including journal articles, conference proceedings, and other scholarly materials. The data has been curated to adhere to open citation principles, ensuring it is structured, separable, and openly accessible. It is provided in accordance with the licenses and terms of use of the original databases.</p> <p>Researchers, practitioners, and policymakers can use this dataset to explore the evolution of open citations and to further their understanding of the connections between scholarly works in the field of open research.</p>

opencc-by-4.0Oct 2024View details →
zenodo48/100

A dataset from a survey investigating disciplinary differences in data citation

<p><strong>GENERAL INFORMATION</strong></p> <p><em>Title of Dataset: </em>&nbsp;A dataset from a survey investigating disciplinary differences in data citation</p> <p><em>Date of data collection:</em><strong> </strong>January to March 2022</p> <p><em>Collection instrument: </em>SurveyMonkey</p> <p><em>Funding:</em> Alfred P. Sloan Foundation</p> <p><br> <strong>SHARING/ACCESS INFORMATION</strong></p> <p><em>Licenses/restrictions placed on the data:&nbsp;&nbsp;</em>These data are available under a <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0 license</a>&nbsp;</p> <p><em>Links to publications that cite or use the data:&nbsp;</em></p> <p>Gregory, K., Ninkov, A., Ripp, C., Peters, I., &amp; Haustein, S. (2022). Surveying practices of data citation and reuse across disciplines. Proceedings of the 26th International Conference on Science and Technology Indicators. <em>International Conference on Science and Technology Indicators</em>, Granada, Spain. https://doi.org/10.5281/ZENODO.6951437</p> <p>Gregory, K., Ninkov, A., Ripp, C., Roblin, E., Peters, I., &amp; Haustein, S. (2023). <em>Tracing data:<br> A survey investigating disciplinary differences in data citation.</em> Zenodo. https://doi.org/10.5281/zenodo.7555266</p> <p><br> <strong>DATA &amp; FILE OVERVIEW</strong></p> <p><em>File List</em></p> <ul> <li>Filename: MDCDatacitationReuse2021Codebookv2.pdf<br> <em>Codebook</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.csv<br> <em>Dataset format in csv</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.sav<br> <em>Dataset format in SPSS</em></li> <li>Filename: MDCDataCitationReuseSurvey2021QNR.pdf<br> <em>Questionnaire</em></li> </ul> <p><em>Additional related data collected that was not included in the current data package:&nbsp;</em>Open ended questions asked to respondents</p> <p><br> <strong>METHODOLOGICAL INFORMATION</strong></p> <p><em>Description of methods used for collection/generation of data:&nbsp;</em></p> <p>The development of the questionnaire (Gregory et al., 2022) was centered around the creation of two main branches of questions for the primary groups of interest in our study: researchers that reuse data (33 questions in total) and researchers that do not reuse data (16 questions in total). The population of interest for this survey consists of researchers from all disciplines and countries, sampled from the corresponding authors of papers indexed in the Web of Science (WoS) between 2016 and 2020.&nbsp;</p> <p>Received 3,632 responses, 2,509 of which were completed, representing a completion rate of 68.6%. Incomplete responses were excluded from the dataset. The final total contains 2,492 complete responses and an uncorrected response rate of 1.57%. Controlling for invalid emails, bounced emails and opt-outs (n=5,201) produced a response rate of 1.62%, similar to surveys using comparable recruitment methods (Gregory et al., 2020).</p> <p><em>Methods for processing the data:&nbsp;</em></p> <p>Results were downloaded from SurveyMonkey in CSV format and were prepared for analysis using Excel and SPSS by recoding ordinal and multiple choice questions and by removing missing values.</p> <p><em>Instrument- or software-specific information needed to interpret the data:&nbsp;</em></p> <p>The dataset is provided in SPSS format, which requires IBM SPSS Statistics. The dataset is also available in a coded format in CSV. The Codebook is required to interpret to values.</p> <p><br> <strong>DATA-SPECIFIC INFORMATION FOR: MDCDataCitationReuse2021surveydata</strong></p> <p><em>Number of variables:</em> 95</p> <p><em>Number of cases/rows:</em> 2,492</p> <p><em>Missing data codes:</em>&nbsp;999 &nbsp; &nbsp; &nbsp; &nbsp;Not asked</p> <p>Refer to MDCDatacitationReuse2021Codebook.pdf for detailed variable information.</p>

opencc-by-4.0Jan 2023View details →
zenodo44/100

Dataset for Machine Learning Assisted Citation Screening for Systematic Reviews

<p>The work "Machine Learning Assisted Citation Screening for Systematic Reviews" explored the problem of citation screening automation using machine-learning (ML) with an aim to accelerate the process of generating <a href="https://en.wikipedia.org/wiki/Systematic_review#:~:text=Systematic%20reviews%20are%20a%20type,synthesize%20findings%20qualitatively%20or%20quantitatively." rel="nofollow">systematic reviews</a>. Manual process of citation screening involve two reviewers manually screening the searched studies using a predefined inclusion criteria. If the study passes the "inclusion" criteria, it is included for further analysis or is excluded. As apparant through manual screening process, the work considered citation screening as a binary classification problem whereby any ML classifier could be trained to separate the searched studies into these two classes (include&nbsp;and&nbsp;exclude).</p> <p>&nbsp;</p> <p>A physiotherapy citation screening dataset was used to test automation approaches and the dataset includes the studies identified for citation screening in an update to the systematic review by Hilfiker <em>et al.</em> The dataset included titles and abstracts (citations) from 31,279 (deduplicated: 25,540) studies identified during the search phase of this SR. These studies were already manually assessed for relevance and labelled by two reviewers into two mutually exclusive labels. The uploaded file consists of 25,540 data samples, with each data sample separated by a new line. It is a tab separated file and the data in it is structured as shown below. This dataset was manually labelled into include and exclude by Hilfiker&nbsp;<em>et al.</em></p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Title</strong></td> <td><strong>PMID</strong></td> <td><strong>Abstract&nbsp;</strong></td> <td><strong>Class</strong></td> <td><strong>MeSH terms (separated by a pipe)</strong></td> </tr> <tr> <td>Structured exercise improves physical functioning in women with stages I and II breast cancer: results of a randomized controlled trial. &nbsp;</td> <td>11157015</td> <td>Abstract PURPOSE: Self-directed and supervised exercise were compared with usual care in a clinical trial designed to evaluate the effect of structured exercise on physical functioning and other dimensions of health-related quality of life in women with stages I and II breast cancer. PATIENTS AND METHODS: One hundred twenty-three women with stages I and II breast cancer completed baseline evaluations of generic and disease- and site-specific health-related quality of life, aerobic capacity, and body weight. Participants were randomly allocated to one of three intervention groups: usual care (control group), self-directed exercise, or supervised exercise. Quality of life, aerobic capacity, and body weight measures were repeated at 26 weeks...</td> <td>include or exclude</td> <td>Clinical Trial | Comparative Study | Randomized Controlled Trial | Research Support, Non-U.S. Gov't | Antineoplastic Combined Chemotherapy Protocols | Breast Neoplasms | Breast Neoplasms | Breast Neoplasms | Chemotherapy, Adjuvant | Exercise | Female | Humans | Middle Aged | Neoplasm Staging | Quality of Life | Radiotherapy, Adjuvant</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>If you use this dataset in your research, please cite our papers.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Preprint Citations in PLOS Dataset

<p>Preprints are research articles that have been published online before undergoing peer review. The role of preprints in the scientific production has been growing in recent years. Our objective is to study these practices and evaluate the differences that exist between citations to preprints and citations to peer-reviewed articles.</p> <p>This dataset contains citation contexts to preprints extracted from the PLOS dataset. We have processed all PLOS articles published up to January 2021. Preprint citations were identified by matching cited source metadata against a list of existing preprint databases. For each citation we have extracted the sentence and its position in the IMRaD structure of the article.</p> <p>The data is presented in a tsv file that contains the following columns :</p> <ul> <li>id: identifier.</li> <li>source_name: name of the preprint database where the preprint is published. In some cases source_name is &quot;preprint kw&quot; which means that it has been identified by the presence of the &quot;preprint&quot; keyword in the source metadata, but could not be linked to a known preprint database.</li> <li>jtitle: title of the PLOS journal from which the citation context is extracted.</li> <li>imrad_code: one of &quot;I&quot;, &quot;M&quot;, &quot;R&quot;, &quot;D&quot;, indicating the name of the section of the citation context in the IMRaD (Introduction, Methods, Results and Discussion) structure.</li> <li>perc: a number between 0 and 100, indicating the position of the citation context in terms of percentage of the text progression of the section in which it appears. This position has been calculated by dividing the number of the sentence of the citation context by the total number of sentences in the section.</li> <li>pub_year: publication year of the article</li> <li>sentence_text: sentence containing the citation to the preprint.</li> </ul> <p>The full description of the dataset and the processing steps to obtain it are described in:</p> <p>Bertin, Marc and Atanassova, Iana (2022). &quot;Preprint Citation Praxis in PLOS&quot;. Scientometrics.</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Datasets and results of the paper titled "Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"

<p>These&nbsp;are&nbsp;the <strong>input&nbsp;datasets</strong> and the <strong>results of the analyses</strong>&nbsp;reported on&nbsp;the paper titled <strong>&quot;Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification&quot;</strong>.</p> <p><strong>Abstract:</strong>&nbsp;</p> <p>The aim of this paper is to study the role of citation network measures in the assessment of scientific maturity. Referring to the case of the Italian national scientific qualification (ASN), we investigate if there is a relationship between citation network indices and the results of the researchers&rsquo; evaluation procedures. In particular, we want to understand if network measures can enhance the prediction accuracy of the results of the evaluation procedures beyond basic performance indices. Moreover, we want to highlight which citation network indices prove to be more relevant in explaining the ASN results, and if quantitative indices used in the citation-based disciplines assessment can replace the citation network measures in non-citation-based disciplines. Data concerning Statistics and Computer Science disciplines are collected from different sources (ASN, Italian Ministry of University and Research, and Scopus) and processed in order to calculate the citation-based measures used in this study. Following, we apply classification models to estimate the effects of network variables. We find that network measures are strongly related to the results of the ASN and significantly improve the explanatory power of the models, especially for the research fields of Statistics. Additionally, citation networks in the specific sub-disciplines are far more relevant than those in the general disciplines. Finally, results show that the citation network measures are not a substitute of the citation-based bibliometric indices.</p> <p><strong>Code</strong></p> <p>The code to collect&nbsp;and process the data used in this paper is available on GitHub at <a href="https://github.com/DigitalDataLab/ASN16-18_CitationNetwork">https://github.com/DigitalDataLab/ASN16-18_CitationNetwork</a><strong>.</strong>&nbsp;</p> <p><strong>Dataset description</strong></p> <p>The files&nbsp;<strong>AdjacencyMatrix_01B1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_09H1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D2.csv</strong> and&nbsp;<strong>AdjacencyMatrix_13D3.csv</strong> are the&nbsp;citation matrices for Italian academics (i.e. ASN candidates and permanent positions in the Italian academic system) in the Recruitment Fields (RFs) 01/B1, 09/H1, 13/D1, 13/D2 and&nbsp;13/D3, respectively.</p> <p>The files&nbsp;<strong>AdjacencyMatrix_CS.csv</strong>&nbsp;and&nbsp;<strong>AdjacencyMatrix_ST.csv</strong> are the citation matrices for the Italian academics in the Computer Science disciplines (i.e. RFs 01/B1 and 09/H1) and the Statistical disciplines (i.e. RFs 13/D1,&nbsp;13/D2 and&nbsp;13/D3), respectively.</p> <p>The files&nbsp;<strong>CS_01B1_1.csv,&nbsp;CS_09H1_1.csv, ST_13D1_1.csv,&nbsp;ST_13D2_1.csv</strong> and&nbsp;<strong>ST_13D3_1.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Full Professor (FP) level.</p> <p>The files&nbsp;<strong>CS_01B1_2.csv,&nbsp;CS_09H1_2.csv, ST_13D1_2.csv,&nbsp;ST_13D2_2.csv</strong> and&nbsp;<strong>ST_13D3_2.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Associate Professor (AP) level.</p> <p>The file&nbsp;<strong>Codebook.pdf</strong>&nbsp;is the codebook of the previous ten files.</p> <p>The file <strong>Appendix.pdf</strong> contains the final results of the stepwise logistic regressions computed for each level (i.e. Full Professor and Associate Professor) and Recruitment Field in the Computer Science and Statistics disciplines.</p> <p>The file&nbsp;<strong>NormalityAssessment.pdf</strong>&nbsp;contains the&nbsp;normality assessment of citation network indices.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo44/100

BIP! NDR (NoDoiRefs): a dataset of citations from papers without DOIs in computer science conferences and workshops

<h2>Overview</h2> <p>In the field of Computer Science, conference and workshop papers serve as important contributions, carrying substantial weight in research assessment processes, compared to other disciplines. However, a considerable number of these papers are not assigned a Digital Object Identifier (DOI), hence their citations are not reported in widely used citation datasets like OpenCitations and Crossref, raising limitations to citation analysis. While the Microsoft Academic Graph (MAG) previously addressed this issue by providing substantial coverage, its discontinuation&nbsp; has created a void in available data.</p> <p>BIP! NDR aims to alleviate this issue and enhance the research assessment processes within the field of Computer Science. To accomplish this, it leverages a workflow that identifies and retrieves Open Science papers lacking DOIs from the DBLP Corpus, and by performing text analysis, it extracts citation information directly from their full text.</p> <p>The current version of the dataset contains&nbsp;<em>~4.3M citations</em> made by approximately <em>211K open access Computer Science conference or workshop papers</em> that, according to DBLP, do not have a DOI. The DBLP snapshot used for this version was the one released on <em>September 2025</em>.&nbsp;</p> <h2>Dataset files</h2> <h3>1. Core Non-DOI Citation Dataset - bip_ndr_{version}.tar.gz</h3> <p>The dataset is formatted as a JSON Lines (JSONL) file (one JSON Object per line) to facilitate file splitting and streaming.&nbsp;</p> <p>Each JSON object has three main fields:</p> <ul> <li> <p>&ldquo;_id&rdquo;: a unique identifier,</p> </li> <li> <p>&ldquo;citing_paper&rdquo;, the &ldquo;dblp_id&rdquo; of the citing paper,</p> </li> <li> <p>&ldquo;cited_papers&rdquo;: array containing the objects that correspond to each reference found in the text of the &ldquo;citing_paper&rdquo;; each object may contain the following fields:</p> <ul> <li> <p>&ldquo;dblp_id&rdquo;: the &ldquo;dblp_id&rdquo; of the cited paper. Optional - this field is required if a &ldquo;doi&rdquo; is not present.</p> </li> <li> <p>&ldquo;doi&rdquo;: the doi of the cited paper. Optional - this field is required if a &ldquo;dblp_id&rdquo; is not present.</p> </li> <li> <p>&ldquo;bibliographic_reference&rdquo;: the raw citation string as it appears in the citing paper.</p> </li> </ul> </li> </ul> <p>Changes from previous version:</p> <ul> <li>Added more papers from DBLP.</li> </ul> <h3>2. Citation Intents Dataset - bip_ndr_ci_{version}.tar.gz</h3> <p>This file enriches the BIP! NDR dataset with citation-level intent classification.<br>It preserves the same base structure of the previous file, while adding a nested array of "citations" with each element of "cited_papers".</p> <p>Each "citation" provides the local textual context, section, and intent of the citation in the following format:</p> <ul> <li>"citation_id": Unique identifier in the format {citing_id}&gt;{cited_id}_CIT{index} linking the citing and cited entities.</li> <li>"section": The section of the citing paper where the citation occurs (e.g., Introduction, Methods, Results).</li> <li>"intent": Inferred purpose of the citation based on textual context (see classification schema below).</li> </ul> <p>The "intent" field follows the SciCite classification schema, which categorizes citations into three high-level functional types:</p> <ol> <li>background information: The citation states, mentions, or points to the background information giving more context about a problem, concept, approach, topic, or importance of the problem in the field.</li> <li>method: Making use of a method, tool, approach or dataset.</li> <li>results comparison: Comparison of the paper's results/findings with the results/findings of other work.</li> </ol> <p>The classification is done with the <a href="https://huggingface.co/sknow-lab/Qwen2.5-14B-CIC-SciCite">Qwen2.5-14B-CIC-SciCite fine-tuned Large Language Model, published by Athena RC</a>.&nbsp;</p> <p>Changes from previous version:&nbsp;</p> <ul> <li>Added more papers with intent</li> </ul>

opencc-zeroMay 2023View details →
zenodo44/100

Coverage of DOAJ journals' citations through OpenCitations - Result DataSet

<p>The dataset contains:&nbsp;</p> <ul> <li><strong>by_journal.json</strong>: a file containing all information extracted by Open Citations about DOAJ journals divide by year and journal name. Inside the file, the metadata about the journal are:&nbsp; <ul> <li>ISSN</li> <li>EISSN</li> <li>number of articles overall in the journal</li> <li>subject(s)&nbsp;</li> <li>number of citations received</li> <li>number of citations done</li> <li>ratio between citations done and received</li> <li>number of citations received from DOAJ journals</li> <li>number of citations done to DOAJ journals</li> <li>ratio between citations done to and received from DOAJ journals.</li> </ul> </li> </ul> <ul> <li><strong>normal.json</strong>: a file containing all information extracted from Open Citations about DOAJ journals divided only by year. Inside the file, the data by year are: <ul> <li>number of citations received.</li> <li>number of citations done.</li> <li>ratio between citations done and received.</li> <li>number of self-citations made by DOAJ inside Open Citations.</li> <li>ratio between the self-citation and the total citations received and done by DOAJ.</li> </ul> </li> </ul> <ul> <li><strong>errors.json</strong>: a file containing the count of all errors obtained from computations. Inside the file: <ul> <li>errors about records that don&#39;t have any specified date (null dates).</li> <li>errors about records that have impossible dates (wrong dates).</li> <li>errors about articles that don&#39;t have any specified Dois.</li> <li>errors about Open Citations records that don&#39;t have any Dois in the citing or cited fields.</li> </ul> </li> <li><strong>DOAJ_metrics.json</strong>: a file containing metrics about DOAJ and Open Citations, obtained by computations. Inside the file are these fields: <ul> <li>number of journals with dois.</li> <li>number of articles which have been processed during computations.</li> <li>number of used Dois. All dois (with no repetition) which are used for the adding journal operation.</li> <li>number of repeated Dois. All dois which are repeated inside the same or in another journal.</li> <li>number of accepted Dois. All articles (with repetition) which have both a defined journal and a defined doi.</li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dataset Citation and Re-use Data

<p>This dataset includes processed citation data for datasets recorded in OpenAlex as of May 2022. It identifies self-citations to these datasets at the individual, institutional, and country level, and includes domain classifications of the citing works using the Science-Metrix classifications.</p>

opencc-by-4.0Aug 2023View details →
zenodo40/100

PatCit: A Comprehensive Dataset of Patent Citations

<p><em><strong>patCit:&nbsp;A Comprehensive Dataset of Patent Citations</strong></em>&nbsp;[<a href="https://tinyletter.com/patcit">Newsletter</a>,&nbsp;<a href="https://github.com/cverluise/PatCit">GitHub</a>]</p> <p>Patents are at the crossroads of many innovation nodes: science, industry, products, competition, etc. Such interactions can be identified through citations&nbsp;<em>in a broad sense</em>.</p> <p>It is now common to use front-page patent citations to study some aspects of the innovation system. However, <strong>there is much more buried in the Non Patent Literature (NPL) citations and in the patent text itself</strong>.&nbsp;<strong>patCit extracts and structures these citations.</strong></p> <blockquote> <p>Want to know more? Read patCit&nbsp;<a href="https://docs.google.com/presentation/d/11COlz64EZn8PipXvnDBBZI_bnDD0fpm6tyx1_EqD6lU/edit?usp=sharing">academic presentation</a>&nbsp;or dive into usage and technical guides on patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p> </blockquote> <p><strong><em>IN PRACTICE</em></strong></p> <p>At patCit, we are building a&nbsp;<em>comprehensive</em>&nbsp;dataset of patent citations to help the community explore this&nbsp;<em>terra incognita</em>. patCit has the following features:</p> <ul> <li>global coverage</li> <li>front-page and in-text citations</li> <li>all categories&nbsp;of NPL documents</li> </ul> <p><strong><em>Front-page</em></strong></p> <p>patCit builds on&nbsp;<a href="https://www.epo.org/searching-for-patents/data/bulk-data-sets/docdb.html#tab-1">DOCDB</a>, the largest database of Non Patent Literature (NPL) citations. First, we deduplicate this corpus and organize it into 10 categories (bibliographical reference, database, norm &amp; standard, etc). Then, we design and apply category specific information extraction models using&nbsp;<a href="https://github.com/explosion/spaCy">spaCy</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p><strong><em>In-text</em></strong></p> <p>patCit builds on Google Patents corpus of&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patents-public-data&amp;d=patents&amp;t=publications&amp;page=table">USPTO full-text patents</a>. First, we extract patent and bibliographical reference citations. Then, we parse detected in-text citations into a series of category dependent attributes using&nbsp;<a href="https://github.com/kermitt2/grobid">grobid</a>. Patent citations are matched with a standard publication number using the Google Patents&nbsp;<a href="https://patents.google.com/api/match">matching API</a>&nbsp;and bibliographical references are matched with a DOI using&nbsp;<a href="https://github.com/kermitt2/biblio-glutton">biblio-glutton</a>. Eventually, when possible, we enrich the data using external domain specific high quality databases (e.g. Crossref for bibliographical references).</p> <p>&nbsp;</p> <p><strong>FAIR</strong></p> <p><strong>Find</strong>&nbsp;- The patCit dataset is available on&nbsp;<a href="https://console.cloud.google.com/bigquery?project=patcit-public-data&amp;p=patcit-public-data&amp;page=project">BigQuery</a>&nbsp;in an interactive environment. For those who have a smattering of SQL, this is the perfect place to explore the data. It can also be downloaded on&nbsp;<a href="https://zenodo.org/record/3710994">Zenodo</a>.</p> <p><strong>Interoperate</strong>&nbsp;- Interoperability is at the core of patCit ambition. We take care to extract unique identifiers whenever it is possible to enable data enrichment for domain specific high quality databases. This includes the DOI, PMID and PMCID for bibliographical references, the Technical Doc Number for standards, the Accession Number for Genetic databases, the publication number for PATSTAT and Claims, etc. See specific table for more details.</p> <p><strong>Reproduce</strong>&nbsp;- Our <a href="https://github.com/cverluise/PatCit">gitHub</a> repository is the project factory. You can learn more about data recipes and models on the patCit&nbsp;<a href="https://cverluise.github.io/PatCit/">documentation website</a>.</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Merged citation library dataset of 3 bibliographic databases for use in testing deduplication tools

<p>Endnote xml file of complete search results - not de-duplicated.&nbsp;</p> <p>Search query (cannabis based medicines, endocannabinoid system modulators and cannabinoids tested in animal models of pathological or injury-related persistent pain) conducted on 9 April 2019.&nbsp;&nbsp;</p> <p>Searched databases: Web of Science, Embase and PubMed</p> <p>The citations retrieved&nbsp;from the searches of each database have been merged together, there is a total of 15216 citations.&nbsp;&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo40/100

COVID-19++: A Citation-Aware Covid-19 Dataset for the Analysis of Research Dynamics

<p>COVID-19++ is a citation-aware COVID-19 dataset for the analysis of research dynamics. In addition to primary COVID-19 related articles and preprints from 2020, it includes citations and the metadata of first-order cited work. All publications are annotated with MeSH terms, either from the ground truth, or via ConceptMapper, if no ground truth was available.&nbsp;</p> <p>The data is organized in CSV files</p> <p>- Paper metadata (paper_id, publdate, title, data_source): paper.csv</p> <p>- Annotation data, mapping paper_id to MeSH terms: annotation.csv&nbsp;</p> <p>- Authorship data, mapping paper_id to author, optionally with ORCID: authorship.csv<br> - Paired DOIs of citing and cited papers: references.csv</p> <p>The column data source within the paper metadata has the value KE (for metadata from ZB MED KE), PP (for preprints) or CR (for cited resources from CrossRef)<br> &nbsp;</p> <p>This work was supported by BMBF within the programme ``Quantitative Wissenschaftsforschung&#39;&#39; under grant numbers 01PU17013A, 01PU17013B, 01PU17013C.<br> &nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Bibliographic data on datasets affiliated to Poznan University of Technology and indexed in Data Citation Index (retrieved by Web of Science service in January 2023))

<p>The file contains the number of datasets published by the researchers affiliated to Poznan University of Technology and indexed in Data Citation Index provided by Web of Science (database updated 10.01.2023). The Search was performed using the name of institution in the &#39;Affiliation&#39; field. Dataset contains two files in two diffrent formats: plain text and xls.</p>

opencc-byJan 2023View details →
zenodo36/100

Coronavirus Open Citations Dataset

<p>The Coronavirus Open Citations Dataset curated by OpenCitations currently contains (as of 16 May 2020) information about 189,697 citations and about the 49,719 citing or cited articles involved in these citations. A subset of these data (stored in the &quot;_partial.json&quot; files introduced below) is used for creating the visualization available at <a href="https://opencitations.github.io/coronavirus/">https://opencitations.github.io/coronavirus/</a>.</p> <p>Each item in both the JSON files storing citations (&#39;citations_full.json&#39; and &#39;citations_partial.json&#39;) contains the following fields:</p> <ul> <li>&quot;id&quot;, a numeric identifier of the citation;</li> <li>&quot;source&quot;, the citing entity;</li> <li>&quot;target&quot;, the cited entity.</li> </ul> <p>Each item in both the JSON files storing article metadata (&#39;metadata_full.json&#39; and &#39;metadata_partial.json&#39;) contains the following fields:</p> <ul> <li>&quot;id&quot;, the DOI of the article;</li> <li>&quot;author&quot;, the surname of all the authors of the article;</li> <li>&quot;year&quot;, the year of publication;</li> <li>&quot;title&quot;, the title of the article;</li> <li>&quot;source_title&quot;, the title of the venue where the article has been published.</li> </ul> <p>In addition, any item in the file &#39;metadata_partial.json&#39; contains another field:</p> <ul> <li>&quot;count&quot;: the overall number of citations that the article received.</li> </ul>

opencc-zeroApr 2020View details →
zenodo36/100

Characterizing Highly Cited Papers in Mass Cytometry through H-Classics: WoS dataset and citation report

<p>Dataset and citation report extracted from Web of Science (WoS) used to characterize highly cited papers in mass cytometry research field from 2010 to 2019.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Citation network dataset covering the work of RP Millar and its citing literature

<p>The following describes the citation network datasets that underpins the manuscript &ldquo;A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research&rdquo; [1].</p> <p><strong>Data collection</strong></p> <p>We retrieved data from the Web of Science Core Collection under the University of Edinburgh&rsquo;s subscription in January 2024. We sought to retrieve all indexed papers of Professor Robert P. Millar (RPM). We searched the following <em>AU = (Millar, R)</em>, and then retrieved records that corresponded to his WoS profile (n=428 records) and an additional 49 paper that were authored by Robert but had not been included in his WoS record &ndash; validating the records against a CV of his published works.</p> <p>We retrieved the full citation history as record by Web of Science to these papers from other indexed records. The 477 RPM papers had been cited 21,677 times by 11,138 documents by date of retrieval, and removing self-citations left 19,256 citations by 10,719 documents. We then retrieved all metadata from WoS concerning the 477 RPM papers and the 10,719 citation papers, resulting in a dataset covering 11,196 documents.</p> <p><strong>Citation network dataset</strong></p> <p>We constructed a citation network dataset by parsing data from each paper&rsquo;s full bibliography consisting of:</p> <p>i. &lsquo;Edge-list&rsquo; that records citation links from a citing to a cited document. This is constructed by assigning unique IDs to each retrieved paper and to every unique reference string contained in their bibliographies. The edge list is composed of a &lsquo;Source&rsquo; column that contains the ID of the <em>citing</em> document and a &lsquo;Target&rsquo; column containing the IDs of its citations, with one record per row. Given that we were only interested in citations between the WoS retrieved documents, we discarded any reference string that represented a document outwith our search.</p> <p>ii. &lsquo;Node-attribute list&rsquo; that contains the ID, with relevant metadata contained in adjacent columns to identify documents, including authors, title of publication, journal, year of publication. We also parsed into this dataset the WoS full citation count for each paper and the total number of references in the bibliographies of each paper.</p> <p>This results in a dataset containing 11,196 nodes and 115,834 edges between nodes. We removed a total of 67 papers for which metadata was incomplete and/or corrupted. We further focussed on the largest interconnected component, removing nodes with no connections (isolates) or smaller components that were detached from the main network. We excluded papers &lt;10 references to remove meeting abstracts and other minor journal items, and papers not published in English. This resulted in a final dataset containing 10,901 nodes and 113,742 edges, and it is this dataset that we share as it is the basis for the analyses within the paper.</p> <p><strong>Description of dataset variables</strong></p> <p><strong>&lsquo;RPM_Edgelist.csv&rsquo;</strong> is a comma-separate values file that consists of all 113,742 citations between the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Source</em>&rsquo;, the unique identifier for the <em>citing </em>document</li> <li>&lsquo;<em>Target&rsquo;</em>, the unique identifier for the <em>cited</em> document</li> <li>&lsquo;<em>Syr</em>&rsquo;, the year of publication of the <em>citing</em> document</li> <li>&lsquo;<em>Tyr</em>&rsquo;, the year of publication of the <em>cited</em> document</li> <li>&lsquo;<em>SC</em>&rsquo;, the cluster ID of the <em>citing</em> document</li> <li>&lsquo;<em>TC</em>&rsquo;, the cluster ID of the <em>cited </em>document</li> </ul> <p><strong>&lsquo;RPM_Nodelist.csv&rsquo;</strong> is a comma-separate values file that consists of the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Id</em>&rsquo;, the unique ID assigned to a document that corresponds with the edgelist</li> <li>&lsquo;<em>Reference string</em>&rsquo;, the reference string of the document</li> <li>&lsquo;<em>WoS ID</em>&rsquo;, the unique accession number assigned to a document by the Web of Science. These can be used to query WoS to find further data on all papers via the &lsquo;UT= &rsquo; field tag.</li> <li>&lsquo;<em>Authors</em>&rsquo;, all authors formatted by full last name and initials</li> <li>&lsquo;<em># of authors&rsquo;</em>, number of authors</li> <li>&lsquo;<em>Title</em>&rsquo;, title of document</li> <li>&lsquo;<em>Publication year</em>&rsquo;, publication year of document</li> <li>&lsquo;<em>Document type</em>&rsquo;, document type defined by WoS (e.g. article, review, etc.)</li> <li>&lsquo;<em>Total references</em>&rsquo;, total number of references within a documents bibliography as recorded by WoS</li> <li>&lsquo;<em>Total WoS citations</em>&rsquo;, total number of citations recorded to a document from other documents indexed in the Web of Science</li> <li>&lsquo;<em>Indegree</em>&rsquo;, total number of within network citations (i.e. counting only citations from other papers retrieved by our query)</li> <li>&lsquo;<em>Outdegree</em>&rsquo;, total number of within network references (i.e. counting only reference to other papers retrieved by our query)</li> <li>&lsquo;<em>Degree</em>&rsquo;, total number of node connections (i.e. indegree + outdegree)</li> <li>&lsquo;<em>Class</em>&rsquo;, variable used to distinguish between RPM&rsquo;s publications (&lsquo;RPM&rsquo;) and the citing documents (&lsquo;CITE&rsquo;)</li> <li>&lsquo;<em>Cluster</em>&rsquo;, provides the cluster membership number as discussed within the manuscript. This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.67 | 25 clusters).</li> </ul> <p><strong>References</strong></p> <p>[1] Leng, R. I., Leng. G. (Under review). A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research. <em>J. Neuroendocrinol</em></p> <p>All bibliographic data included in this study are derived originally from Clarivate&trade; (Web of Science&trade;) and downloaded in January 2024. &copy; Clarivate 2024. All rights reserved.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Dataset for Spence et al., "Availability of study protocols for randomized trials published in high-impact medical journals: cross-sectional analysis" (CITATION)

<p>Contains our extraction sheets (as SAS data files), code to calculate the values in the tables in our manuscript, and a supplemental file with additional notes on methods used in our study.</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Dataset for Citation Network Analysis on Quality Cues for Meat Purchases

<p>Data are retrieved from two databases permissible to the Vosviewer&reg; software: Scopus and Web of Science. Three sets of Boolean search strings are used to find articles in the databases:</p> <p>(i) [<em>(&ldquo;certification&rdquo; AND &ldquo;meat&rdquo;) OR (&ldquo;certification&rdquo; AND &ldquo;dairy&rdquo;)</em>]</p> <p>(ii) [<em>(&ldquo;credence&rdquo;&nbsp;AND&nbsp;meat) OR&nbsp;(&ldquo;experience&rdquo;&nbsp;AND meat)</em>]<em> </em></p> <p><em>(iii) </em>[<em>(&ldquo;meat&rdquo; AND &ldquo;intrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;extrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;quality cue&rdquo;)</em>].</p> <p>The searches targeted the title, abstract and keywords sections for the Scopus database, and the topic section for the Web of Science database. The search was limited to peer-reviewed articles published in the English language; no timeline restrictions were set for the initial search.&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"

<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record