Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.7.1
Dataset results
76 results for “data citation”
Data on a citation context analysis focusing on natural sciences and social sciences and humanities
<p>This dataset contains data on citation context analysis between natural sciences (NS) and social sciences and humanities (SSH). In particular, the data were created through manual coding of each citation between papers related to SDG7 (renewable energy) and SDG13 (climate change) and papers cited by them. This dataset consists of 9 files, associated with the article: Nishikawa, K. How and why are citations between disciplines made? A citation context analysis focusing on natural sciences and social sciences and humanities. Scientometrics (2023). <a href="https://doi.org/10.1007/s11192-023-04664-y">https://doi.org/10.1007/s11192-023-04664-y</a></p> <p> </p> <p>The files are numbered as follows:</p> <ul> <li>00 – README</li> <li>01 – Data by citation pair for SDG7 (original)</li> <li>02 – Data by citation pair for SDG13 (original)</li> <li>03 – Data by mention location for SDG7 (original)</li> <li>04 – Data by mention location for SDG13 (original)</li> <li>05 – Data by citation pair for SDG7 (additional)</li> <li>06 – Data by citation pair for SDG13 (additional)</li> <li>07 – Data by mention location for SDG7 (additional)</li> <li>08 – Data by mention location for SDG13 (additional)</li> </ul> <p>See README for more information.</p>
IPBES Data Management Tutorials - Session 5.5: References and citation manager: Zotero
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>This session, <em>References and citation manager: Zotero, </em>reviews why IPBES recommends Zotero to manage references and provides links to key resources.</p>
Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics
<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv. </p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint’s title, author, abstract, subject category and the arXiv ID (the arXiv’s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received <em>c</em><sub>pre</sub> and <em>c</em><sub>pub</sub> citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as <em>c</em><sub>pre</sub> + <em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (‘astro-ph’), Computer Science (‘comp-sci’), Condensed Matter Physics (‘cond-mat’), High Energy Physics (‘hep’), Mathematics (‘math’) and Other Physics (‘oth-phys’), which collectively accounted for 98% of all the eprints. Those eprints tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints. </p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p> </p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> : arXiv ID</li> <li><strong>category</strong> : Research discipline</li> <li><strong>pre_year</strong> : Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> : Year of DOI acquisition</li> <li><strong>c_tot</strong> : No. of citations acquired during 1991–2019</li> <li><strong>c_pre</strong> : No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> : No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong> (<em>yyyy</em> = 1991, …, 2019) : No. of citations acquired in the year <em>yyyy</em> (with ‘<em>yyyy</em>’ running from 1991 to 2019)</li> <li><strong>gamma</strong> : The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> : The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (γ; ‘gamma’) and that of the standardised citation index (γ*; ‘gamma_star’) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times. </p> <p> </p> <p><strong>Data files</strong></p> <p>A comma-separated values file (‘<strong>arXiv_impact.csv</strong>’) and a Stata file (‘<strong>arXiv_impact.dta</strong>’) are provided, both containing the same information.</p> <p> </p>
Citation network data sets for 'Oxytocin – a social peptide? Deconstructing the evidence'
<p><strong>Introduction</strong></p> <p>This note describes the data sets used for all analyses contained in the manuscript 'Oxytocin - a social peptide?’<a href="#_ftn1">[1]</a> </p> <p><strong>Data Collection</strong></p> <p>The datasets described here were originally retrieved from Web of Science (WoS) Core Collection via the University of Edinburgh’s library subscription <a href="#_ftn2">[2]</a>. The aim of the original study for which these data were gathered was to survey peer-reviewed primary studies on oxytocin and social behaviour. To capture relevant papers, we used the following query:</p> <p><em>TI = (“oxytocin” OR “pitocin” OR “syntocinon”) AND TS = (“social*” OR “pro$social” OR “anti$social”)</em></p> <p>The final search was performed on the 13 September 2021. This returned a total of 2,747 records, of which 2,049 were classified by WoS as ‘articles’. Given our interest in primary studies <em>only</em> – articles reporting original data – we excluded all other document types. We further excluded all articles sub-classified as ‘book chapters’ or as ‘proceeding papers’ in order to limit our analysis to primary studies published in peer-reviewed academic journals. This reduced the set to 1,977 articles. All of these were published in the English language, and no further language refinements were unnecessary.</p> <p>All available metadata on these 1,977 articles was exported as plain text ‘flat’ format files in four batches, which we later merged together via Notepad++. Upon manually examination, we discovered examples of papers classified as ‘articles’ by WoS that were, in fact, reviews. To further filter our results, we searched all available PMIDs in PubMed (1,903 had associated PMIDs - ~96% of set). We then filtered results to identify all records classified as ‘review’, ‘systematic review’, or ‘meta-analysis’, identifying 75 records <a href="#_ftn3">[3]</a> (thus, ~4% of records classified by WoS were classified as reviews in PubMed). After examining a sample and agreeing with the PubMed classification, these were removed these from our dataset - leaving a total of 1,902 articles.</p> <p>From these data, we constructed two datasets via parsing out relevant reference data via the Sci2 Tool <a href="#_ftn4">[4]</a>. First, we constructed a ‘node-attribute-list’ by first linking unique reference strings (‘Cite Me As’ column in WoS data files) to unique identifiers, we then parsed into this dataset information on the identify of a paper, including the title of the article, all authors, journal publication, year of publication, total citations as recorded from WoS, and WoS accession number. Second, we constructed an ‘edge-list’ that records the citations from a <em>citing paper</em> in the ‘Source’ column and identifies the <em>cited paper</em> in the ‘Target’ column, using the unique identifies as described previously to link these data to the node-attribute-list.</p> <p>We then constructed a network in which papers are nodes, and citation links between nodes are directed edges between nodes. We used Gephi Version 0.9.2 <a href="#_ftn5">[5]</a> to manually clean these data by merging duplicate references that are caused by different reference formats or by referencing errors. To do this, we needed to retain both all retrieved records (1,902) as well as including <em>all</em> of their references to papers whether these were included in our original search or not. In total, this produced a network of 46,633 nodes (unique reference strings) and 112,520 edges (citation links). Thus, the average reference list size of these articles is ~59 references. The mean indegree (within network citations) is 2.4 (median is 1) for the entire network reflecting a great diversity in referencing choices among our 1,902 articles.</p> <p>After merging duplicates, we then restricted the network to include <em>only</em> articles fully retrieved (1,902), and retrained <em>only</em> those that were connected together by citations links in a large interconnected network (i.e. the largest component). In total, 1,892 (99.5%) of our initial set were connected together via citation links, meaning a total of ten papers were removed from the following analysis – and these were neither connected to the largest component, nor did they form connections with one another (i.e. these were ‘isolates’).</p> <p>This left us with a network of 1,892 nodes connected together by 26,019 edges. <strong><em>It is this network that is described by the ‘node-attribute-list’ and ‘edge-list’ provided here</em></strong>. This network has a mean in-degree of 13.76 (median in-degree of 4). By restricting our analysis in this way, we lose 44,741 unique references (96%) and 86,501 citations (77%) from the full network, but retain a set of articles tightly knitted together, all of which have been fully retrieved due to possessing certain terms related to oxytocin AND social behaviour in their title, abstract, or associated keywords.</p> <p>Before moving on, we calculated indegree for all nodes in this network – this counts the number of citations to a given paper from other papers within this network – and have included this in the <em>node-attribute-list</em>. We further clustered this network via modularity maximisation via the Leiden algorithm <a href="#_ftn6">[6]</a>. We set the algorithm to resolution 1, and allowed the algorithm to run over 100 iterations and 100 restarts. This gave <em>Q</em>=0.43 and identified seven clusters, which we describe in detail within the body of the paper. We have included cluster membership as an attribute in the node-attribute-list.</p> <p>For additional analysis, we also analysed the full reference list data to examine the most commonly cited references between 2016 and 2021 - the results of this are described in OTSOC_Cited_2016-2021.csv. This takes the reference lists of all retrieved papers within the network and examines their full reference lists (including references to other papers not contained within the network). These data were cleaned by matching DOIs and manual cleansing. </p> <p><strong>Data description</strong></p> <p>We include here two network datasets: (i) ‘OTSOC-node-attribute-list.csv’ consists of the attributes of 1,892 primary articles retrieved from WoS that include terms indicating a focus on oxytocin and social behaviour; (ii) ‘OTSOC-edge-list.csv’ records the citations between these papers. Together, these can be imported into a range of different software for network analysis; however, we have formatted these for ease of upload into Gephi 0.9.2. Finally, we include (iii) 'OTSOC_Cited_2016-2021' that lists all papers cited by >10 papers in the OTSOC network following any analysis of the bibliographies of retrieved papers. Below, we detail their contents:</p> <p><strong>1. ‘OTSOC-node-attribute-list.csv’</strong> is a comma-separate values file that contains all node attributes for the citation network (n=1,892) analysed in the paper. The columns refer to:</p> <p><em>Id</em>, the unique identifier</p> <p><em>Label</em>, the reference string of the paper to which the attributes in this row correspond. This is taken from the ‘Cite Me As’ column from the original WoS download. The reference string is in the following format: last name of first author, publication year, journal, volume, start page, and DOI (if available). </p> <p><em>Wos_id</em>, unique Web of Science (WoS) accession number. These can be used to query WoS to find further data on all papers via the ‘UT= ’ field tag.</p> <p><em>Title</em>, paper title.</p> <p><em>Authors</em>, all named authors.</p> <p><em>Journal, </em>journal of publication.</p> <p><em>Pub_year</em>, year of publication.</p> <p><em>Wos_citations</em>, total number of citations recorded by WoS Core Collection to a given paper as of 13 September 2021</p> <p><em>Indegree</em>, the number of within network citations to a given paper, calculated for the network shown in Figure 1 of the manuscript.</p> <p><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Figure 1). This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.43|7 clusters)</p> <p><strong>2. ‘OTSOC-edge -list.csv’</strong> is a comma-separated values file that contains all citation links between the 1,892 articles (n=26,019). The columns refer to:</p> <p><em>Source</em>, the unique identifier of the citing paper.</p> <p><em>Target, </em>the unique identifier of the cited paper.</p> <p><em>Type, </em>edges are ‘Directed’, and this column tells Gephi to regard all edges as such.</p> <p><em>Syr_date, </em>this contains the date of publication of the citing paper.</p> <p><em>Tyr_date, </em>this contains the date of publication of the cited paper.</p> <p><strong>3. 'OTSOC_Cited_2016-2021.csv'</strong> is a comma-separated values file that contain citations to all cited references that were cited by at least 10 of the retrieved papers within the OTSOC network published from 2016 onwards. The columns refer to: </p> <p><em>Reference, </em>the cited reference string extracted from the bibliographies of retrieved papers.</p> <p><em>Publication year, </em>the publication year of the cited reference.</p> <p><em>DOI</em>, the DOI of the cited reference. </p> <p><em>indegree_2016, </em>the total number of citations to a cited reference from papers published in 2016 and contained within the OTSOC network. </p> <p><em>indegree_2017, </em>the total number of citations to a cited reference from papers published in 2017 and contained within the OTSOC network. </p> <p><em>indegree_2018, </em>the total number of citations to a cited reference from papers published in 2018 and contained within the OTSOC network. </p> <p><em>indegree_2019, </em>the total number of citations to a cited reference from papers published in 2019 and contained within the OTSOC network. </p> <p><em>indegree_2020, </em>the total number of citations to a cited reference from papers published in 2020 and contained within the OTSOC network. </p> <p><em>indegree_2021, </em>the total number of citations to a cited reference from papers published in 2021 and contained within the OTSOC network. </p> <p><em>total indegree 2016-21</em>, the total number of citation to a cited reference from papers published between 2016-2021 and contained within the OTSOC network. </p> <p><strong>Software recommended for analysis</strong></p> <p>Gephi version 0.9.2 was used for the visualisations within the manuscript, and both files can be read and into Gephi without modification.</p> <p><strong>Notes</strong></p> <p><a href="#_ftnref1">[1]</a> Leng, G., Leng, R. I., Ludwig, M. (Submitted). Oxytocin – a social peptide? Deconstructing the evidence.</p> <p><a href="#_ftnref2">[2]</a> Edinburgh University’s subscription to Web of Science covers the following databases: (i) Science Citation Index Expanded, 1900-present; (ii) Social Sciences Citation Index, 1900-present; (iii) Arts & Humanities Citation Index, 1975-present; (iv) Conference Proceedings Citation Index- Science, 1990-present; (v) Conference Proceedings Citation Index- Social Science & Humanities, 1990-present; (vi) Book Citation Index– Science, 2005-present; (vii) Book Citation Index– Social Sciences & Humanities, 2005-present; (viii) Emerging Sources Citation Index, 2015-present.</p> <p><a href="#_ftnref3">[3]</a> For those interested, the following PMIDs were identified as ‘articles’ by WoS, but as ‘reviews’ by PubMed: ‘34502097’ ‘33400920’ ‘32060678’ ‘31925983’ ‘31734142’ ‘30496762’ ‘30253045’ ‘29660735’ ‘29518698’ ‘29065361’ ‘29048602’ ‘28867943’ ‘28586471’ ‘28301323’ ‘27974283’ ‘27626613’ ‘27603523’ ‘27603327’ ‘27513442’ ‘27273834’ ‘27071789’ ‘26940141’ ‘26932552’ ‘26895254’ ‘26869847’ ‘26788924’ ‘26581735’ ‘26548910’ ‘26317636’ ‘26121678’ ‘26094200’ ‘25997760’ ‘25631363’ ‘25526824’ ‘25446893’ ‘25153535’ ‘25092245’ ‘25086828’ ‘24946432’ ‘24637261’ ‘24588761’ ‘24508579’ ‘24486356’ ‘24462936’ ‘24239932’ ‘24239931’ ‘24231551’ ‘24216134’ ‘23955310’ ‘23856187’ ‘23686025’ ‘23589638’ ‘23575742’ ‘23469841’ ‘23055480’ ‘22981649’ ‘22406388’ ‘22373652’ ‘22141469’ ‘21960250’ ‘21881219’ ‘21802859’ ‘21714746’ ‘21618004’ ‘21150165’ ‘20435805’ ‘20173685’ ‘19840865’ ‘19546570’ ‘19309413’ ‘15288368’ ‘12359512’ ‘9401603’ ‘9213136’ ‘7630585’</p> <p><a href="#_ftnref4">[4]</a> Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></p> <p><a href="#_ftnref5">[5]</a> Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media. Gephi is available via <a href="https://gephi.org/">https://gephi.org/</a></p> <p><a href="#_ftnref6">[6]</a> Traag, V. A., Waltman, L., & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports, 9(1), 5233. <a href="https://doi.org/10.1038/s41598-019-41695-z">https://doi.org/10.1038/s41598-019-41695-z</a></p>
A dataset from a survey investigating disciplinary differences in data citation
<p><strong>GENERAL INFORMATION</strong></p> <p><em>Title of Dataset: </em> A dataset from a survey investigating disciplinary differences in data citation</p> <p><em>Date of data collection:</em><strong> </strong>January to March 2022</p> <p><em>Collection instrument: </em>SurveyMonkey</p> <p><em>Funding:</em> Alfred P. Sloan Foundation</p> <p><br> <strong>SHARING/ACCESS INFORMATION</strong></p> <p><em>Licenses/restrictions placed on the data: </em>These data are available under a <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0 license</a> </p> <p><em>Links to publications that cite or use the data: </em></p> <p>Gregory, K., Ninkov, A., Ripp, C., Peters, I., & Haustein, S. (2022). Surveying practices of data citation and reuse across disciplines. Proceedings of the 26th International Conference on Science and Technology Indicators. <em>International Conference on Science and Technology Indicators</em>, Granada, Spain. https://doi.org/10.5281/ZENODO.6951437</p> <p>Gregory, K., Ninkov, A., Ripp, C., Roblin, E., Peters, I., & Haustein, S. (2023). <em>Tracing data:<br> A survey investigating disciplinary differences in data citation.</em> Zenodo. https://doi.org/10.5281/zenodo.7555266</p> <p><br> <strong>DATA & FILE OVERVIEW</strong></p> <p><em>File List</em></p> <ul> <li>Filename: MDCDatacitationReuse2021Codebookv2.pdf<br> <em>Codebook</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.csv<br> <em>Dataset format in csv</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.sav<br> <em>Dataset format in SPSS</em></li> <li>Filename: MDCDataCitationReuseSurvey2021QNR.pdf<br> <em>Questionnaire</em></li> </ul> <p><em>Additional related data collected that was not included in the current data package: </em>Open ended questions asked to respondents</p> <p><br> <strong>METHODOLOGICAL INFORMATION</strong></p> <p><em>Description of methods used for collection/generation of data: </em></p> <p>The development of the questionnaire (Gregory et al., 2022) was centered around the creation of two main branches of questions for the primary groups of interest in our study: researchers that reuse data (33 questions in total) and researchers that do not reuse data (16 questions in total). The population of interest for this survey consists of researchers from all disciplines and countries, sampled from the corresponding authors of papers indexed in the Web of Science (WoS) between 2016 and 2020. </p> <p>Received 3,632 responses, 2,509 of which were completed, representing a completion rate of 68.6%. Incomplete responses were excluded from the dataset. The final total contains 2,492 complete responses and an uncorrected response rate of 1.57%. Controlling for invalid emails, bounced emails and opt-outs (n=5,201) produced a response rate of 1.62%, similar to surveys using comparable recruitment methods (Gregory et al., 2020).</p> <p><em>Methods for processing the data: </em></p> <p>Results were downloaded from SurveyMonkey in CSV format and were prepared for analysis using Excel and SPSS by recoding ordinal and multiple choice questions and by removing missing values.</p> <p><em>Instrument- or software-specific information needed to interpret the data: </em></p> <p>The dataset is provided in SPSS format, which requires IBM SPSS Statistics. The dataset is also available in a coded format in CSV. The Codebook is required to interpret to values.</p> <p><br> <strong>DATA-SPECIFIC INFORMATION FOR: MDCDataCitationReuse2021surveydata</strong></p> <p><em>Number of variables:</em> 95</p> <p><em>Number of cases/rows:</em> 2,492</p> <p><em>Missing data codes:</em> 999 Not asked</p> <p>Refer to MDCDatacitationReuse2021Codebook.pdf for detailed variable information.</p>
A2.2a Digital repositories data citation practices. Supplementary material
<p>Data to complement the quantitative analysis of data citation practices in digital repositories based on metadata records from the re3data.org repositories registry.</p> <p>Data was retrieved using re3data.org API on 23-02-2023 and 06-03-2023 and processed using the OpenRefine software.</p> <p>Part of "A FAIR-enabling citation model for Cultural Heritage Objects" project activities.</p>
Data Citation Corpus Data File
<p>Data file for the fourth release of the Data Citation Corpus, produced by DataCite and Make Data Count as part of an ongoing grant project funded by the Wellcome Trust. <a href="https://makedatacount.org/data-citation/">Read more about the project</a>.</p> <p>The data file includes 10,697,745 data citation records (of which 9,682,257 represent unique dataset-publication pairs) in JSON and CSV formats. The JSON file is the version of record.</p> <p>Data is provided in batches of approximately 1 million records each. The publication date and batch number are included in the file name, ex: 2025-08-15-data-citation-corpus-01-v4.1.json.</p> <p>The data citations in the file originate from the following sources:</p> <ul> <li>DataCite Event Data</li> <li>Chan Zuckerberg Initiative (CZI) Science Knowledge Graph</li> <li>Aligning Science Across Parkinson’s (ASAP)</li> <li>Europe PMC</li> </ul> <p>Each data citation record is comprised of:</p> <ul> <li> <p>A pair of identifiers: An identifier for the dataset (a DOI or an accession number) and the DOI of the publication (journal article or preprint) in which the dataset is cited </p> </li> <li> <p>Metadata for the cited dataset and for the citing publication </p> </li> </ul> <p>The data file includes the following fields:</p> <div> <table> <tbody> <tr> <td> <p>Field</p> </td> <td> <p>Description</p> </td> <td> <p>Required?</p> </td> </tr> <tr> <td> <p>id</p> </td> <td> <p>Internal identifier for the citation</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>created</p> </td> <td> <p>Date of item's incorporation into the corpus</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>updated</p> </td> <td> <p>Date of item's most recent update in corpus</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>repository</p> </td> <td> <p>Repository where cited data is stored</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>publisher</p> </td> <td> <p>Publisher for the article citing the data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>journal</p> </td> <td> <p>Journal for the article citing the data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>title</p> </td> <td> <p>Title of cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>publication</p> </td> <td> <p>DOI of article where data is cited</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>dataset</p> </td> <td> <p>DOI or accession number of cited data</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>publishedDate</p> </td> <td> <p>Date when citing article was published</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>source</p> </td> <td> <p>Source where citation was harvested</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>subjects</p> </td> <td> <p>Subject information for cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>affiliations</p> </td> <td> <p>Affiliation information for creator of cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>funders</p> </td> <td> <p>Funding information for cited data</p> </td> <td> <p>No</p> </td> </tr> </tbody> </table> </div> <p><strong> </strong></p> <p>Additional documentation about the citations and metadata in the file is available on the <a href="https://makedatacount.org/find-a-tool/data-citation-corpus-documentation/">Make Data Count website</a>. </p> <p><strong>Notes on v4.1:</strong></p> <p>Version 4.1 of the Data Citation Corpus is a minor update to v4.0 that corrects (1) an error that occurred when a portion of DOI-DOI citations originating from Europe PMC were attributed to the wrong repository, and (2) a small number of DOI formatting errors in the "publication" field.</p> <p><strong>Notes on v4.0:</strong></p> <p>The fourth release of the Data Citation Corpus data file adds new citations from the following sources:</p> <ul> <li> <p dir="ltr">5.2 million data citations from <a href="https://europepmc.org/">Europe PMC</a> identified as "eupmc" in the source field. Ingest of these citations was performed 9 July 2025.</p> </li> <li> <p dir="ltr">139,647 data citations from DataCite Event Data for the period 1 January 2025 through 30 June 2025.</p> </li> </ul> <p dir="ltr">This release also includes the following new metadata enhancements:</p> <ul> <li> <p dir="ltr">Affiliation information for cited data from the Gene Expression Omnibus (GEO) repository, reonciled to Research Organization Registry (ROR) IDs where possible.</p> </li> <li> <p dir="ltr">Reconciliation of organization and funders names with the Research Organization Registry (ROR) for new citations from Event Data.</p> </li> <li> <p dir="ltr">Application of Field of Science subject terms to citation records originating from Europe PMC, based on disciplinary area of data repository.</p> </li> </ul> <p>Additional details about the above changes, including scripts used to perform the above tasks, are available in <a href="https://github.com/Make-Data-Count-Community/corpus-data-file" target="_blank" rel="noopener">GitHub</a>. </p> <p>Additional enhancements to the corpus are ongoing and will be addressed in the course of subsequent releases. Users are invited to submit feedback via <a href="https://github.com/Make-Data-Count-Community/data-citation-corpus-feedback">GitHub</a>. For general questions, email <a href="mailto:info@makedatacount.org">info@makedatacount.org</a>.</p>
Citations to software and data in Zenodo via open sources
<p>In January 2019, the Asclepias Broker harvested citation links to Zenodo objects from three discovery systems: the NASA Astrophysics Datasystem (ADS), Crossref Event Data and Europe PMC. Each row of our dataset represents one unique link between a citing publication and a Zenodo DOI. Both endpoints are described by basic metadata. The second dataset contains usage metrics for every cited Zenodo DOI of our data sample. </p> <p> </p>
Source Data for Manuscript: Identifying genomic data use with the Data Citation Explorer
<p>This page contains the source data for the manuscript describing the Data Citation Explorer, currently in review for publication. The preprint version can be found on this page.</p> <p>Files:</p> <p><strong>DCE_manual_eval_sample.xlsx:</strong></p> <p>This file was used to manually evaluate hits generated by the Data Citation Explorer. There are two separate sheets: one with publications returned by searches in PubMed and PubMed Central and another with publications returned by searches in Dimensions. Column descriptions can be found in the file itself. Each row in each evaluation sheet refers to a pair between a JAMO record and a linked publication.</p> <p><strong>DCE_citation_report.csv</strong></p> <p>Contains JAMO record IDs and PubMed IDs from the initial 2020 DCE trial run. There are 238,994 unique JAMO IDs and 30,641 unique PubMed IDs. 78,104 JAMO records are linked with publications.</p> <p>Columns:</p> <ul> <li>jamo_id - unique JAMO record ID</li> <li>sample_group - Sample strata from which manually evaluated records were pulled</li> <li>citation_count - Number of citations associated with each record</li> <li>citations - comma-delimited PubMed IDs for linked publications</li> <li>sampled - True/False, denoting which records were included in the initial evaluation sample</li> <li>notes - descriptions for why certain sampled records were excluded from manual evaluation</li> <li>unprocessed - True/False. These 7,890 records contained anomalous fields that caused them to be rejected for processing. They are represented as zero-length files in the archive.</li> </ul> <p><strong>DCE_source_files.zip:</strong></p> <p>This folder contains 3 files for each JAMO record in DCE_citation_report.tsv. For each JAMO record listed in the citation report, three files are provided:</p> <ol> <li>JAMO_ID_source.yaml - The fields extracted from the JAMO record that were relevant to the citation search, including any previously known PMIDs (manually curated).</li> <li>JAMO_ID_expand.yaml - The source record augmented with additional metadata discovered in other resources, including the citations that were discovered based on querying PubMed Central for the values in those metadata fields.</li> <li>JAMO_ID_audit.json - The audit path as a directed acyclic graph, in JSON.</li> </ol>
Dataset Citation and Re-use Data
<p>This dataset includes processed citation data for datasets recorded in OpenAlex as of May 2022. It identifies self-citations to these datasets at the individual, institutional, and country level, and includes domain classifications of the citing works using the Science-Metrix classifications.</p>
Data and code for EDI overview paper, data collection characteristics, FAIR evaluation, downloads, and citations
The Environmental Data Initiative (EDI) is a trustworthy, stable data repository and data management support organization for the environmental scientist. EDI provides tools and support that allow the environmental researcher to easily integrate data publishing into the research workflow. Almost ten years since going into production, these data and code were used to provide a general description of EDI’s collection of data and its data management philosophy and placement in the repository landscape. They show how comprehensive metadata and the repository infrastructure lead to highly findable, accessible, interoperable, and reusable (FAIR) data by evaluating compliance with specific community proposed FAIR criteria. Finally, they provide measures and patterns of data (re)use, assuring that EDI is fulfilling its stated premise.
Data for "Open Access impact on citations: a case study"
<p>This dataset is a list of 347 papers published in 2010 and retrieved from the Web of Science, Scopus and Google Scholar. For each paper, the number of citations and the citation date(s) have been collected. If the full-text is available online, the date of "liberation" and the URL of the file have been retrieved as well. The objective was to assess the impact of Open access on citation rate and more particularly the impact before and after full-text "liberation".</p> <p> </p>
PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)
<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>
Open access in Africa: scopus citation data
<p>The following citation dataset was retrieved from Scopus in June 24, 2017 (3am, Western Indonesian time).</p> <p>It consists of 3 sets of data based on our searches. Each search was saved both in 'csv' and 'bib':</p> <ol> <li>OA_Africa_inTitle.xxx: "Open Access" AND Africa IN TITLE</li> <li>OA_Africa_inTitle_inAbstract_inKeywords.xxx: "Open Access" AND Africa IN TITLE, IN ABSTRACT, IN KEYWORDS</li> <li>OAmovement_Africa_inTitle_inAbstract_inKeywords.xxx: "Open Access movement" AND Africa IN TITLE, IN ABSTRACT, IN KEYWORDS</li> </ol> <p>The access to Scopus was provided by The Central Library of Institut Teknologi Bandung (Indonesia)</p>
Citations to Astronomy Journals 1: The growth of interdisciplinarity - Data Supplement
<p>This repository contains the data used in the blog "Citations to Astronomy Journals 1: The growth of interdisciplinarity", Michael J. Kurtz and Edwin Henneken. Each file has a header with a description of the data contained in the file. Table 1 consists of the bibstems for the journals in the main sample. Table 2 contains the individual data for each journal in the format journal, indicator, value, year. Table 3 has these data for all refereed journals (including the journals in the main sample, and all the rest). All three data files are ASCII files with space-separated columns.</p> <p>The term "bibstem" is the journal abbreviation used within the Astrophysics Data System. The complete list of bibstems is provided here: http://adsabs.harvard.edu/abs_doc/journals2.html. The bibstem is used in the ADS bibliographic identifier ("bibcode") with the convention that ampersands are replaced by plus signs (see: http://adsabs.github.io/help/actions/bibcode).</p>
Data and code for: Community structure in co-inventor networks affects time to first citation for patents
<p>This package provides the datasets and programming code needed to reproduce the results reported in the article "Community structure in co-inventor networks affects time to first citation for patents".</p> <p>v2: Added data and code pertaining to randomized-community-association test and updated README file.</p>
Data set of the article: Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus
<p>Data of investigation published in the article "Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus".</p> <p>Abstract of the article:</p> <p>Search engine optimization (SEO) constitutes the set of methods designed to increase the visibility of, and the number of visits to, a web page by means of its ranking on the search engine results pages. Recently, SEO has also been applied to academic databases and search engines, in a trend that is in constant growth. This new approach, known as academic SEO (ASEO), has generated a field of study with considerable future growth potential due to the impact of open science. The study reported here forms part of this new field of analysis. The ranking of results is a key aspect in any information system since it determines the way in which these results are presented to the user. The aim of this study is to analyse and compare the relevance ranking algorithms employed by various academic platforms to identify the importance of citations received in their algorithms. Specifically, we analyse two search engines and two bibliographic databases: Google Scholar and Microsoft Academic, on the one hand, and Web of Science and Scopus, on the other. A reverse engineering methodology is employed based on the statistical analysis of Spearman’s correlation coefficients. The results indicate that the ranking algorithms used by Google Scholar and Microsoft are the two that are most heavily influenced by citations received. Indeed, citation counts are clearly the main SEO factor in these academic search engines. An unexpected finding is that, at certain points in time, WoS used citations received as a key ranking factor, despite the fact that WoS support documents claim this factor does not intervene.</p>
Zenodo data and software citation links captured by the Asclepias Broker
<p>The dataset was retrieved from the Asclepias Broker early January 2019 after having performed a full harvesting and deduplication cycle from a clean database with zero citation links.</p> <p>The dataset contains citation links from three discovery systems: the NASA Astrophysics Datasystem (ADS), Crossref Event Data and Europe PMC. Only citation links with a target DOI in the DOI prefix 10.5281 (Zenodo’s DOI prefix) were kept.</p>
Citation count error data for "Data inaccuracy quantification and uncertainty propagation for bibliometric indicators"
<p>This is the original collected data on citation count errors resulting from citation matching errors in Web of Science data for the publication "Data inaccuracy quantification and uncertainty propagation for<br>bibliometric indicators". The first column, <code>CITCOUNT_ALL</code>, gives the total (corrected) citation count for a publication, which is the citation count according to WoS plus the additionally manually identified citations (missed by WoS's algorithm). The second column, <code>CITCOUNT_WOS</code>, is the WoS citation count. The numeric difference between the two column values in one row is the number of additionally manually identified citations.</p>
Diversity in citations to a single study: Supplementary data set for citation context network analysis
<p><strong>Introduction</strong></p> <p>This document describes the data set used for all analyses in 'Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited' accepted for publication in <em>Quantitative Science Studies</em> [1].</p> <p><strong>Data Collection</strong></p> <p>The data collection procedure has been fully described [1]. Concisely, the data set contains bibliometric data collected from Web of Science Core Collection via the University of Edinburgh’s Library subscription concerning all papers that cited a cohort study, Paul <em>et al.</em> [2], in the period <1985. This includes a full list of citing papers, and the citations between these papers. Additionally, it includes textual passages (citation contexts) from 343 citing papers, which were manually recovered from the full-text documents accessible via the University of Edinburgh’s Library subscription. These data have been cleaned, converted into network readable datasets, and are coded into particular classifications reflecting content, which are described fully in the supplied code book and within the manuscript [1]. </p> <p><strong>Data description</strong></p> <p>All relevant data can be found in the attached file 'Supplementary_material_Leng_QSS_2021.xlsx', which contains the following five workbooks:</p> <ul> <li><strong>“Overview”</strong> includes a list of the content of the workbooks.</li> <li><strong>“Code Book”</strong> contains the coding rules and definitions used for the classification of findings and paper titles.</li> <li><strong>“Node attribute list”</strong> includes a workbook containing all node attributes for the citation network, which includes Paul et al. [2] and its citing papers as of 1984. Highlighted in yellow at the bottom of this workbook is two papers that were discarded due to duplication - remove these if analysing this dataset in a network analysis. The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Label</em>, the formal citation of the paper to which data within this row corresponds. Citation is in the following format: last name of first author, year of publication, journal of publication, volume number, start page, and DOI (if available). </li> <li><em>Title</em>, the paper title for the paper in question.</li> <li><em>Publication_year</em>, the year of publication.</li> <li><em>Document_type, </em>the document type (e.g. review, article)</li> <li><em>WoS_ID</em>, the paper’s unique Web of Science accession number.</li> <li><em>Citation_context</em>, a column specifying whether citation context data is available from that paper</li> <li><em>Explanans</em>, the title explanans terms for that paper;</li> <li><em>Explanandum</em>, the explanandum terms for that paper.</li> <li><em>Combined_Title_Classification</em>, the combined terms used for fig 2 of the published manuscript.</li> <li><em>Serum_cholesterol_(SC)</em>, a column identifying papers that cited the serum cholesterol findings.</li> <li><em>Blood_Pressure_(BP), </em>a column identifying papers that cited the blood pressure findings.</li> <li><em>Coffee_(C),</em> a column identifying papers that cited the coffee findings.</li> <li><em>Diet_(D), </em>a column identifying papers that cited the dietary findings.</li> <li><em>Smoking_(S), </em>a column identifying papers that cited the smoking findings.</li> <li><em>Alcohol_(A), </em>a column identifying papers that cited the alcohol findings.</li> <li><em>Physical_Activity_(PA),</em> a column identifying papers that cited the physical activity findings.</li> <li><em>Body_Fatness (BF), </em>a column identifying papers that cited the body fatness findings.</li> <li><em>Indegree,</em> the number of within network citations to that paper, calculated for the network shown in Fig 4 of the manuscript.</li> <li><em>Outdegree</em>, the number of within network references of that paper as calculated for the network in Fig 4.</li> <li><em>Main_component</em>, a column specifying whether a node is contained in the largest weakly connect component as shown in Fig 4 of the manuscript.</li> <li><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Fig 5).</li> </ol> <ul> <li><strong>“Edge list”</strong> includes a workbook including the edges for the network. The columns refer to:</li> </ul> <ol> <li><em>Source</em>, contains the node identifier of the citing paper.</li> <li><em>Target,</em> contains the node identifier of the cited paper.</li> </ol> <ul> <li><strong>“Citation context classification</strong>” includes a workbook containing the WoS accession number for the paper analysed, and any finding category discussed in that paper established via context analysis (see the code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Finding_Class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <ul> <li><strong> “Citation context data”</strong> includes a workbook containing the WoS accession number for papers in which citation context data was available, the citation context passages, the reference number or format of Paul et al. within the citing paper, and the finding categories discussed in those contexts (see code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Citation_context</em>, the passage copied from the full text of the citing paper containing discussion of the findings of Paul et al.</li> <li><em>Reference_in_citing_article</em>, the reference number or format of Paul et al. within the citing paper.</li> <li><em>Finding_class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <p><strong>Software recommended for analysis</strong></p> <p>For the analyses performed within the manuscript, Gephi version 0.9.2 was used [3], and both the edge and node lists are in a format that is easily read into this software. The Sci2 tool was used to parse data initially [4].</p> <p><strong>Notes</strong></p> <ol> <li>Leng, R. I. (Forthcoming). Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited. Quantitative Science Studies.</li> <li>Paul, O., Lepper, M. H., Phelan, W. H., Dupertuis, G. W., Macmillan, A., McKean, H., <em>et al.</em> (1963). A longitudinal study of coronary heart disease. <em>Circulation, </em><strong>28</strong>, 20-31. <a href="https://doi.org/10.1161/01.cir.28.1.20">https://doi.org/10.1161/01.cir.28.1.20</a>.</li> <li>Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media.</li> <li>Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.