Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

20

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

20 results for “citation network”

Learn how ShareScore rates datasets ↗
zenodo48/100

Citation network data sets for 'Oxytocin – a social peptide? Deconstructing the evidence'

<p><strong>Introduction</strong></p> <p>This note describes the data sets used for all analyses contained in the manuscript &#39;Oxytocin - a social peptide?&rsquo;<a href="#_ftn1">[1]</a>&nbsp;</p> <p><strong>Data Collection</strong></p> <p>The datasets described here were originally retrieved from Web of Science (WoS) Core Collection via the University of Edinburgh&rsquo;s library subscription&nbsp;<a href="#_ftn2">[2]</a>. The aim of the original study for which these data were gathered was to survey peer-reviewed primary studies on oxytocin and social behaviour. To capture relevant papers, we used the following query:</p> <p><em>TI = (&ldquo;oxytocin&rdquo; OR &ldquo;pitocin&rdquo; OR &ldquo;syntocinon&rdquo;)&nbsp;AND&nbsp;TS&nbsp;=&nbsp;(&ldquo;social*&rdquo; OR &ldquo;pro$social&rdquo; OR &ldquo;anti$social&rdquo;)</em></p> <p>The final search was performed on the 13 September 2021. This returned a total of 2,747 records, of which 2,049 were classified by WoS as &lsquo;articles&rsquo;. Given our interest in primary studies <em>only</em> &ndash; articles reporting original data &ndash; we excluded all other document types. We further excluded all articles sub-classified as &lsquo;book chapters&rsquo; or as &lsquo;proceeding papers&rsquo; in order to limit our analysis to primary studies published in peer-reviewed academic journals. This reduced the set to 1,977 articles. All of these were published in the English language, and no further language refinements were unnecessary.</p> <p>All available metadata on these 1,977 articles was exported as plain text &lsquo;flat&rsquo; format files in four batches, which we later merged together via Notepad++. Upon manually examination, we discovered examples of papers classified as &lsquo;articles&rsquo; by WoS that were, in fact, reviews. To further filter our results, we searched all available PMIDs in PubMed (1,903 had associated PMIDs - ~96% of set). We then filtered results to identify all records classified as &lsquo;review&rsquo;, &lsquo;systematic review&rsquo;, or &lsquo;meta-analysis&rsquo;, identifying 75 records&nbsp;<a href="#_ftn3">[3]</a> (thus, ~4% of records classified by WoS were classified as reviews in PubMed). After examining a sample and agreeing with the PubMed classification, these were removed these from our dataset - leaving a total of 1,902 articles.</p> <p>From these data, we constructed two datasets via parsing out relevant reference data via the Sci2 Tool&nbsp;<a href="#_ftn4">[4]</a>. First, we constructed a &lsquo;node-attribute-list&rsquo; by first linking unique reference strings (&lsquo;Cite Me As&rsquo; column in WoS data files) to unique identifiers, we then parsed into this dataset information on the identify of a paper, including the title of the article, all authors, journal publication, year of publication, total citations as recorded from WoS, and WoS accession number. Second, we constructed an &lsquo;edge-list&rsquo; that records the citations from a <em>citing paper</em> in the &lsquo;Source&rsquo; column and identifies the <em>cited paper</em> in the &lsquo;Target&rsquo; column, using the unique identifies as described previously to link these data to the node-attribute-list.</p> <p>We then constructed a network in which papers are nodes, and citation links between nodes are directed edges between nodes. We used Gephi Version 0.9.2&nbsp;<a href="#_ftn5">[5]</a> to manually clean these data by merging duplicate references that are caused by different reference formats or by referencing errors. To do this, we needed to retain both all retrieved records (1,902) as well as including <em>all</em> of their references to papers whether these were included in our original search or not. In total, this produced a network of 46,633 nodes (unique reference strings) and 112,520 edges (citation links). Thus, the average reference list size of these articles is ~59 references. The mean indegree (within network citations) is 2.4 (median is 1) for the entire network reflecting a great diversity in referencing choices among our 1,902 articles.</p> <p>After merging duplicates, we then restricted the network to include <em>only</em> articles fully retrieved (1,902), and retrained <em>only</em> those that were connected together by citations links in a large interconnected network (i.e. the largest component). In total, 1,892 (99.5%) of our initial set were connected together via citation links, meaning a total of ten papers were removed from the following analysis &ndash; and these were neither connected to the largest component, nor did they form connections with one another (i.e. these were &lsquo;isolates&rsquo;).</p> <p>This left us with a network of 1,892 nodes connected together by 26,019 edges. <strong><em>It is this network that is described by the &lsquo;node-attribute-list&rsquo; and &lsquo;edge-list&rsquo; provided here</em></strong>. This network has a mean in-degree of 13.76 (median in-degree of 4). By restricting our analysis in this way, we lose 44,741 unique references (96%) and 86,501 citations (77%) from the full network, but retain a set of articles tightly knitted together, all of which have been fully retrieved due to possessing certain terms related to oxytocin AND social behaviour in their title, abstract, or associated keywords.</p> <p>Before moving on, we calculated indegree for all nodes in this network &ndash; this counts the number of citations to a given paper from other papers within this network &ndash; and have included this in the <em>node-attribute-list</em>. We further clustered this network via modularity maximisation via the Leiden algorithm&nbsp;<a href="#_ftn6">[6]</a>. We set the algorithm to resolution 1, and allowed the algorithm to run over 100 iterations and 100 restarts. This gave <em>Q</em>=0.43 and identified seven clusters, which we describe in detail within the body of the paper. We have included cluster membership as an attribute in the node-attribute-list.</p> <p>For additional analysis, we also analysed the full reference list data to examine the most commonly cited references between 2016 and 2021 - the results of this are described in OTSOC_Cited_2016-2021.csv. This takes the reference lists of all retrieved papers within the network and examines their full reference lists (including references to other papers not contained within the network). These data were cleaned by matching DOIs and manual cleansing.&nbsp;</p> <p><strong>Data description</strong></p> <p>We include here two network datasets: (i) &lsquo;OTSOC-node-attribute-list.csv&rsquo; consists of the attributes of 1,892 primary articles retrieved from WoS that include terms indicating a focus on oxytocin and social behaviour; (ii) &lsquo;OTSOC-edge-list.csv&rsquo; records the citations between these papers. Together, these can be imported into a range of different software for network analysis; however, we have formatted these for ease of upload into Gephi 0.9.2. Finally, we include (iii) &#39;OTSOC_Cited_2016-2021&#39; that lists all papers cited by &gt;10 papers in the OTSOC network following any analysis of the bibliographies of retrieved papers. Below, we detail their contents:</p> <p><strong>1. &lsquo;OTSOC-node-attribute-list.csv&rsquo;</strong> is a comma-separate values file that contains all node attributes for the citation network (n=1,892) analysed in the paper. The columns refer to:</p> <p><em>Id</em>, the unique identifier</p> <p><em>Label</em>, the reference string of the paper to which the attributes in this row correspond. This is taken from the &lsquo;Cite Me As&rsquo; column from the original WoS download. The reference string is in the following format: last name of first author, publication year, journal, volume, start page, and DOI (if available).&nbsp;</p> <p><em>Wos_id</em>, unique Web of Science (WoS) accession number. These can be used to query WoS to find further data on all papers via the &lsquo;UT= &rsquo; field tag.</p> <p><em>Title</em>, paper title.</p> <p><em>Authors</em>, all named authors.</p> <p><em>Journal, </em>journal of publication.</p> <p><em>Pub_year</em>, year of publication.</p> <p><em>Wos_citations</em>, total number of citations recorded by WoS Core Collection to a given paper as of 13 September 2021</p> <p><em>Indegree</em>, the number of within network citations to a given paper, calculated for the network shown in Figure 1 of the manuscript.</p> <p><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Figure 1). This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.43|7 clusters)</p> <p><strong>2. &lsquo;OTSOC-edge -list.csv&rsquo;</strong> is a comma-separated values file that contains all citation links between the 1,892 articles (n=26,019). The columns refer to:</p> <p><em>Source</em>, the unique identifier of the citing paper.</p> <p><em>Target, </em>the unique identifier of the cited paper.</p> <p><em>Type, </em>edges are &lsquo;Directed&rsquo;, and this column tells Gephi to regard all edges as such.</p> <p><em>Syr_date, </em>this contains the date of publication of the citing paper.</p> <p><em>Tyr_date, </em>this contains the date of publication of the cited paper.</p> <p><strong>3. &#39;OTSOC_Cited_2016-2021.csv&#39;</strong>&nbsp;is a comma-separated values file that contain citations to all cited references that were cited by at least 10 of the&nbsp;retrieved papers within the OTSOC network&nbsp;published from 2016 onwards. The columns refer to:&nbsp;</p> <p><em>Reference,&nbsp;</em>the cited reference string extracted from the&nbsp;bibliographies of retrieved papers.</p> <p><em>Publication year,&nbsp;</em>the publication year of the cited reference.</p> <p><em>DOI</em>, the DOI of the cited reference.&nbsp;</p> <p><em>indegree_2016,&nbsp;</em>the total number of citations to a cited reference from papers published in 2016 and contained within the OTSOC network.&nbsp;</p> <p><em>indegree_2017,&nbsp;</em>the total number of citations to a cited reference from papers published in 2017 and contained within the OTSOC network.&nbsp;</p> <p><em>indegree_2018,&nbsp;</em>the total number of citations to a cited reference from papers published in 2018 and contained within the OTSOC network.&nbsp;</p> <p><em>indegree_2019,&nbsp;</em>the total number of citations to a cited reference from papers published in 2019&nbsp;and contained within the OTSOC network.&nbsp;</p> <p><em>indegree_2020,&nbsp;</em>the total number of citations to a cited reference from papers published in 2020&nbsp;and contained within the OTSOC network.&nbsp;</p> <p><em>indegree_2021,&nbsp;</em>the total number of citations to a cited reference from papers published in 2021&nbsp;and contained within the OTSOC network.&nbsp;</p> <p><em>total indegree 2016-21</em>, the total number of citation to a cited reference from papers published between 2016-2021 and contained within the OTSOC network.&nbsp;</p> <p><strong>Software recommended for analysis</strong></p> <p>Gephi version 0.9.2 was used for the visualisations within the manuscript, and both files can be read and into Gephi without modification.</p> <p><strong>Notes</strong></p> <p><a href="#_ftnref1">[1]</a> Leng, G., Leng, R. I., Ludwig, M. (Submitted). Oxytocin &ndash; a social peptide? Deconstructing the evidence.</p> <p><a href="#_ftnref2">[2]</a> Edinburgh University&rsquo;s subscription to Web of Science covers the following databases: (i) Science Citation Index Expanded, 1900-present; (ii) Social Sciences Citation Index, 1900-present; (iii) Arts &amp; Humanities Citation Index, 1975-present; (iv) Conference Proceedings Citation Index- Science, 1990-present; (v) Conference Proceedings Citation Index- Social Science &amp; Humanities, 1990-present; (vi) Book Citation Index&ndash; Science, 2005-present; (vii) Book Citation Index&ndash; Social Sciences &amp; Humanities, 2005-present; (viii) Emerging Sources Citation Index, 2015-present.</p> <p><a href="#_ftnref3">[3]</a> For those interested, the following PMIDs were identified as &lsquo;articles&rsquo; by WoS, but as &lsquo;reviews&rsquo; by PubMed: &lsquo;34502097&rsquo; &lsquo;33400920&rsquo; &lsquo;32060678&rsquo; &lsquo;31925983&rsquo; &lsquo;31734142&rsquo; &lsquo;30496762&rsquo; &lsquo;30253045&rsquo; &lsquo;29660735&rsquo; &lsquo;29518698&rsquo; &lsquo;29065361&rsquo; &lsquo;29048602&rsquo; &lsquo;28867943&rsquo; &lsquo;28586471&rsquo; &lsquo;28301323&rsquo; &lsquo;27974283&rsquo; &lsquo;27626613&rsquo; &lsquo;27603523&rsquo; &lsquo;27603327&rsquo; &lsquo;27513442&rsquo; &lsquo;27273834&rsquo; &lsquo;27071789&rsquo; &lsquo;26940141&rsquo; &lsquo;26932552&rsquo; &lsquo;26895254&rsquo; &lsquo;26869847&rsquo; &lsquo;26788924&rsquo; &lsquo;26581735&rsquo; &lsquo;26548910&rsquo; &lsquo;26317636&rsquo; &lsquo;26121678&rsquo; &lsquo;26094200&rsquo; &lsquo;25997760&rsquo; &lsquo;25631363&rsquo; &lsquo;25526824&rsquo; &lsquo;25446893&rsquo; &lsquo;25153535&rsquo; &lsquo;25092245&rsquo; &lsquo;25086828&rsquo; &lsquo;24946432&rsquo; &lsquo;24637261&rsquo; &lsquo;24588761&rsquo; &lsquo;24508579&rsquo; &lsquo;24486356&rsquo; &lsquo;24462936&rsquo; &lsquo;24239932&rsquo; &lsquo;24239931&rsquo; &lsquo;24231551&rsquo; &lsquo;24216134&rsquo; &lsquo;23955310&rsquo; &lsquo;23856187&rsquo; &lsquo;23686025&rsquo; &lsquo;23589638&rsquo; &lsquo;23575742&rsquo; &lsquo;23469841&rsquo; &lsquo;23055480&rsquo; &lsquo;22981649&rsquo; &lsquo;22406388&rsquo; &lsquo;22373652&rsquo; &lsquo;22141469&rsquo; &lsquo;21960250&rsquo; &lsquo;21881219&rsquo; &lsquo;21802859&rsquo; &lsquo;21714746&rsquo; &lsquo;21618004&rsquo; &lsquo;21150165&rsquo; &lsquo;20435805&rsquo; &lsquo;20173685&rsquo; &lsquo;19840865&rsquo; &lsquo;19546570&rsquo; &lsquo;19309413&rsquo; &lsquo;15288368&rsquo; &lsquo;12359512&rsquo; &lsquo;9401603&rsquo; &lsquo;9213136&rsquo; &lsquo;7630585&rsquo;</p> <p><a href="#_ftnref4">[4]</a> Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></p> <p><a href="#_ftnref5">[5]</a> Bastian, M., Heymann, S., &amp; Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media. Gephi is available via <a href="https://gephi.org/">https://gephi.org/</a></p> <p><a href="#_ftnref6">[6]</a> Traag, V. A., Waltman, L., &amp; van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports, 9(1), 5233. <a href="https://doi.org/10.1038/s41598-019-41695-z">https://doi.org/10.1038/s41598-019-41695-z</a></p>

opencc-by-4.0Oct 2021View details →
zenodo44/100

Datasets and results of the paper titled "Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification"

<p>These&nbsp;are&nbsp;the <strong>input&nbsp;datasets</strong> and the <strong>results of the analyses</strong>&nbsp;reported on&nbsp;the paper titled <strong>&quot;Are citation networks relevant to explain academic promotions? An empirical analysis of the Italian national scientific qualification&quot;</strong>.</p> <p><strong>Abstract:</strong>&nbsp;</p> <p>The aim of this paper is to study the role of citation network measures in the assessment of scientific maturity. Referring to the case of the Italian national scientific qualification (ASN), we investigate if there is a relationship between citation network indices and the results of the researchers&rsquo; evaluation procedures. In particular, we want to understand if network measures can enhance the prediction accuracy of the results of the evaluation procedures beyond basic performance indices. Moreover, we want to highlight which citation network indices prove to be more relevant in explaining the ASN results, and if quantitative indices used in the citation-based disciplines assessment can replace the citation network measures in non-citation-based disciplines. Data concerning Statistics and Computer Science disciplines are collected from different sources (ASN, Italian Ministry of University and Research, and Scopus) and processed in order to calculate the citation-based measures used in this study. Following, we apply classification models to estimate the effects of network variables. We find that network measures are strongly related to the results of the ASN and significantly improve the explanatory power of the models, especially for the research fields of Statistics. Additionally, citation networks in the specific sub-disciplines are far more relevant than those in the general disciplines. Finally, results show that the citation network measures are not a substitute of the citation-based bibliometric indices.</p> <p><strong>Code</strong></p> <p>The code to collect&nbsp;and process the data used in this paper is available on GitHub at <a href="https://github.com/DigitalDataLab/ASN16-18_CitationNetwork">https://github.com/DigitalDataLab/ASN16-18_CitationNetwork</a><strong>.</strong>&nbsp;</p> <p><strong>Dataset description</strong></p> <p>The files&nbsp;<strong>AdjacencyMatrix_01B1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_09H1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D1.csv</strong>,&nbsp;<strong>AdjacencyMatrix_13D2.csv</strong> and&nbsp;<strong>AdjacencyMatrix_13D3.csv</strong> are the&nbsp;citation matrices for Italian academics (i.e. ASN candidates and permanent positions in the Italian academic system) in the Recruitment Fields (RFs) 01/B1, 09/H1, 13/D1, 13/D2 and&nbsp;13/D3, respectively.</p> <p>The files&nbsp;<strong>AdjacencyMatrix_CS.csv</strong>&nbsp;and&nbsp;<strong>AdjacencyMatrix_ST.csv</strong> are the citation matrices for the Italian academics in the Computer Science disciplines (i.e. RFs 01/B1 and 09/H1) and the Statistical disciplines (i.e. RFs 13/D1,&nbsp;13/D2 and&nbsp;13/D3), respectively.</p> <p>The files&nbsp;<strong>CS_01B1_1.csv,&nbsp;CS_09H1_1.csv, ST_13D1_1.csv,&nbsp;ST_13D2_1.csv</strong> and&nbsp;<strong>ST_13D3_1.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Full Professor (FP) level.</p> <p>The files&nbsp;<strong>CS_01B1_2.csv,&nbsp;CS_09H1_2.csv, ST_13D1_2.csv,&nbsp;ST_13D2_2.csv</strong> and&nbsp;<strong>ST_13D3_2.csv</strong>&nbsp;contain the data used to build the&nbsp;logistic regression models presented in the paper for the Italian academics at the Associate Professor (AP) level.</p> <p>The file&nbsp;<strong>Codebook.pdf</strong>&nbsp;is the codebook of the previous ten files.</p> <p>The file <strong>Appendix.pdf</strong> contains the final results of the stepwise logistic regressions computed for each level (i.e. Full Professor and Associate Professor) and Recruitment Field in the Computer Science and Statistics disciplines.</p> <p>The file&nbsp;<strong>NormalityAssessment.pdf</strong>&nbsp;contains the&nbsp;normality assessment of citation network indices.&nbsp;</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

KGlove+Glove embedding for MAKG citation network

<p>Entity embedding by using KGlove+Glove on MAKG citation network (citation network from&nbsp;<a href="https://zenodo.org/record/4617285/files/08.PaperReferences.nt.bz2?download=1">https://zenodo.org/record/4617285/files/08.PaperReferences.nt.bz2?download=1</a>)</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Data and code for: Community structure in co-inventor networks affects time to first citation for patents

<p>This package provides the datasets and programming code&nbsp;needed to reproduce the results reported in the article &quot;Community structure in co-inventor networks affects time to first citation for patents&quot;.</p> <p>v2: Added data and code pertaining to randomized-community-association test and updated README file.</p>

opencc-by-4.0Feb 2019View details →
zenodo40/100

Diversity in citations to a single study: Supplementary data set for citation context network analysis

<p><strong>Introduction</strong></p> <p>This document describes the data set used for all analyses in &#39;Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited&#39; accepted for publication in&nbsp;<em>Quantitative&nbsp;Science Studies</em> [1].</p> <p><strong>Data Collection</strong></p> <p>The data collection procedure has been fully described [1]. Concisely, the data set contains bibliometric data collected from Web of Science Core Collection via the University of Edinburgh&rsquo;s Library subscription concerning all papers that cited a cohort study, Paul <em>et al.</em> [2], in the period &lt;1985. This includes a full list of citing papers, and the citations between these papers. Additionally, it includes textual passages (citation contexts) from 343 citing papers, which were manually recovered from the full-text documents accessible via the University of Edinburgh&rsquo;s Library subscription. These data have been cleaned, converted into network readable datasets, and are coded into particular classifications reflecting content, which are described fully in the supplied code book and within the manuscript [1].&nbsp;</p> <p><strong>Data description</strong></p> <p>All relevant data can be found in the attached file &#39;Supplementary_material_Leng_QSS_2021.xlsx&#39;, which contains the following five workbooks:</p> <ul> <li><strong>&ldquo;Overview&rdquo;</strong> includes a list of the content of the workbooks.</li> <li><strong>&ldquo;Code Book&rdquo;</strong> contains the coding rules and definitions used for the classification of findings and paper titles.</li> <li><strong>&ldquo;Node attribute list&rdquo;</strong> includes a workbook containing all node attributes for the citation network, which includes Paul et al. [2] and its citing papers as of 1984. Highlighted in yellow at the bottom of this workbook is two papers that were discarded due to duplication - remove these if analysing this dataset in a network analysis. The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Label</em>, the formal citation of the paper to which data within this row corresponds. Citation is in the following format: last name of first author, year of publication, journal of publication, volume number, start page, and DOI (if available). &nbsp;</li> <li><em>Title</em>, the paper title for the paper in question.</li> <li><em>Publication_year</em>, the year of publication.</li> <li><em>Document_type, </em>the document type (e.g. review, article)</li> <li><em>WoS_ID</em>, the paper&rsquo;s unique Web of Science accession number.</li> <li><em>Citation_context</em>, a column specifying whether citation context data is available from that paper</li> <li><em>Explanans</em>, the title explanans terms for that paper;</li> <li><em>Explanandum</em>, the explanandum terms for that paper.</li> <li><em>Combined_Title_Classification</em>, the combined terms used for fig 2 of the published manuscript.</li> <li><em>Serum_cholesterol_(SC)</em>, a column identifying papers that cited the serum cholesterol findings.</li> <li><em>Blood_Pressure_(BP), </em>a column identifying papers that cited the blood pressure findings.</li> <li><em>Coffee_(C),</em> a column identifying papers that cited the coffee findings.</li> <li><em>Diet_(D), </em>a column identifying papers that cited the dietary findings.</li> <li><em>Smoking_(S), </em>a column identifying papers that cited the smoking findings.</li> <li><em>Alcohol_(A), </em>a column identifying papers that cited the alcohol findings.</li> <li><em>Physical_Activity_(PA),</em> a column identifying papers that cited the physical activity findings.</li> <li><em>Body_Fatness (BF), </em>a column identifying papers that cited the body fatness findings.</li> <li><em>Indegree,</em> the number of within network citations to that paper, calculated for the network shown in Fig 4 of the manuscript.</li> <li><em>Outdegree</em>, the number of within network references of that paper as calculated for the network in Fig 4.</li> <li><em>Main_component</em>, a column specifying whether a node is contained in the largest weakly connect component as shown in Fig 4 of the manuscript.</li> <li><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Fig 5).</li> </ol> <ul> <li><strong>&ldquo;Edge list&rdquo;</strong> includes a workbook including the edges for the network. The columns refer to:</li> </ul> <ol> <li><em>Source</em>, contains the node identifier of the citing paper.</li> <li><em>Target,</em> contains the node identifier of the cited paper.</li> </ol> <ul> <li><strong>&ldquo;Citation context classification</strong>&rdquo; includes a workbook containing the WoS accession number for the paper analysed, and any finding category discussed in that paper established via context analysis (see the code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Finding_Class, </em>the findings discussed from Paul et al. within the body of the citing paper. &nbsp;</li> </ol> <ul> <li><strong>&nbsp;&ldquo;Citation context data&rdquo;</strong> includes a workbook containing the WoS accession number for papers in which citation context data was available, the citation context passages, the reference number or format of Paul et al. within the citing paper, and the finding categories discussed in those contexts (see code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Citation_context</em>, the passage copied from the full text of the citing paper containing discussion of the findings of Paul et al.</li> <li><em>Reference_in_citing_article</em>, the reference number or format of Paul et al. within the citing paper.</li> <li><em>Finding_class, </em>the findings discussed from Paul et al. within the body of the citing paper.&nbsp;</li> </ol> <p><strong>Software recommended for analysis</strong></p> <p>For the analyses performed within the manuscript, Gephi version 0.9.2 was used [3], and both the edge and node lists are in a format that is easily read into this software. The Sci2 tool was used to parse data initially [4].</p> <p><strong>Notes</strong></p> <ol> <li>Leng, R. I. (Forthcoming). Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited. Quantitative Science Studies.</li> <li>Paul, O., Lepper, M. H., Phelan, W. H., Dupertuis, G. W., Macmillan, A., McKean, H., <em>et al.</em> (1963). A longitudinal study of coronary heart disease. <em>Circulation, </em><strong>28</strong>, 20-31. <a href="https://doi.org/10.1161/01.cir.28.1.20">https://doi.org/10.1161/01.cir.28.1.20</a>.</li> <li>Bastian, M., Heymann, S., &amp; Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media.</li> <li>Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></li> </ol>

opencc-by-4.0Aug 2021View details →
zenodo40/100

DATABASE: Electric Vehicle, Battery and Smart Grid patent citation networks and main paths.

<p>This dataset comprises the original patent citation networks that were created to calculate&nbsp;the main citation paths for the technologies of Electric Vehicle, Battery and Smart Grid.</p> <p>For each technology (1- Electric Vehicle, 2- Battery, 3- Smart Grid), four outputs are provided:</p> <p>a- Patent extraction: USPTO patents filtered by IPC or CPC and found in the Triadic Patent Families database (OECD, 2021)&nbsp;&nbsp;</p> <p>b- Full nodes and links reconstructed by following patent citations through a snowball method&nbsp;(until no further patents found)</p> <p>c- Filtered nodes and links according to keywords</p> <p>d- Main path nodes and links (with citation weights).</p> <p>For a detailed explanation of the methodology please refer to the submitted paper:</p> <p><strong>Transitions as a coevolutionary process: the urban emergence of electric vehicle inventions</strong></p>

opencc-by-4.0Sep 2021View details →
zenodo36/100

Interactive Citation Networks for Texas Archaeology

<p>This dataset includes the JavaScript used to generate two interactive citation networks. IntNetFig1 includes the network filtered by the giant component and nodes sized by OutDegree, and IntNetFig2 is the same network filtered by a degree of two,&nbsp;nodes sized by Eigenvector Centrality and colored by modularity class.</p>

opencc-by-4.0Jul 2016View details →
zenodo36/100

Citation network dataset covering the work of RP Millar and its citing literature

<p>The following describes the citation network datasets that underpins the manuscript &ldquo;A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research&rdquo; [1].</p> <p><strong>Data collection</strong></p> <p>We retrieved data from the Web of Science Core Collection under the University of Edinburgh&rsquo;s subscription in January 2024. We sought to retrieve all indexed papers of Professor Robert P. Millar (RPM). We searched the following <em>AU = (Millar, R)</em>, and then retrieved records that corresponded to his WoS profile (n=428 records) and an additional 49 paper that were authored by Robert but had not been included in his WoS record &ndash; validating the records against a CV of his published works.</p> <p>We retrieved the full citation history as record by Web of Science to these papers from other indexed records. The 477 RPM papers had been cited 21,677 times by 11,138 documents by date of retrieval, and removing self-citations left 19,256 citations by 10,719 documents. We then retrieved all metadata from WoS concerning the 477 RPM papers and the 10,719 citation papers, resulting in a dataset covering 11,196 documents.</p> <p><strong>Citation network dataset</strong></p> <p>We constructed a citation network dataset by parsing data from each paper&rsquo;s full bibliography consisting of:</p> <p>i. &lsquo;Edge-list&rsquo; that records citation links from a citing to a cited document. This is constructed by assigning unique IDs to each retrieved paper and to every unique reference string contained in their bibliographies. The edge list is composed of a &lsquo;Source&rsquo; column that contains the ID of the <em>citing</em> document and a &lsquo;Target&rsquo; column containing the IDs of its citations, with one record per row. Given that we were only interested in citations between the WoS retrieved documents, we discarded any reference string that represented a document outwith our search.</p> <p>ii. &lsquo;Node-attribute list&rsquo; that contains the ID, with relevant metadata contained in adjacent columns to identify documents, including authors, title of publication, journal, year of publication. We also parsed into this dataset the WoS full citation count for each paper and the total number of references in the bibliographies of each paper.</p> <p>This results in a dataset containing 11,196 nodes and 115,834 edges between nodes. We removed a total of 67 papers for which metadata was incomplete and/or corrupted. We further focussed on the largest interconnected component, removing nodes with no connections (isolates) or smaller components that were detached from the main network. We excluded papers &lt;10 references to remove meeting abstracts and other minor journal items, and papers not published in English. This resulted in a final dataset containing 10,901 nodes and 113,742 edges, and it is this dataset that we share as it is the basis for the analyses within the paper.</p> <p><strong>Description of dataset variables</strong></p> <p><strong>&lsquo;RPM_Edgelist.csv&rsquo;</strong> is a comma-separate values file that consists of all 113,742 citations between the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Source</em>&rsquo;, the unique identifier for the <em>citing </em>document</li> <li>&lsquo;<em>Target&rsquo;</em>, the unique identifier for the <em>cited</em> document</li> <li>&lsquo;<em>Syr</em>&rsquo;, the year of publication of the <em>citing</em> document</li> <li>&lsquo;<em>Tyr</em>&rsquo;, the year of publication of the <em>cited</em> document</li> <li>&lsquo;<em>SC</em>&rsquo;, the cluster ID of the <em>citing</em> document</li> <li>&lsquo;<em>TC</em>&rsquo;, the cluster ID of the <em>cited </em>document</li> </ul> <p><strong>&lsquo;RPM_Nodelist.csv&rsquo;</strong> is a comma-separate values file that consists of the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Id</em>&rsquo;, the unique ID assigned to a document that corresponds with the edgelist</li> <li>&lsquo;<em>Reference string</em>&rsquo;, the reference string of the document</li> <li>&lsquo;<em>WoS ID</em>&rsquo;, the unique accession number assigned to a document by the Web of Science. These can be used to query WoS to find further data on all papers via the &lsquo;UT= &rsquo; field tag.</li> <li>&lsquo;<em>Authors</em>&rsquo;, all authors formatted by full last name and initials</li> <li>&lsquo;<em># of authors&rsquo;</em>, number of authors</li> <li>&lsquo;<em>Title</em>&rsquo;, title of document</li> <li>&lsquo;<em>Publication year</em>&rsquo;, publication year of document</li> <li>&lsquo;<em>Document type</em>&rsquo;, document type defined by WoS (e.g. article, review, etc.)</li> <li>&lsquo;<em>Total references</em>&rsquo;, total number of references within a documents bibliography as recorded by WoS</li> <li>&lsquo;<em>Total WoS citations</em>&rsquo;, total number of citations recorded to a document from other documents indexed in the Web of Science</li> <li>&lsquo;<em>Indegree</em>&rsquo;, total number of within network citations (i.e. counting only citations from other papers retrieved by our query)</li> <li>&lsquo;<em>Outdegree</em>&rsquo;, total number of within network references (i.e. counting only reference to other papers retrieved by our query)</li> <li>&lsquo;<em>Degree</em>&rsquo;, total number of node connections (i.e. indegree + outdegree)</li> <li>&lsquo;<em>Class</em>&rsquo;, variable used to distinguish between RPM&rsquo;s publications (&lsquo;RPM&rsquo;) and the citing documents (&lsquo;CITE&rsquo;)</li> <li>&lsquo;<em>Cluster</em>&rsquo;, provides the cluster membership number as discussed within the manuscript. This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.67 | 25 clusters).</li> </ul> <p><strong>References</strong></p> <p>[1] Leng, R. I., Leng. G. (Under review). A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research. <em>J. Neuroendocrinol</em></p> <p>All bibliographic data included in this study are derived originally from Clarivate&trade; (Web of Science&trade;) and downloaded in January 2024. &copy; Clarivate 2024. All rights reserved.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Dataset for Citation Network Analysis on Quality Cues for Meat Purchases

<p>Data are retrieved from two databases permissible to the Vosviewer&reg; software: Scopus and Web of Science. Three sets of Boolean search strings are used to find articles in the databases:</p> <p>(i) [<em>(&ldquo;certification&rdquo; AND &ldquo;meat&rdquo;) OR (&ldquo;certification&rdquo; AND &ldquo;dairy&rdquo;)</em>]</p> <p>(ii) [<em>(&ldquo;credence&rdquo;&nbsp;AND&nbsp;meat) OR&nbsp;(&ldquo;experience&rdquo;&nbsp;AND meat)</em>]<em> </em></p> <p><em>(iii) </em>[<em>(&ldquo;meat&rdquo; AND &ldquo;intrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;extrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;quality cue&rdquo;)</em>].</p> <p>The searches targeted the title, abstract and keywords sections for the Scopus database, and the topic section for the Web of Science database. The search was limited to peer-reviewed articles published in the English language; no timeline restrictions were set for the initial search.&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

unarXive: All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation Network (open subset)

<h2><strong>Description</strong></h2><p>unarXive is a scholarly data set containing publications' structured full-text, annotated in-text citations, linked non-text content (mathematical notation, figure/table captions) and a citation network.</p><p>The data is generated from all LaTeX sources on <a href="https://arxiv.org/">arXiv</a> and therefore of higher quality than data generated from PDF files.</p><p>Typical uses are</p><ul><li>Training of ML models (citation recommendation, summarization, LLMs)</li><li>Citation context analysis</li><li>Bibliographic analyses</li></ul><h2><strong>Access</strong></h2><p>┏━━━━━━━━━━━━━━━━━━━━━━━━━━┓<br>┃ &nbsp;<a href="https://github.com/IllDepence/unarXive/raw/master/doc/unarXive_data_sample.tar.gz"><strong>D O W N L O A D &nbsp; S A M P L E</strong></a> &nbsp; ┃<br>┗━━━━━━━━━━━━━━━━━━━━━━━━━━┛</p><p>Regarding the full data set, please note the following:</p><blockquote><p><strong>Note</strong>: this Zenodo record is the "open subset" of unarXive, which contains all permissively licensed papers from arXiv.org. You can find the <a href="https://doi.org/10.5281/zenodo.7752754">full version here</a>.</p></blockquote><p>The code used for generating the data set is <a href="https://github.com/IllDepence/unarXive">publicly available</a>.</p>

opencc-by-sa-4.0Mar 2023View details →
dryad36/100

The disruption index suffers from citation inflation and is confounded by shifts in scholarly citation practice: synthetic citation networks for bibliometric null models

Open the record for dataset details and reuse information.

publicFeb 2025View details →
zenodo32/100

Citation network nodelist, edgelist, and visualization (UNIGE MA Thesis)

<p>The data in this deposit were used in the <a href="https://archive-ouverte.unige.ch/unige:140928">author&#39;s Thesis</a> for the degree of Master of Arts in Philosophy with Specialization in the Philosophy of Science, under the supervision of Prof. Marcel Weber (Department of Philosophy, University of Geneva, Switzerland).</p> <p>This deposit contains three files: (i) a list of nodes; (ii) a list of edges; and (iii) a digital image of the network as it appears in the Appendix A of the Thesis. The lists of nodes and edges contained here were manually built by the author, and the visualization was obtained using the software Gephi.</p> <p>The nodes represent specific texts related to the historical development of Expected Utility Theory (as explained in section 3.1 of the Thesis), corresponding to the texts registered in the Citation Network Texts section of the References of the Thesis. The data and visualization in this deposit use the convention &quot;authorsORIGINALYEAR&quot;. For example, Pareto (1909/1979) is represented as &quot;pareto1909&quot;, and Safra et al. (1990a) as &quot;safra.etal1990a&quot;. The color of each node depends on the subsection of the historical description (section 3.2) they appear in, corresponding to the categories of the &quot;topic&quot; column of the nodelist. And their sizes are proportional to how many times they were mentioned by others in the network.</p> <p>The edges represent citations between texts and their colors represent <em>mention types</em> (as defined in section 3.1), such that agreement are in green, disagreements are in light blue, and neutral mentions are in light gray. The reasons for the final categorization of potentially unclear mention types are noted in the &quot;reasons&quot; column.</p>

opencc-by-4.0Aug 2020View details →
dryad32/100

Data from: A stochastic generative model for citation networks among academic papers

<p>We propose a stochastic generative model to represent a directed graph constructed by citations among academic papers, where nodes and directed edges represent papers with discrete publication time and citations respectively. The proposed model assumes that a citation between two papers occurs with a probability based on the type of the citing paper, the importance of cited paper, and the difference between their publication times, like the existing models. We consider the out-degrees of citing paper as its type, because, for example, survey paper cites many papers. We approximate the importance of a cited paper by its in-degrees. In our model, we adopt three functions: a logistic function for illustrating the numbers of papers published in discrete time, an inverse Gaussian probability distribution function to express the aging effect based on the difference between publication times, and an exponential distribution (or a generalized Pareto distribution) for describing the out-degree distribution. We consider that our model is a more reasonable and appropriate stochastic model than other existing models and can perform complete simulations without using original data. In this paper, we first use the Web of Science database and see the features used in our model. By using the proposed model, we can generate simulated graphs and demonstrate that they are similar to the original data concerning the in- and out-degree distributions, and node triangle participation. In addition, we analyze two other citation networks derived from physics papers in the arXiv database and verify the effectiveness of the model.</p>

opencc-zeroJun 2022View details →
zenodo32/100

Results from "The Field-Dependent Nature of PageRank Values in Citation Networks"

<p>This repo contains a gzipped archive of the resulting dataframes from the analyses discussed in our manuscript</p>

opencc-by-4.0Dec 2022View details →
dryad32/100

Data from: A stochastic generative model for citation networks among academic papers

Open the record for dataset details and reuse information.

publicJun 2022View details →
dryad28/100

Research on key generic technology prediction based on graph neural networks under the perspective of patent citation - An example from the field of genetic engineering

<p>In this research, we adopted graph neural network models for key generic prediction based on cited patent data. Through the construction of the patent citation network and the design of a key generic evaluation system, 20879 relevant patents and 51,610 irrelevant patents were screened out. Further, we utilized the LDA topic model to interpret technical topics at a finer granularity. Finally, to test the effectiveness of this method, we took the field of genetic engineering as an example for key generic technology prediction, with an accuracy rate of 95%.</p>

opencc-zeroJan 2024View details →
zenodo28/100

3DCP.fyi - A Comprehensive Citation Network Graph on the State of the Art in 3D Concrete Printing

<p><span>Research in digital fabrication, specifically in 3D concrete printing (3DCP), has seen a substantial increase in publication output in the past five years, making it hard to keep up with the latest developments. The 3dcp.fyi database aims to provide the research community with a comprehensive, up-to-date, and manually curated literature dataset documenting the development of the field from its early beginnings in the late 1990s to its resurgence in the 2010s until today. The data set is compiled using a systematic approach. A thorough literature search was conducted in scientific databases, following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) scheme. This was then enhanced iteratively with non-indexed literature through a snowball citation search. The authors of the articles were assigned unique and persistent identifiers (ORCID&reg; IDs) through a systematic process that combined querying APIs systematically and manually curating data. The works in the data set also include references to other works, as long as those referenced works are also included within the same data set. A citation network graph is created where scientific articles are represented as vertices, and their citations to other scientific articles are the edges. The constructed network graph is subjected to detailed analysis using specific graph-theoretic algorithms, like PageRank. These algorithms evaluate the structure and connections within the graph, yielding quantitative metrics. Currently, the high-quality dataset contains more than 2600 manually curated scientific works, including journal articles, conference articles, books, and theses, with more than 40000 cross-references and 2000 authors, opening up the possibility for more detailed analysis. The data is published on <a href="https://3dcp.fyi">https://3dcp.fyi</a>, ready for import into several reference managers, and is continuously updated. We encourage researchers to enrich the database by submitting their publications, adding missing works, or suggesting new features.</span></p>

restrictedcc-by-nc-4.0Apr 2024View details →
dryad28/100

Data from: How new concepts become universal scientific approaches – insights from citation network analysis of agent-based complex systems science

Open the record for dataset details and reuse information.

publicFeb 2018View details →
dryad28/100

Key generic technology prediction in patent citation using graph neural networks

Open the record for dataset details and reuse information.

publicJan 2024View details →
zenodo12/100

unarXive: All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation Network (full)

<h2><strong>Description</strong></h2><p>unarXive is a scholarly data set containing publications' structured full-text, annotated in-text citations, linked non-text content (mathematical notation, figure/table captions) and a citation network.</p><p>The data is generated from all LaTeX sources on <a href="https://arxiv.org/">arXiv</a> and therefore of higher quality than data generated from PDF files.</p><p>Typical uses are</p><ul><li>Training of ML models (citation recommendation, summarization, LLMs)</li><li>Citation context analysis</li><li>Bibliographic analyses</li></ul><h2><strong>Access</strong></h2><p>┏━━━━━━━━━━━━━━━━━━━━━━━━━━┓<br>┃ &nbsp;<a href="https://github.com/IllDepence/unarXive/raw/master/doc/unarXive_data_sample.tar.gz"><strong>D O W N L O A D &nbsp; S A M P L E</strong></a> &nbsp; ┃<br>┗━━━━━━━━━━━━━━━━━━━━━━━━━━┛</p><p>To download the whole data set send an access request and note the following:</p><blockquote><p><strong>Note</strong>: this Zenodo record is the "full" version of unarXive, which was generated from all of arXiv.org <i>including non-permissively licensed papers</i>. Make sure that your use of the data is compliant with the paper's licensing terms.¹<br>Alternatively you can use the <a href="https://doi.org/10.5281/zenodo.7752615">unarXive open subset</a>.</p><p>¹ For information on papers' licenses use <a href="https://info.arxiv.org/help/bulk_data/index.html">arXiv's bulk metadata access</a>.</p></blockquote><p>The code for generating the data set is <a href="https://github.com/IllDepence/unarXive">publicly available</a>.</p>

restrictedMar 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record