Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
214
datasets available to search
ShareScore release 0.7.1
Dataset results
214 results for “citations”
Semantic annotation of PLoS journal citation contexts
<p>Dataset </p>
F I G U R E 1 The relationship between 10 in Characteristics of papers that affect citations in the Journal of Fish Biology
F I G U R E 1 The relationship between 10 of the most influential variables extracted from papers published in the Journal of Fish Biology (between January 2010 and March 2021) and paper impact (high, medium, low) (±95% confidence interval). Papers were categorized as high impact if they fell into the top quartile of normalized citation counts and low impact if they fell into the bottom quartile of normalized citation counts; all other papers were classified as medium impact.
Reliability of citations of medRxiv preprints in articles published on COVID-19 in the world leading medical journals
<p>Articles published on COVID in 2020 in the BMJ, The Lancet, the JAMA and the NEJM were manually screened to identify all articles citing at least one preprint from medRxiv. We searched PubMed, Google and Google Scholar to assess if the preprint had been published in a peer-reviewed journal, and when. Published articles were screened to assess if the title, data or conclusions were identical to the preprint version.</p>
KGlove+Glove embedding for MAKG citation network
<p>Entity embedding by using KGlove+Glove on MAKG citation network (citation network from <a href="https://zenodo.org/record/4617285/files/08.PaperReferences.nt.bz2?download=1">https://zenodo.org/record/4617285/files/08.PaperReferences.nt.bz2?download=1</a>)</p>
Supplementary Material of "How Does Author Affiliation Affect Preprint Citation Count? Analyzing Citation Bias at the Institution and Country Level"
<p>The source code and dataset for the following paper:</p> <p>Nishioka, C., Färber, M., and Saier, T. How Does Author Affiliation Affect Preprint Citation Count? Analyzing Citation Bias at the Institution and Country Level. In Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2022 (JCDL '22), 2022.</p> <p> </p>
A guiding diagram for the selection of a CiTO citation function for a given in-text citation
<p>A decision model for the selection of a Citation Typing Ontology (CiTO, <a href="http://purl.org/spar/cito">http://purl.org/spar/cito</a>) citation function to use for the annotation of the citation intent of an examined in-text citation based on its context. The first large row of the diagram contains three macro-categories: (1) “Reviewing …”, (2) “Affecting …”, and (3) “Referring …”. Each macro-category has at least two subcategories, and each subcategory refers to a set of citation functions. The first row defines the suitable citation functions for it with the help of a guiding sentence to be completed according to the chosen sub-category and citation function. </p> <p>To annotate the in-text citation intent using the guiding scheme (see Figure), we use a priority ranked strategy that works as follows: </p> <ol> <li>we match each in-text citation against at least one of the three macro-categories, i.e. “Reviewing ...”, “Affecting ...” and “Referring ...” (first row in Figure);</li> <li>for each macro-category we have selected in (1), we choose one or more citation functions from those provided by CiTO;</li> <li>in case we select only one citation function, then we annotate the in-text citation intent with such a value; otherwise</li> <li>we calculate the priority of each citation function we have selected by summing its value in parenthesis (from 0.1 to 0.6) with the corresponding value in the y-axis (from 10 to 50) and in the x-axis (from 1 to 8) shown in Figure. The smaller the sum, the more priority the citation function has. For instance, the priority of the citation function “confirms” is 11.2 that is higher than the one of the citation function “describes”, which is 43.2. Finally, we select the citation function that has higher priority and annotate the in-text citation function with it.</li> </ol>
PLOS ONE – a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)
<p>This is a dataset used in and produced by research described in article "PLOS ONE - a case study of citation analysis of research papers based on the data in an open citation index (The OpenCitations Corpus)" that is translation of the original Polish text "PLOS ONE – studium przypadku analizy cytowań prac naukowych na podstawie danych otwartego indeksu cytowań (OpenCitations Corpus)" published by EBiB bulletin (2017, No 176).</p> <p>Data were extracted, as nodes (PLOS_cited_nodes.csv) and edges (PLOS_edges.csv) files from the OpenCitations Corpus (http://opencitations.net/download) on 2017.07.25 and describe all cited papers published by PLOS ONE (nodes), and all citing relations (edges). The research was conducted using Gephi (https://gephi.org/) platform so the same source data are also avaiable as GEXF file (for "one-click" import capabilities). In addition, the same data are published in NET format (but be warned that due to this format limitations, information about the publication year of papers has been lost) used by PAJEK platform, as it is very popular tool for analysis of network data.</p> <p>Published figures have prefix names corresponding to figures captions in the original paper, where they have been thoroughly discussed. This data set contains also the additional figure not published in the article, showing most cited paper with citing chains of articles of lenght not greater than 3.<br> These pictures have much better quality than those published in the article, which allows for "drill down"/zoom-in analysis and large format printing.</p>
NCBI genes with citation counts
<p>This dataset is a tab-separated file containing a list of human genes as listed on NCBI's Gene database and annotated with citation counts. The columns are:</p> <p>taxid gene_id unused gene_symbol gene_description gene_type citation_count</p> <p>Example:</p> <p>9606 3553 - IL1B interleukin 1 beta protein-coding 2462</p>
Open access in Africa: scopus citation data
<p>The following citation dataset was retrieved from Scopus in June 24, 2017 (3am, Western Indonesian time).</p> <p>It consists of 3 sets of data based on our searches. Each search was saved both in 'csv' and 'bib':</p> <ol> <li>OA_Africa_inTitle.xxx: "Open Access" AND Africa IN TITLE</li> <li>OA_Africa_inTitle_inAbstract_inKeywords.xxx: "Open Access" AND Africa IN TITLE, IN ABSTRACT, IN KEYWORDS</li> <li>OAmovement_Africa_inTitle_inAbstract_inKeywords.xxx: "Open Access movement" AND Africa IN TITLE, IN ABSTRACT, IN KEYWORDS</li> </ol> <p>The access to Scopus was provided by The Central Library of Institut Teknologi Bandung (Indonesia)</p>
An example of citation graphs showing year by year dynamics
<p>This is an example of citation graphs showing year by year dynamics. On x-axis there are years of publication while on y-axis the total number of citing articles for a given paper. Size of node also shows the total number of citing articles. This visualization is made with help of the Gephi application.</p>
Citations to Astronomy Journals 1: The growth of interdisciplinarity - Data Supplement
<p>This repository contains the data used in the blog "Citations to Astronomy Journals 1: The growth of interdisciplinarity", Michael J. Kurtz and Edwin Henneken. Each file has a header with a description of the data contained in the file. Table 1 consists of the bibstems for the journals in the main sample. Table 2 contains the individual data for each journal in the format journal, indicator, value, year. Table 3 has these data for all refereed journals (including the journals in the main sample, and all the rest). All three data files are ASCII files with space-separated columns.</p> <p>The term "bibstem" is the journal abbreviation used within the Astrophysics Data System. The complete list of bibstems is provided here: http://adsabs.harvard.edu/abs_doc/journals2.html. The bibstem is used in the ADS bibliographic identifier ("bibcode") with the convention that ampersands are replaced by plus signs (see: http://adsabs.github.io/help/actions/bibcode).</p>
Evolution of citations received by the Journal of Contemporary Administration (RAC) from journals indexed in Scopus (and their respective fields of knowledge, 1999-2018)
<p>This figure shows the evolution of the number of citations received by the RAC among the journals indexed in Scopus. * The query was performed on the Scopus platform on 11/15/2018, so the last year of the illustrated period still did not count on the total number of citations. You can see this figure and other comments by using the Editorial for Issue 22(6)2018 of The Journal of Contemporary Administration (Revista de Administração Contemporânea-RAC). www.rac.anpad.org.br</p>
DOI-to-DOI citations of articles '10.12685/027.7-4-2-155', '10.18452/9093' and '10.11588/PB.2012.1.9398'
<p>This data set contains the references to publications with a DOI from my publications:</p> <ul> <li>Baierer, Konstantin ; Zumstein, Philipp (2016): Verbesserung der OCR in digitalen Sammlungen von Bibliotheken. <em>027.7 : Zeitschrift für Bibliothekskultur = Journal for Library Culture</em>, 4(2):72-83. <a href="https://doi.org/10.12685/027.7-4-2-155">https://doi.org/10.12685/027.7-4-2-155</a></li> <li>Kim, Timotheus Chang-whae ; Zumstein, Philipp (2016): Semiautomatische Katalogisierung und Normdatenverknüpfung mit Zotero im Index Theologicus. <em>LIBREAS. Library ideas,</em> [12]:29 47-56. <a href="https://doi.org/10.18452/9093">https://doi.org/10.18452/9093</a></li> <li>Zumstein, Philipp (2012): Die Rolle des Semantic Web für Bibliotheken: Linked Open Data und mehr: Welche Strategien können hier die Bibliotheken in die Zukunft führen? <em>Perspektive Bibliothek ,</em> 1(1):81-102. <a href="https://doi.org/10.11588/PB.2012.1.9398">https://doi.org/10.11588/PB.2012.1.9398</a></li> </ul>
Data and code for: Community structure in co-inventor networks affects time to first citation for patents
<p>This package provides the datasets and programming code needed to reproduce the results reported in the article "Community structure in co-inventor networks affects time to first citation for patents".</p> <p>v2: Added data and code pertaining to randomized-community-association test and updated README file.</p>
Data set of the article: Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus
<p>Data of investigation published in the article "Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus".</p> <p>Abstract of the article:</p> <p>Search engine optimization (SEO) constitutes the set of methods designed to increase the visibility of, and the number of visits to, a web page by means of its ranking on the search engine results pages. Recently, SEO has also been applied to academic databases and search engines, in a trend that is in constant growth. This new approach, known as academic SEO (ASEO), has generated a field of study with considerable future growth potential due to the impact of open science. The study reported here forms part of this new field of analysis. The ranking of results is a key aspect in any information system since it determines the way in which these results are presented to the user. The aim of this study is to analyse and compare the relevance ranking algorithms employed by various academic platforms to identify the importance of citations received in their algorithms. Specifically, we analyse two search engines and two bibliographic databases: Google Scholar and Microsoft Academic, on the one hand, and Web of Science and Scopus, on the other. A reverse engineering methodology is employed based on the statistical analysis of Spearman’s correlation coefficients. The results indicate that the ranking algorithms used by Google Scholar and Microsoft are the two that are most heavily influenced by citations received. Indeed, citation counts are clearly the main SEO factor in these academic search engines. An unexpected finding is that, at certain points in time, WoS used citations received as a key ranking factor, despite the fact that WoS support documents claim this factor does not intervene.</p>
Zenodo data and software citation links captured by the Asclepias Broker
<p>The dataset was retrieved from the Asclepias Broker early January 2019 after having performed a full harvesting and deduplication cycle from a clean database with zero citation links.</p> <p>The dataset contains citation links from three discovery systems: the NASA Astrophysics Datasystem (ADS), Crossref Event Data and Europe PMC. Only citation links with a target DOI in the DOI prefix 10.5281 (Zenodo’s DOI prefix) were kept.</p>
Merged citation library dataset of 3 bibliographic databases for use in testing deduplication tools
<p>Endnote xml file of complete search results - not de-duplicated. </p> <p>Search query (cannabis based medicines, endocannabinoid system modulators and cannabinoids tested in animal models of pathological or injury-related persistent pain) conducted on 9 April 2019. </p> <p>Searched databases: Web of Science, Embase and PubMed</p> <p>The citations retrieved from the searches of each database have been merged together, there is a total of 15216 citations. </p> <p> </p>
All Computer Science Papers @ arXiv.org -- A High-Quality Gold Standard for Citation-based Tasks
<p>We propose a newly-created gold standard <strong>data set for citation-based tasks</strong>. This gold standard is based on <strong>all computer science papers in arXiv.org</strong>.</p> <p><strong>Abstract</strong>. Analyzing and recommending citations with their specific citation contexts have recently received much attention due to the growing number of available publications. Although data sets such as CiteSeerX have been created for evaluating approaches for such tasks, those data sets exhibit striking defects. This is understandable if one considers that both information extraction and entity linking as well as entity resolution need to be performed. In this paper, we propose a new evaluation data set for citation-dependent tasks based on arXiv.org publications. Our data set is characterized by the fact that it exhibits almost zero noise in the extracted content and that all citations are linked to their correct publications. Besides the pure content, available on a sentence-basis, cited publications are annotated directly in the text via global identifiers. As far as possible, referenced publications are further linked to DBLP. Our data set consists of over 15M sentences and is freely available for research purposes. It can be used for training and testing citation-based tasks, such as recommending citations, determining the functions or importance of citations, and summarizing documents based on their citations.</p> <p> </p> <p>More information can be found in our <strong>publication "<a href="http://www.lrec-conf.org/proceedings/lrec2018/pdf/283.pdf">A High-Quality Gold Standard for Citation-based Tasks</a>" (LREC'18)</strong>.</p> <p>You can cite the data set as follows:</p> <pre><code>@inproceedings{DBLP:conf/lrec/0001TJ18, author = {Michael F{\"{a}}rber and Alexander Thiemann and Adam Jatowt}, title = "{A High-Quality Gold Standard for Citation-based Tasks}", booktitle = "{Proceedings of the Eleventh International Conference on Language Resources and Evaluation}", series = "{LREC'18}", location = "{Miyazaki, Japan}", year = {2018}, url = {http://www.lrec-conf.org/proceedings/lrec2018/summaries/283.html} } </code></pre> <p> </p>
Citation count error data for "Data inaccuracy quantification and uncertainty propagation for bibliometric indicators"
<p>This is the original collected data on citation count errors resulting from citation matching errors in Web of Science data for the publication "Data inaccuracy quantification and uncertainty propagation for<br>bibliometric indicators". The first column, <code>CITCOUNT_ALL</code>, gives the total (corrected) citation count for a publication, which is the citation count according to WoS plus the additionally manually identified citations (missed by WoS's algorithm). The second column, <code>CITCOUNT_WOS</code>, is the WoS citation count. The numeric difference between the two column values in one row is the number of additionally manually identified citations.</p>
Diversity in citations to a single study: Supplementary data set for citation context network analysis
<p><strong>Introduction</strong></p> <p>This document describes the data set used for all analyses in 'Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited' accepted for publication in <em>Quantitative Science Studies</em> [1].</p> <p><strong>Data Collection</strong></p> <p>The data collection procedure has been fully described [1]. Concisely, the data set contains bibliometric data collected from Web of Science Core Collection via the University of Edinburgh’s Library subscription concerning all papers that cited a cohort study, Paul <em>et al.</em> [2], in the period <1985. This includes a full list of citing papers, and the citations between these papers. Additionally, it includes textual passages (citation contexts) from 343 citing papers, which were manually recovered from the full-text documents accessible via the University of Edinburgh’s Library subscription. These data have been cleaned, converted into network readable datasets, and are coded into particular classifications reflecting content, which are described fully in the supplied code book and within the manuscript [1]. </p> <p><strong>Data description</strong></p> <p>All relevant data can be found in the attached file 'Supplementary_material_Leng_QSS_2021.xlsx', which contains the following five workbooks:</p> <ul> <li><strong>“Overview”</strong> includes a list of the content of the workbooks.</li> <li><strong>“Code Book”</strong> contains the coding rules and definitions used for the classification of findings and paper titles.</li> <li><strong>“Node attribute list”</strong> includes a workbook containing all node attributes for the citation network, which includes Paul et al. [2] and its citing papers as of 1984. Highlighted in yellow at the bottom of this workbook is two papers that were discarded due to duplication - remove these if analysing this dataset in a network analysis. The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Label</em>, the formal citation of the paper to which data within this row corresponds. Citation is in the following format: last name of first author, year of publication, journal of publication, volume number, start page, and DOI (if available). </li> <li><em>Title</em>, the paper title for the paper in question.</li> <li><em>Publication_year</em>, the year of publication.</li> <li><em>Document_type, </em>the document type (e.g. review, article)</li> <li><em>WoS_ID</em>, the paper’s unique Web of Science accession number.</li> <li><em>Citation_context</em>, a column specifying whether citation context data is available from that paper</li> <li><em>Explanans</em>, the title explanans terms for that paper;</li> <li><em>Explanandum</em>, the explanandum terms for that paper.</li> <li><em>Combined_Title_Classification</em>, the combined terms used for fig 2 of the published manuscript.</li> <li><em>Serum_cholesterol_(SC)</em>, a column identifying papers that cited the serum cholesterol findings.</li> <li><em>Blood_Pressure_(BP), </em>a column identifying papers that cited the blood pressure findings.</li> <li><em>Coffee_(C),</em> a column identifying papers that cited the coffee findings.</li> <li><em>Diet_(D), </em>a column identifying papers that cited the dietary findings.</li> <li><em>Smoking_(S), </em>a column identifying papers that cited the smoking findings.</li> <li><em>Alcohol_(A), </em>a column identifying papers that cited the alcohol findings.</li> <li><em>Physical_Activity_(PA),</em> a column identifying papers that cited the physical activity findings.</li> <li><em>Body_Fatness (BF), </em>a column identifying papers that cited the body fatness findings.</li> <li><em>Indegree,</em> the number of within network citations to that paper, calculated for the network shown in Fig 4 of the manuscript.</li> <li><em>Outdegree</em>, the number of within network references of that paper as calculated for the network in Fig 4.</li> <li><em>Main_component</em>, a column specifying whether a node is contained in the largest weakly connect component as shown in Fig 4 of the manuscript.</li> <li><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Fig 5).</li> </ol> <ul> <li><strong>“Edge list”</strong> includes a workbook including the edges for the network. The columns refer to:</li> </ul> <ol> <li><em>Source</em>, contains the node identifier of the citing paper.</li> <li><em>Target,</em> contains the node identifier of the cited paper.</li> </ol> <ul> <li><strong>“Citation context classification</strong>” includes a workbook containing the WoS accession number for the paper analysed, and any finding category discussed in that paper established via context analysis (see the code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Finding_Class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <ul> <li><strong> “Citation context data”</strong> includes a workbook containing the WoS accession number for papers in which citation context data was available, the citation context passages, the reference number or format of Paul et al. within the citing paper, and the finding categories discussed in those contexts (see code book for definitions). The columns refer to:</li> </ul> <ol> <li><em>Id</em>, the node identifier</li> <li><em>Citation_context</em>, the passage copied from the full text of the citing paper containing discussion of the findings of Paul et al.</li> <li><em>Reference_in_citing_article</em>, the reference number or format of Paul et al. within the citing paper.</li> <li><em>Finding_class, </em>the findings discussed from Paul et al. within the body of the citing paper. </li> </ol> <p><strong>Software recommended for analysis</strong></p> <p>For the analyses performed within the manuscript, Gephi version 0.9.2 was used [3], and both the edge and node lists are in a format that is easily read into this software. The Sci2 tool was used to parse data initially [4].</p> <p><strong>Notes</strong></p> <ol> <li>Leng, R. I. (Forthcoming). Diversity in citations to a single study: A citation context network analysis of how evidence from a prospective cohort study was cited. Quantitative Science Studies.</li> <li>Paul, O., Lepper, M. H., Phelan, W. H., Dupertuis, G. W., Macmillan, A., McKean, H., <em>et al.</em> (1963). A longitudinal study of coronary heart disease. <em>Circulation, </em><strong>28</strong>, 20-31. <a href="https://doi.org/10.1161/01.cir.28.1.20">https://doi.org/10.1161/01.cir.28.1.20</a>.</li> <li>Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media.</li> <li>Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></li> </ol>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.