Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
214
datasets available to search
ShareScore release 0.7.1
Dataset results
214 results for “citations”
Foundation Species Revisited: Citation Analysis of Ellison et al. 2005
Ecologists and environmental scientists often prioritize research efforts with conservation importance. Dominant, widespread, or locally abundant species at low risk of extinction receive relatively little attention unless they are invasive. Native foundation species create habitats and environmental conditions that support many associated species and modulate local-scale ecosystem processes, but the generally high local or regional abundance of foundation species may lead to less research about them. We used citation analysis (2005-2014) to examine research following from a suggestion to identify and study foundation species while they were still common and not threatened. We explored the use and expanding definition of the foundation species concept, as well as the trajectory and ecological focus of research on foundation species throughout the world in 378 papers published in this nine-year span. Contemporary authors who cite key papers defining a foundation species pay little attention to its actual definition and species studied in this context rarely were identified as foundation species. Although functions and roles of foundation species, such as creating unique microclimates or supporting dependent species, are being studied, less research is focused on identifying them before they are threatened or lost from the ecosystem that they otherwise define. Invasive species were identified as the most common threat to foundation species. Our citation analysis and synthesis provides a new conceptual framework linking identification of and research about foundation species with their functional roles and our ability to manage emerging threats to them.
Corpus of critical citations contexts
<p>We present here a corpus of 505 critical citation contexts, i.e. a set of sentences or propositions that contain at least one citation of a study towards which the author(s) has/have a negative opinion. Those contexts come from other existing annotated corpora, from our readings about critical citation and disagreement in science, and from contexts manually annotated by native speakers of English. We have re-annotated all those contexts in order to be sure that they match our definition of critical citations. This corpus can be helpful to train tools dedicated to the automatic retrieval of critical citations. English (2024-02-20)</p>
Synthetic Dataset of Citation Strings in 12 Styles
<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland. </p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br> title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}}, <br> author = {Iana Atanassova and Marc Bertin},<br> year = {2024},<br> booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br> address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git </p>
Dataset related to the manuscript: "An open-source integrated framework for the automation of citation collection and screening in systematic reviews"
<p>Dataset related to the manuscript: “An open-source integrated framework for the automation of citation collection and screening in systematic reviews”, to be used together with the code stored at https://github.com/AD-Papers-Material/BART_SystReviewClassifier to reproduce the results.</p> <p>There are three datasets:<br> - The Record data collected from the online scientific databases;<br> - The session journal which describes the search session, i.e., how many records were collected and from which source, for each query/session pairs.<br> - The session data which is the outcome of the classification and review tasks;</p>
Data on a citation context analysis focusing on natural sciences and social sciences and humanities
<p>This dataset contains data on citation context analysis between natural sciences (NS) and social sciences and humanities (SSH). In particular, the data were created through manual coding of each citation between papers related to SDG7 (renewable energy) and SDG13 (climate change) and papers cited by them. This dataset consists of 9 files, associated with the article: Nishikawa, K. How and why are citations between disciplines made? A citation context analysis focusing on natural sciences and social sciences and humanities. Scientometrics (2023). <a href="https://doi.org/10.1007/s11192-023-04664-y">https://doi.org/10.1007/s11192-023-04664-y</a></p> <p> </p> <p>The files are numbered as follows:</p> <ul> <li>00 – README</li> <li>01 – Data by citation pair for SDG7 (original)</li> <li>02 – Data by citation pair for SDG13 (original)</li> <li>03 – Data by mention location for SDG7 (original)</li> <li>04 – Data by mention location for SDG13 (original)</li> <li>05 – Data by citation pair for SDG7 (additional)</li> <li>06 – Data by citation pair for SDG13 (additional)</li> <li>07 – Data by mention location for SDG7 (additional)</li> <li>08 – Data by mention location for SDG13 (additional)</li> </ul> <p>See README for more information.</p>
IPBES Data Management Tutorials - Session 5.5: References and citation manager: Zotero
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>This session, <em>References and citation manager: Zotero, </em>reviews why IPBES recommends Zotero to manage references and provides links to key resources.</p>
Topics in Research on International Relations as Clusters of Citation Links
<p>Data, scripts, and results of a memetic topic clustering of citation links in papers published 2011-2015 in the specialty of political science that is dealing with international relations </p> <p>Supplementary Information to the paper about "Topics as clusters of citation links to highly cited sources: The case of research on international relation" by Frank Havemann<em> </em>(published 2021 in the OA-journal<em> Quantitative Science Studies</em> 2 (1): 204–223). <a href="https://doi.org/10.1162/qss_a_00108">https://doi.org/10.1162/qss_a_00108</a></p>
Citation data of arXiv eprints and the associated quantitatively-and-temporally normalised impact metrics
<p><strong>Data collection</strong></p> <p>This dataset contains information on the eprints posted on arXiv from its launch in 1991 until the end of 2019 (1,589,006 unique eprints), plus the data on their citations and the associated impact metrics. Here, eprints include preprints, conference proceedings, book chapters, data sets and commentary, i.e. every electronic material that has been posted on arXiv. </p> <p>The content and metadata of the arXiv eprints were retrieved from the arXiv API (https://arxiv.org/help/api/) as of 21st January 2020, where the metadata included data of the eprint’s title, author, abstract, subject category and the arXiv ID (the arXiv’s original eprint identifier). In addition, the associated citation data were derived from the Semantic Scholar API (https://api.semanticscholar.org/) from 24th January 2020 to 7th February 2020, containing the citation information in and out of the arXiv eprints and their published versions (if applicable). Here, whether an eprint has been published in a journal or other means is assumed to be inferrable, albeit indirectly, from the status of the digital object identifier (DOI) assignment. It is also assumed that if an arXiv eprint received <em>c</em><sub>pre</sub> and <em>c</em><sub>pub</sub> citations until the data retrieval date (7th February 2020) before and after it is assigned a DOI, respectively, then the citation count of this eprint is recorded in the Semantic Scholar dataset as <em>c</em><sub>pre</sub> + <em>c</em><sub>pub</sub>. Both the arXiv API and the Semantic Scholar datasets contained the arXiv ID as metadata, which served as a key variable to merge the two datasets.</p> <p>The classification of research disciplines is based on that described in the arXiv.org website (https://arxiv.org/help/stats/2020_by_area/). There, the arXiv subject categories are aggregated into several disciplines, of which we restrict our attention to the following six disciplines: Astrophysics (‘astro-ph’), Computer Science (‘comp-sci’), Condensed Matter Physics (‘cond-mat’), High Energy Physics (‘hep’), Mathematics (‘math’) and Other Physics (‘oth-phys’), which collectively accounted for 98% of all the eprints. Those eprints tagged to multiple arXiv disciplines were counted independently for each discipline. Due to this overlapping feature, the current dataset contains a cumulative total of 2,011,216 eprints. </p> <p>Some general statistics and visualisations per research discipline are provided in the original article (Okamura, 2022), where the validity and limitations associated with the dataset are also discussed.</p> <p> </p> <p><strong>Description of columns (variables)</strong></p> <ul> <li><strong>arxiv_id</strong> : arXiv ID</li> <li><strong>category</strong> : Research discipline</li> <li><strong>pre_year</strong> : Year of posting v1 on arXiv</li> <li><strong>pub_year</strong> : Year of DOI acquisition</li> <li><strong>c_tot</strong> : No. of citations acquired during 1991–2019</li> <li><strong>c_pre</strong> : No. of citations acquired before and including the year of DOI acquisition</li> <li><strong>c_pub</strong> : No. of citations acquired after the year of DOI acquisition</li> <li><strong>c_<em>yyyy</em></strong> (<em>yyyy</em> = 1991, …, 2019) : No. of citations acquired in the year <em>yyyy</em> (with ‘<em>yyyy</em>’ running from 1991 to 2019)</li> <li><strong>gamma</strong> : The quantitatively-and-temporally normalised citation index</li> <li><strong>gamma_star</strong> : The quantitatively-and-temporally standardised citation index</li> </ul> <p><em>Note:</em> The definition of the quantitatively-and-temporally normalised citation index (γ; ‘gamma’) and that of the standardised citation index (γ*; ‘gamma_star’) are provided in the original article (Okamura, 2022). Both indices can be used to compare the citational impact of papers/eprints published in different research disciplines at different times. </p> <p> </p> <p><strong>Data files</strong></p> <p>A comma-separated values file (‘<strong>arXiv_impact.csv</strong>’) and a Stata file (‘<strong>arXiv_impact.dta</strong>’) are provided, both containing the same information.</p> <p> </p>
Citation network data sets for 'Oxytocin – a social peptide? Deconstructing the evidence'
<p><strong>Introduction</strong></p> <p>This note describes the data sets used for all analyses contained in the manuscript 'Oxytocin - a social peptide?’<a href="#_ftn1">[1]</a> </p> <p><strong>Data Collection</strong></p> <p>The datasets described here were originally retrieved from Web of Science (WoS) Core Collection via the University of Edinburgh’s library subscription <a href="#_ftn2">[2]</a>. The aim of the original study for which these data were gathered was to survey peer-reviewed primary studies on oxytocin and social behaviour. To capture relevant papers, we used the following query:</p> <p><em>TI = (“oxytocin” OR “pitocin” OR “syntocinon”) AND TS = (“social*” OR “pro$social” OR “anti$social”)</em></p> <p>The final search was performed on the 13 September 2021. This returned a total of 2,747 records, of which 2,049 were classified by WoS as ‘articles’. Given our interest in primary studies <em>only</em> – articles reporting original data – we excluded all other document types. We further excluded all articles sub-classified as ‘book chapters’ or as ‘proceeding papers’ in order to limit our analysis to primary studies published in peer-reviewed academic journals. This reduced the set to 1,977 articles. All of these were published in the English language, and no further language refinements were unnecessary.</p> <p>All available metadata on these 1,977 articles was exported as plain text ‘flat’ format files in four batches, which we later merged together via Notepad++. Upon manually examination, we discovered examples of papers classified as ‘articles’ by WoS that were, in fact, reviews. To further filter our results, we searched all available PMIDs in PubMed (1,903 had associated PMIDs - ~96% of set). We then filtered results to identify all records classified as ‘review’, ‘systematic review’, or ‘meta-analysis’, identifying 75 records <a href="#_ftn3">[3]</a> (thus, ~4% of records classified by WoS were classified as reviews in PubMed). After examining a sample and agreeing with the PubMed classification, these were removed these from our dataset - leaving a total of 1,902 articles.</p> <p>From these data, we constructed two datasets via parsing out relevant reference data via the Sci2 Tool <a href="#_ftn4">[4]</a>. First, we constructed a ‘node-attribute-list’ by first linking unique reference strings (‘Cite Me As’ column in WoS data files) to unique identifiers, we then parsed into this dataset information on the identify of a paper, including the title of the article, all authors, journal publication, year of publication, total citations as recorded from WoS, and WoS accession number. Second, we constructed an ‘edge-list’ that records the citations from a <em>citing paper</em> in the ‘Source’ column and identifies the <em>cited paper</em> in the ‘Target’ column, using the unique identifies as described previously to link these data to the node-attribute-list.</p> <p>We then constructed a network in which papers are nodes, and citation links between nodes are directed edges between nodes. We used Gephi Version 0.9.2 <a href="#_ftn5">[5]</a> to manually clean these data by merging duplicate references that are caused by different reference formats or by referencing errors. To do this, we needed to retain both all retrieved records (1,902) as well as including <em>all</em> of their references to papers whether these were included in our original search or not. In total, this produced a network of 46,633 nodes (unique reference strings) and 112,520 edges (citation links). Thus, the average reference list size of these articles is ~59 references. The mean indegree (within network citations) is 2.4 (median is 1) for the entire network reflecting a great diversity in referencing choices among our 1,902 articles.</p> <p>After merging duplicates, we then restricted the network to include <em>only</em> articles fully retrieved (1,902), and retrained <em>only</em> those that were connected together by citations links in a large interconnected network (i.e. the largest component). In total, 1,892 (99.5%) of our initial set were connected together via citation links, meaning a total of ten papers were removed from the following analysis – and these were neither connected to the largest component, nor did they form connections with one another (i.e. these were ‘isolates’).</p> <p>This left us with a network of 1,892 nodes connected together by 26,019 edges. <strong><em>It is this network that is described by the ‘node-attribute-list’ and ‘edge-list’ provided here</em></strong>. This network has a mean in-degree of 13.76 (median in-degree of 4). By restricting our analysis in this way, we lose 44,741 unique references (96%) and 86,501 citations (77%) from the full network, but retain a set of articles tightly knitted together, all of which have been fully retrieved due to possessing certain terms related to oxytocin AND social behaviour in their title, abstract, or associated keywords.</p> <p>Before moving on, we calculated indegree for all nodes in this network – this counts the number of citations to a given paper from other papers within this network – and have included this in the <em>node-attribute-list</em>. We further clustered this network via modularity maximisation via the Leiden algorithm <a href="#_ftn6">[6]</a>. We set the algorithm to resolution 1, and allowed the algorithm to run over 100 iterations and 100 restarts. This gave <em>Q</em>=0.43 and identified seven clusters, which we describe in detail within the body of the paper. We have included cluster membership as an attribute in the node-attribute-list.</p> <p>For additional analysis, we also analysed the full reference list data to examine the most commonly cited references between 2016 and 2021 - the results of this are described in OTSOC_Cited_2016-2021.csv. This takes the reference lists of all retrieved papers within the network and examines their full reference lists (including references to other papers not contained within the network). These data were cleaned by matching DOIs and manual cleansing. </p> <p><strong>Data description</strong></p> <p>We include here two network datasets: (i) ‘OTSOC-node-attribute-list.csv’ consists of the attributes of 1,892 primary articles retrieved from WoS that include terms indicating a focus on oxytocin and social behaviour; (ii) ‘OTSOC-edge-list.csv’ records the citations between these papers. Together, these can be imported into a range of different software for network analysis; however, we have formatted these for ease of upload into Gephi 0.9.2. Finally, we include (iii) 'OTSOC_Cited_2016-2021' that lists all papers cited by >10 papers in the OTSOC network following any analysis of the bibliographies of retrieved papers. Below, we detail their contents:</p> <p><strong>1. ‘OTSOC-node-attribute-list.csv’</strong> is a comma-separate values file that contains all node attributes for the citation network (n=1,892) analysed in the paper. The columns refer to:</p> <p><em>Id</em>, the unique identifier</p> <p><em>Label</em>, the reference string of the paper to which the attributes in this row correspond. This is taken from the ‘Cite Me As’ column from the original WoS download. The reference string is in the following format: last name of first author, publication year, journal, volume, start page, and DOI (if available). </p> <p><em>Wos_id</em>, unique Web of Science (WoS) accession number. These can be used to query WoS to find further data on all papers via the ‘UT= ’ field tag.</p> <p><em>Title</em>, paper title.</p> <p><em>Authors</em>, all named authors.</p> <p><em>Journal, </em>journal of publication.</p> <p><em>Pub_year</em>, year of publication.</p> <p><em>Wos_citations</em>, total number of citations recorded by WoS Core Collection to a given paper as of 13 September 2021</p> <p><em>Indegree</em>, the number of within network citations to a given paper, calculated for the network shown in Figure 1 of the manuscript.</p> <p><em>Cluster</em>, provides the cluster membership number as discussed within the manuscript (Figure 1). This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.43|7 clusters)</p> <p><strong>2. ‘OTSOC-edge -list.csv’</strong> is a comma-separated values file that contains all citation links between the 1,892 articles (n=26,019). The columns refer to:</p> <p><em>Source</em>, the unique identifier of the citing paper.</p> <p><em>Target, </em>the unique identifier of the cited paper.</p> <p><em>Type, </em>edges are ‘Directed’, and this column tells Gephi to regard all edges as such.</p> <p><em>Syr_date, </em>this contains the date of publication of the citing paper.</p> <p><em>Tyr_date, </em>this contains the date of publication of the cited paper.</p> <p><strong>3. 'OTSOC_Cited_2016-2021.csv'</strong> is a comma-separated values file that contain citations to all cited references that were cited by at least 10 of the retrieved papers within the OTSOC network published from 2016 onwards. The columns refer to: </p> <p><em>Reference, </em>the cited reference string extracted from the bibliographies of retrieved papers.</p> <p><em>Publication year, </em>the publication year of the cited reference.</p> <p><em>DOI</em>, the DOI of the cited reference. </p> <p><em>indegree_2016, </em>the total number of citations to a cited reference from papers published in 2016 and contained within the OTSOC network. </p> <p><em>indegree_2017, </em>the total number of citations to a cited reference from papers published in 2017 and contained within the OTSOC network. </p> <p><em>indegree_2018, </em>the total number of citations to a cited reference from papers published in 2018 and contained within the OTSOC network. </p> <p><em>indegree_2019, </em>the total number of citations to a cited reference from papers published in 2019 and contained within the OTSOC network. </p> <p><em>indegree_2020, </em>the total number of citations to a cited reference from papers published in 2020 and contained within the OTSOC network. </p> <p><em>indegree_2021, </em>the total number of citations to a cited reference from papers published in 2021 and contained within the OTSOC network. </p> <p><em>total indegree 2016-21</em>, the total number of citation to a cited reference from papers published between 2016-2021 and contained within the OTSOC network. </p> <p><strong>Software recommended for analysis</strong></p> <p>Gephi version 0.9.2 was used for the visualisations within the manuscript, and both files can be read and into Gephi without modification.</p> <p><strong>Notes</strong></p> <p><a href="#_ftnref1">[1]</a> Leng, G., Leng, R. I., Ludwig, M. (Submitted). Oxytocin – a social peptide? Deconstructing the evidence.</p> <p><a href="#_ftnref2">[2]</a> Edinburgh University’s subscription to Web of Science covers the following databases: (i) Science Citation Index Expanded, 1900-present; (ii) Social Sciences Citation Index, 1900-present; (iii) Arts & Humanities Citation Index, 1975-present; (iv) Conference Proceedings Citation Index- Science, 1990-present; (v) Conference Proceedings Citation Index- Social Science & Humanities, 1990-present; (vi) Book Citation Index– Science, 2005-present; (vii) Book Citation Index– Social Sciences & Humanities, 2005-present; (viii) Emerging Sources Citation Index, 2015-present.</p> <p><a href="#_ftnref3">[3]</a> For those interested, the following PMIDs were identified as ‘articles’ by WoS, but as ‘reviews’ by PubMed: ‘34502097’ ‘33400920’ ‘32060678’ ‘31925983’ ‘31734142’ ‘30496762’ ‘30253045’ ‘29660735’ ‘29518698’ ‘29065361’ ‘29048602’ ‘28867943’ ‘28586471’ ‘28301323’ ‘27974283’ ‘27626613’ ‘27603523’ ‘27603327’ ‘27513442’ ‘27273834’ ‘27071789’ ‘26940141’ ‘26932552’ ‘26895254’ ‘26869847’ ‘26788924’ ‘26581735’ ‘26548910’ ‘26317636’ ‘26121678’ ‘26094200’ ‘25997760’ ‘25631363’ ‘25526824’ ‘25446893’ ‘25153535’ ‘25092245’ ‘25086828’ ‘24946432’ ‘24637261’ ‘24588761’ ‘24508579’ ‘24486356’ ‘24462936’ ‘24239932’ ‘24239931’ ‘24231551’ ‘24216134’ ‘23955310’ ‘23856187’ ‘23686025’ ‘23589638’ ‘23575742’ ‘23469841’ ‘23055480’ ‘22981649’ ‘22406388’ ‘22373652’ ‘22141469’ ‘21960250’ ‘21881219’ ‘21802859’ ‘21714746’ ‘21618004’ ‘21150165’ ‘20435805’ ‘20173685’ ‘19840865’ ‘19546570’ ‘19309413’ ‘15288368’ ‘12359512’ ‘9401603’ ‘9213136’ ‘7630585’</p> <p><a href="#_ftnref4">[4]</a> Sci2 Team. (2009). Science of Science (Sci2) Tool. Indiana University and SciTech Strategies. Stable URL: <a href="https://sci2.cns.iu.edu">https://sci2.cns.iu.edu</a></p> <p><a href="#_ftnref5">[5]</a> Bastian, M., Heymann, S., & Jacomy, M. (2009). Gephi: an open source software for exploring and manipulating networks. International AAAI Conference on Weblogs and Social Media. Gephi is available via <a href="https://gephi.org/">https://gephi.org/</a></p> <p><a href="#_ftnref6">[6]</a> Traag, V. A., Waltman, L., & van Eck, N. J. (2019). From Louvain to Leiden: guaranteeing well-connected communities. Scientific reports, 9(1), 5233. <a href="https://doi.org/10.1038/s41598-019-41695-z">https://doi.org/10.1038/s41598-019-41695-z</a></p>
Triangle of Biomedicine Framework to Analyze the Citations' Impact on Categories Dissemination in the PubMed Database
<p>This is the data and the most relevant script of the paper 'Triangle of Biomedicine Framework to Analyze the Citations’ Impact on Categories Dissemination in the PubMed Database'.</p>
Types, open citations, closed citations, publishers, and participation reports of Crossref entities
<p>This publication contains several datasets that have been used in the paper "Crowdsourcing open citations with CROCI – An analysis of the current status of open citations, and a proposal" submitted to the <a href="https://www.issi2019.org/">17th International Conference on Scientometrics and Bibliometrics (ISSI 2019)</a>, available at <a href="https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/">https://opencitations.wordpress.com/2019/02/07/crowdsourcing-open-citations-with-croci/</a>.</p> <p>Additional information about the analyses described in the paper, including the code and the data we have used to compute all the figures, is available as a Jupyter notebook at <a href="https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb">https://github.com/sosgang/pushing-open-citations-issi2019/blob/master/script/croci_nb.ipynb</a>. The datasets contain the following information.</p> <p><strong>non_open.zip:</strong> it is a zipped (~5 GB unzipped) CSV file containing the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, dated October 2018. All the entity types retrieved from Crossref were aligned to one of following five categories: journal, book, proceedings, dataset, other. The open CC0 citation data we used came from the CSV dump of <a href="https://doi.org/10.6084/m9.figshare.6741422.v3">most recent release of COCI dated 12 November 2018</a>. The number of closed citations was calculated by subtracting the number of open citations to each entity available within COCI from the value “is-referenced-by-count” available in the Crossref metadata for that particular cited entity, which reports all the DOI-to-DOI citation links that point to the cited entity from within the whole Crossref database (including those present in the Crossref ‘closed’ dataset).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>doi:</em> the DOI of the publication in Crossref;</li> <li><em>type:</em> the type of the publication as indicated in Crossref;</li> <li><em>cited_by:</em> the number of open citations received by the publication according to COCI;</li> <li><em>non_open:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>croci_types.csv:</strong> it is a CSV file that contains the numbers of open citations and closed citations received by the entities in the Crossref dump used in our computation, as collected in the previous CSV file, alligned in five classes depening on the entity types retrieved from Crossref: <em>journal</em> (Crossref types: journal-article, journal-issue, journal-volume, journal), <em>book</em> (Crossref types: book, book-chapter, book-section, monograph, book track, book-part, book-set, reference-book, dissertation, book series, edited book), <em>proceedings</em> (Crossref types: proceedings-article, proceedings, proceedings-series), <em>dataset</em> (Crossref types: dataset), <em>other</em> (Crossref types: other, report, peer review, reference-entry, component, report-series, standard, posted-content, standard-series).</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>type:</em> the type publication between "journal", "book", "proceedings", "dataset", "other";</li> <li><em>label:</em> the label assigned to the type for visualisation purposes;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publication type according to COCI;</li> <li><em>crossref_close_cit:</em> the number of closed citations received by the publication according to Crossref + COCI.</li> </ul> <p><strong>publishers_cits.csv:</strong> it is a CSV file that contains the top twenty publishers that received the greatest number of open citations. The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher</em>: the name of the publisher;</li> <li><em>doi_prefix</em>: the list of DOI prefixes used assigned by the publisher;</li> <li><em>coci_open_cit</em>: the number of open citations received by the publications of the publisher according to COCI;</li> <li><em>crossref_close_cit</em>: the number of closed citations received by the publications of the publishers according to Crossref + COCI;</li> <li><em>total_cit</em>: the total number of citations received by the publications of the publisher (= <em>coci_open_cit</em> + <em>crossref_close_cit</em>).</li> </ul> <p><strong>20publishers_cr.csv: </strong>it is a CSV file that contains the numbers of the contributions to open citations made by the twenty publishers introduced in the previous CSV file as of 24 January 2018, according to the data available through the Crossref API. The counts listed in this file refers to the number of publications for which each publisher has submitted metadata to Crossref that include the publication’s reference list. The categories 'closed', 'limited' and 'open' refer to publications for which the reference lists are not visible to anyone outside the Crossref Cited-by membership, are visible only to them and to Crossref Metadata Plus members, or are visible to all, respectively. In addition, the file also record the total number of publications for which the publisher has submitted metadata to Crossref, whether or not those metadata include the reference lists of those publications.</p> <p>The columns of the CSV file are the following ones:</p> <ul> <li><em>publisher: </em>the name of the publisher;</li> <li><em>open: </em>the number of publications in Crossref with an 'open' visibility for their reference lists;</li> <li><em>limited: </em>the number of publications in Crossref with an 'limited' visibility for their reference lists;</li> <li><em>closed: </em>the number of publications in Crossref with an 'closed' visibility for their reference lists;</li> <li><em>overall_deposited:</em> the overall number of publications for which the publisher has submitted metadata to Crossref.</li> </ul>
Checklists for Software Citation: what you need to know
<p>The FORCE11 Software Citation Implementation working group has been developing practical guidance in the form of checklists, which are aimed at authors, reviewers, developers, editors and publishers. The checklists help these audiences to implement software citation principles in their workflows. This talk will give a quick introduction to these checklists and explain how they can be used to improve the practice of open research by giving credit for software.</p>
[review paper] Sustaining the 'Frozen Footprints' of Scholarly Communication through Open Citations_Dataset
<p>This dataset belongs to the review titled <em>Sustaining the ‘Frozen Footprints’ of Scholarly Communication through Open Citations</em>. The review explores the developments in the open citations movement, the OpenCitations infrastructure, and the Initiative for Open Citations (I4OC), providing a comprehensive overview of key milestones and initiatives.</p> <p>The dataset includes bibliographic and citation data for 174 scholarly outputs and 149 blogposts analyzed in the review. These outputs were drawn from a range of sources, including journal articles, conference proceedings, and other scholarly materials. The data has been curated to adhere to open citation principles, ensuring it is structured, separable, and openly accessible. It is provided in accordance with the licenses and terms of use of the original databases.</p> <p>Researchers, practitioners, and policymakers can use this dataset to explore the evolution of open citations and to further their understanding of the connections between scholarly works in the field of open research.</p>
A dataset from a survey investigating disciplinary differences in data citation
<p><strong>GENERAL INFORMATION</strong></p> <p><em>Title of Dataset: </em> A dataset from a survey investigating disciplinary differences in data citation</p> <p><em>Date of data collection:</em><strong> </strong>January to March 2022</p> <p><em>Collection instrument: </em>SurveyMonkey</p> <p><em>Funding:</em> Alfred P. Sloan Foundation</p> <p><br> <strong>SHARING/ACCESS INFORMATION</strong></p> <p><em>Licenses/restrictions placed on the data: </em>These data are available under a <a href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0 license</a> </p> <p><em>Links to publications that cite or use the data: </em></p> <p>Gregory, K., Ninkov, A., Ripp, C., Peters, I., & Haustein, S. (2022). Surveying practices of data citation and reuse across disciplines. Proceedings of the 26th International Conference on Science and Technology Indicators. <em>International Conference on Science and Technology Indicators</em>, Granada, Spain. https://doi.org/10.5281/ZENODO.6951437</p> <p>Gregory, K., Ninkov, A., Ripp, C., Roblin, E., Peters, I., & Haustein, S. (2023). <em>Tracing data:<br> A survey investigating disciplinary differences in data citation.</em> Zenodo. https://doi.org/10.5281/zenodo.7555266</p> <p><br> <strong>DATA & FILE OVERVIEW</strong></p> <p><em>File List</em></p> <ul> <li>Filename: MDCDatacitationReuse2021Codebookv2.pdf<br> <em>Codebook</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.csv<br> <em>Dataset format in csv</em></li> <li>Filename: MDCDataCitationReuse2021surveydatav2.sav<br> <em>Dataset format in SPSS</em></li> <li>Filename: MDCDataCitationReuseSurvey2021QNR.pdf<br> <em>Questionnaire</em></li> </ul> <p><em>Additional related data collected that was not included in the current data package: </em>Open ended questions asked to respondents</p> <p><br> <strong>METHODOLOGICAL INFORMATION</strong></p> <p><em>Description of methods used for collection/generation of data: </em></p> <p>The development of the questionnaire (Gregory et al., 2022) was centered around the creation of two main branches of questions for the primary groups of interest in our study: researchers that reuse data (33 questions in total) and researchers that do not reuse data (16 questions in total). The population of interest for this survey consists of researchers from all disciplines and countries, sampled from the corresponding authors of papers indexed in the Web of Science (WoS) between 2016 and 2020. </p> <p>Received 3,632 responses, 2,509 of which were completed, representing a completion rate of 68.6%. Incomplete responses were excluded from the dataset. The final total contains 2,492 complete responses and an uncorrected response rate of 1.57%. Controlling for invalid emails, bounced emails and opt-outs (n=5,201) produced a response rate of 1.62%, similar to surveys using comparable recruitment methods (Gregory et al., 2020).</p> <p><em>Methods for processing the data: </em></p> <p>Results were downloaded from SurveyMonkey in CSV format and were prepared for analysis using Excel and SPSS by recoding ordinal and multiple choice questions and by removing missing values.</p> <p><em>Instrument- or software-specific information needed to interpret the data: </em></p> <p>The dataset is provided in SPSS format, which requires IBM SPSS Statistics. The dataset is also available in a coded format in CSV. The Codebook is required to interpret to values.</p> <p><br> <strong>DATA-SPECIFIC INFORMATION FOR: MDCDataCitationReuse2021surveydata</strong></p> <p><em>Number of variables:</em> 95</p> <p><em>Number of cases/rows:</em> 2,492</p> <p><em>Missing data codes:</em> 999 Not asked</p> <p>Refer to MDCDatacitationReuse2021Codebook.pdf for detailed variable information.</p>
A2.2a Digital repositories data citation practices. Supplementary material
<p>Data to complement the quantitative analysis of data citation practices in digital repositories based on metadata records from the re3data.org repositories registry.</p> <p>Data was retrieved using re3data.org API on 23-02-2023 and 06-03-2023 and processed using the OpenRefine software.</p> <p>Part of "A FAIR-enabling citation model for Cultural Heritage Objects" project activities.</p>
FAIR-CHO Citation Model Zotero Group Library Bibliography. Supplementary material
<p>This is a selected bibliography created during the project <em>A FAIR-enabling citation model for Cultural Heritage Objects</em>.</p> <p>This bibliography has been set up via a Zotero Library Group, by organizing it into subject folders representative of the project content. An initial list of descriptors was also defined to 'semantically' label the bibliographic references as they were collected.</p> <p>The dataset represents the Zotero Library on 1st August 2023.</p> <p>The dataset is published in .csv and .ris formats.</p> <p>See also on Zotero Groups: <a href="https://www.zotero.org/groups/4883319/cho_citation_model/library">https://www.zotero.org/groups/4883319/cho_citation_model/library</a>.</p>
Structured citations in the English Wikipedia
<p>This dataset contains the metadata of citations in the English Wikipedia that editors have input as citation templates (which use the Citation Style 1). It has been obtained from an XML dump of Wikipedia (2016-05-01), and was parsed with the wikiciteparser library. This library runs the Lua code used in Wikipedia to format such citations and generate structured metadata (such as COinS) from them.</p> <p>https://github.com/dissemin/wikiciteparser</p> <p>Sample:</p> <blockquote> <p>716551092 12 2016-04-22T10:19:33Z Anarchism cite journal {"PublisherName": "International Group of San Francisco", "Title": "Towards Anarchism", "URL": "http://www.marxists.org/archive/malatesta/1930s/xx/toanarchy.htm", "Authors": [{"link": "Errico Malatesta", "last": "Malatesta", "first": "Errico"}], "ID_list": {"OCLC": "3930443"}, "Periodical": "MAN!", "PublicationPlace": "Los Angeles"}<br /> 716551092 12 2016-04-22T10:19:33Z Anarchism cite journal {"Date": "2007-05-14", "URL": "http://www.theglobeandmail.com/servlet/story/RTGAM.20070514.wxlanarchist14/BNStory/lifeWork/home/", "Title": "Working for The Man", "Periodical": "The Globe and Mail", "Authors": [{"last": "Agrell", "first": "Siri"}]}<br /> 716551092 12 2016-04-22T10:19:33Z Anarchism cite web {"Date": "2006", "URL": "http://www.britannica.com/eb/article-9117285", "PublisherName": "Encyclop\u00e6dia Britannica Premium Service", "Periodical": "Encyclop\u00e6dia Britannica", "Title": "Anarchism"}<br /> 716551092 12 2016-04-22T10:19:33Z Anarchism cite journal {"Date": "2005", "Pages": "14", "Periodical": "The Shorter Routledge Encyclopedia of Philosophy", "Title": "Anarchism"}</p> </blockquote> <p>Credit: these citations have been input by Wikipedia editors (this dataset is therefore distributed under a CC-BY-SA license).</p> <p>Acknowledgments: this extraction process was carried on one of OpenJournal's servers.</p> <p>See also: citation identifiers extracted from Wikipedia (so, not looking at specific templates, but using regular expressions): http://dx.doi.org/10.6084/m9.figshare.1299540</p>
Data Citation Corpus Data File
<p>Data file for the fourth release of the Data Citation Corpus, produced by DataCite and Make Data Count as part of an ongoing grant project funded by the Wellcome Trust. <a href="https://makedatacount.org/data-citation/">Read more about the project</a>.</p> <p>The data file includes 10,697,745 data citation records (of which 9,682,257 represent unique dataset-publication pairs) in JSON and CSV formats. The JSON file is the version of record.</p> <p>Data is provided in batches of approximately 1 million records each. The publication date and batch number are included in the file name, ex: 2025-08-15-data-citation-corpus-01-v4.1.json.</p> <p>The data citations in the file originate from the following sources:</p> <ul> <li>DataCite Event Data</li> <li>Chan Zuckerberg Initiative (CZI) Science Knowledge Graph</li> <li>Aligning Science Across Parkinson’s (ASAP)</li> <li>Europe PMC</li> </ul> <p>Each data citation record is comprised of:</p> <ul> <li> <p>A pair of identifiers: An identifier for the dataset (a DOI or an accession number) and the DOI of the publication (journal article or preprint) in which the dataset is cited </p> </li> <li> <p>Metadata for the cited dataset and for the citing publication </p> </li> </ul> <p>The data file includes the following fields:</p> <div> <table> <tbody> <tr> <td> <p>Field</p> </td> <td> <p>Description</p> </td> <td> <p>Required?</p> </td> </tr> <tr> <td> <p>id</p> </td> <td> <p>Internal identifier for the citation</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>created</p> </td> <td> <p>Date of item's incorporation into the corpus</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>updated</p> </td> <td> <p>Date of item's most recent update in corpus</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>repository</p> </td> <td> <p>Repository where cited data is stored</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>publisher</p> </td> <td> <p>Publisher for the article citing the data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>journal</p> </td> <td> <p>Journal for the article citing the data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>title</p> </td> <td> <p>Title of cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>publication</p> </td> <td> <p>DOI of article where data is cited</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>dataset</p> </td> <td> <p>DOI or accession number of cited data</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>publishedDate</p> </td> <td> <p>Date when citing article was published</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>source</p> </td> <td> <p>Source where citation was harvested</p> </td> <td> <p>Yes</p> </td> </tr> <tr> <td> <p>subjects</p> </td> <td> <p>Subject information for cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>affiliations</p> </td> <td> <p>Affiliation information for creator of cited data</p> </td> <td> <p>No</p> </td> </tr> <tr> <td> <p>funders</p> </td> <td> <p>Funding information for cited data</p> </td> <td> <p>No</p> </td> </tr> </tbody> </table> </div> <p><strong> </strong></p> <p>Additional documentation about the citations and metadata in the file is available on the <a href="https://makedatacount.org/find-a-tool/data-citation-corpus-documentation/">Make Data Count website</a>. </p> <p><strong>Notes on v4.1:</strong></p> <p>Version 4.1 of the Data Citation Corpus is a minor update to v4.0 that corrects (1) an error that occurred when a portion of DOI-DOI citations originating from Europe PMC were attributed to the wrong repository, and (2) a small number of DOI formatting errors in the "publication" field.</p> <p><strong>Notes on v4.0:</strong></p> <p>The fourth release of the Data Citation Corpus data file adds new citations from the following sources:</p> <ul> <li> <p dir="ltr">5.2 million data citations from <a href="https://europepmc.org/">Europe PMC</a> identified as "eupmc" in the source field. Ingest of these citations was performed 9 July 2025.</p> </li> <li> <p dir="ltr">139,647 data citations from DataCite Event Data for the period 1 January 2025 through 30 June 2025.</p> </li> </ul> <p dir="ltr">This release also includes the following new metadata enhancements:</p> <ul> <li> <p dir="ltr">Affiliation information for cited data from the Gene Expression Omnibus (GEO) repository, reonciled to Research Organization Registry (ROR) IDs where possible.</p> </li> <li> <p dir="ltr">Reconciliation of organization and funders names with the Research Organization Registry (ROR) for new citations from Event Data.</p> </li> <li> <p dir="ltr">Application of Field of Science subject terms to citation records originating from Europe PMC, based on disciplinary area of data repository.</p> </li> </ul> <p>Additional details about the above changes, including scripts used to perform the above tasks, are available in <a href="https://github.com/Make-Data-Count-Community/corpus-data-file" target="_blank" rel="noopener">GitHub</a>. </p> <p>Additional enhancements to the corpus are ongoing and will be addressed in the course of subsequent releases. Users are invited to submit feedback via <a href="https://github.com/Make-Data-Count-Community/data-citation-corpus-feedback">GitHub</a>. For general questions, email <a href="mailto:info@makedatacount.org">info@makedatacount.org</a>.</p>
Citation analysis of Brown 1988 and Charnov 1976, for Figure 1 of manuscript "Halloween Charnov", Calcagno et al. 2023
<p>This contains the R script (bibliom.txt) and the ciatation data files (three .csv files) needed to generate Figure 1 in manuscript "Taking fear back into the Marginal Value Theorem: the risk-MVT and optimal boldness", by Calcagno, Gorgnard, Hamelin and Mailleret, 2023.</p>
Dataset for Machine Learning Assisted Citation Screening for Systematic Reviews
<p>The work "Machine Learning Assisted Citation Screening for Systematic Reviews" explored the problem of citation screening automation using machine-learning (ML) with an aim to accelerate the process of generating <a href="https://en.wikipedia.org/wiki/Systematic_review#:~:text=Systematic%20reviews%20are%20a%20type,synthesize%20findings%20qualitatively%20or%20quantitatively." rel="nofollow">systematic reviews</a>. Manual process of citation screening involve two reviewers manually screening the searched studies using a predefined inclusion criteria. If the study passes the "inclusion" criteria, it is included for further analysis or is excluded. As apparant through manual screening process, the work considered citation screening as a binary classification problem whereby any ML classifier could be trained to separate the searched studies into these two classes (include and exclude).</p> <p> </p> <p>A physiotherapy citation screening dataset was used to test automation approaches and the dataset includes the studies identified for citation screening in an update to the systematic review by Hilfiker <em>et al.</em> The dataset included titles and abstracts (citations) from 31,279 (deduplicated: 25,540) studies identified during the search phase of this SR. These studies were already manually assessed for relevance and labelled by two reviewers into two mutually exclusive labels. The uploaded file consists of 25,540 data samples, with each data sample separated by a new line. It is a tab separated file and the data in it is structured as shown below. This dataset was manually labelled into include and exclude by Hilfiker <em>et al.</em></p> <p> </p> <table> <tbody> <tr> <td><strong>Title</strong></td> <td><strong>PMID</strong></td> <td><strong>Abstract </strong></td> <td><strong>Class</strong></td> <td><strong>MeSH terms (separated by a pipe)</strong></td> </tr> <tr> <td>Structured exercise improves physical functioning in women with stages I and II breast cancer: results of a randomized controlled trial. </td> <td>11157015</td> <td>Abstract PURPOSE: Self-directed and supervised exercise were compared with usual care in a clinical trial designed to evaluate the effect of structured exercise on physical functioning and other dimensions of health-related quality of life in women with stages I and II breast cancer. PATIENTS AND METHODS: One hundred twenty-three women with stages I and II breast cancer completed baseline evaluations of generic and disease- and site-specific health-related quality of life, aerobic capacity, and body weight. Participants were randomly allocated to one of three intervention groups: usual care (control group), self-directed exercise, or supervised exercise. Quality of life, aerobic capacity, and body weight measures were repeated at 26 weeks...</td> <td>include or exclude</td> <td>Clinical Trial | Comparative Study | Randomized Controlled Trial | Research Support, Non-U.S. Gov't | Antineoplastic Combined Chemotherapy Protocols | Breast Neoplasms | Breast Neoplasms | Breast Neoplasms | Chemotherapy, Adjuvant | Exercise | Female | Humans | Middle Aged | Neoplasm Staging | Quality of Life | Radiotherapy, Adjuvant</td> </tr> </tbody> </table> <p> </p> <p>If you use this dataset in your research, please cite our papers.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.