Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

214

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

214 results for “citation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Statistical Analysis of the Effect of Equations on Citations

<p>Statistical analysis of a data set of number of equations and number of citations of papers published in volumes 94 and 104 of the journal <em>Physical Review Letters</em>. This analysis is referred to by the paper <strong>Equation-dense papers receive fewer citations—in physics as well as biology</strong> in the <em>New Journal of Physics </em>(vol. 18, article 118003) by Andrew D Higginson and Tim W Fawcett.  http://iopscience.iop.org/article/10.1088/1367-2630/18/11/118003</p>

opencc-zeroJul 2016View details →
zenodo36/100

Interactive Citation Networks for Texas Archaeology

<p>This dataset includes the JavaScript used to generate two interactive citation networks. IntNetFig1 includes the network filtered by the giant component and nodes sized by OutDegree, and IntNetFig2 is the same network filtered by a degree of two,&nbsp;nodes sized by Eigenvector Centrality and colored by modularity class.</p>

opencc-by-4.0Jul 2016View details →
zenodo36/100

Crossref open citations

<p>This text file contains citations between documents classified as journal article, book content, conference paper, or preprint in Crossref. Each line in the file contains the DOI of a citing document followed by the DOI of a cited document. There are 147,161,531 documents classified as journal article, book content, conference paper, or preprint in Crossref. There are 1,723,702,594 citations between these documents in Crossref. The data has been extracted from Crossref&rsquo;s XML Metadata Plus Snapshot downloaded on February 10, 2025.</p>

opencc-zeroMay 2021View details →
dryad36/100

Research on the benefits of nature to people: How much overlap is there in citations and terms for 'nature' across disciplines?

<ol> <li class="ThesisAbstractCxSpFirst">Research on the diverse benefits of nature to people is characterised by a broad range of disciplines involved, encompassing a variety of approaches, methods and terminologies. While a diversity of approaches is valuable, it can lead to difficulties in integrating and sharing findings and could form a barrier to effective knowledge exchange, hindering the development and applications of research outputs.</li> <li class="ThesisAbstractCxSpMiddle">As a starting point for this scoping review, we chose four broad research areas (medicine, psychology, education and environment), selected to represent disparate approaches to research on the benefits of nature to people, within and across which to explore overlap in citations and terms used to describe nature.</li> <li class="ThesisAbstractCxSpMiddle">We conducted expert consultation and a snowball-based approach to source publications, resulting in a sample of 210 papers, spanning multiple disciplines within each of our four research areas. For each paper, we recorded the discipline of the journal in which it was published (publishing discipline), the discipline of its first author (first-author discipline), the number of times journals of each discipline were cited in its bibliography (cited discipline) and the term(s) used in the paper's title or abstract to describe the aspect of nature being explored (nature term).</li> <li class="ThesisAbstractCxSpMiddle">Cited disciplines were significantly different between publishing and first-author disciplines, with papers from psychology, education and public health citing distinct communities of papers. However, disciplines generally cited a wide range of other disciplines, with articles in medical journals being particularly broadly cited.</li> <li class="ThesisAbstractCxSpMiddle">Nature terms were significantly different between publishing and first-author disciplines, with some degree of consistency within disciplines (e.g., education papers consistently used a narrow range of nature terms, such as 'outdoor learning'). However, there was a notably high range of nature terms used within psychology and public health papers, indicating that research from these disciplines may be particularly prone to being overlooked by search strings.</li> <li class="ThesisAbstractCxSpLast">The wide range of disciplines cited is encouraging, since this indicates that diverse research areas are generally aware of each other's work. However, to avoid unnecessary expansion of nature terms and support searchability, we propose four key terms for nature: ('outdoor learning' OR 'outdoor education'), ('nature' OR 'natural'), ('green space' OR 'greenspace') and ('biodiversity' or 'trees'), which could be used across disciplines. We particularly propose that at least one of these be included in every paper, and all four should be included in review search strings. This is likely to result in better understanding of the valuable, disparate contributions made by different disciplines to this expanding and important topic.</li> </ol>

opencc-zeroDec 2023View details →
zenodo36/100

Reference Manager Data Citation Analysis

<p>DESCRIPTION:</p> <p>This package contains data used to analyze citation metadata completeness and correctness for several common reference managers used in scholarly research and several common repositories in the Earth, space, and environmental sciences.</p> <p>METHODS:</p> <p>Metadata fields for import and export methods and for 8 metadata fields (authors/creators, publisher, DOI, dataset title, version, access date, publication date, and resource type) were collected from reference managers via all import methods available (app or wizard and plugin) during summer 2024 from most recent software versions of all. To encode data, citation information for each dataset as imported by Reference Manager was compared to that registered for the DOI with DataCite. Correct metadata for each of 8 fields for both import and export was encoded as 0, incorrect as 1, and missing as '' or nan. See publication and software package for more information.</p> <p>FILES:</p> <p>FOLDER 'coded-data' contains files that include information (DOIs) about the data examined in this study, preserved copies of exported data citations used in the data interpretation and processing, and the processed data itself encoded in columns.</p> <p>FOLDER 'datacite-metadata-profiles' includes the raw metadata from each dataset DOI at the time of analysis, included for reproducibility purposes. &nbsp;</p> <p>FOLDER 'bibtex-files' includes the downloaded .bib files, where available, for each dataset DOI examined.</p> <p>See README file for more information.</p>

opencc-by-4.0Nov 2024View details →
zenodo36/100

Current and future sources of open citations

<p>Video recording of the presentation done in the context of the Austrian DataCite Consortium on 19 November 2021. The video finishes after the last question.</p>

opencc-by-4.0Nov 2021View details →
zenodo36/100

Inputs and results of "A qualitative and quantitative analysis of open citations to retracted articles: the Wakefield 1998 et al.'s case"

<p>This repository contains the datasets and visualizations generated in our work:&nbsp;<strong>&quot;A qualitative and quantitative analysis of open citations to retracted articles: the Wakefield 1998 et al.&rsquo;s case&quot;</strong>.</p> <p><strong>Note:</strong>&nbsp;the data are all contained inside the&nbsp;<strong><em>data.zip</em>&nbsp;</strong>file. You need to unzip the container to get access to all the files and directories listed below.</p> <p>The data (citations) gathered accompanied by&nbsp;their annotated characteristics&nbsp;are stored in&nbsp;<strong><em>data/</em>:</strong></p> <ul> <li><em><strong>&quot;cits_features.csv&quot;:&nbsp;</strong></em>a dataset containing&nbsp;all the entities (rows in the CSV) which have cited the&nbsp;Wakefield et al.&nbsp;retracted article, and a set of&nbsp;features characterizing each citing entity&nbsp;(columns in the CSV). The features included are:&nbsp;DOI (&quot;doi&quot;), year of publication (&quot;year&quot;), the title (&quot;title&quot;), the venue identifier (&quot;source_id&quot;), the title of the venue (&quot;source_title&quot;), yes/no value in case the entity is retracted as well (&quot;retracted&quot;), the subject area (&quot;area&quot;), the subject category (&quot;category&quot;), the sections of the in-text citations (&quot;intext_citation.section&quot;), the value of the reference pointer (&quot;intext_citation.pointer&quot;), the in-text citation function (&quot;intext_citation.intent&quot;), the in-text citation perceived sentiment (&quot;intext_citation.sentiment&quot;), and a yes/no value to denote whether the in-text citation context mentions the retraction of the cited entity&nbsp;&nbsp; &nbsp;(&quot;intext_citation.section.ret_mention&quot;).<br> <strong>Note:&nbsp;</strong>this dataset is licensed under a&nbsp;<a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.</li> <li><em><strong>&quot;cits_text.csv&quot;:&nbsp;</strong>this dataset stores the abstract (&quot;abstract&quot;) and the in-text citations context (&quot;intext_citation.context&quot;)&nbsp;</em>for&nbsp;each citing entity identified using the DOI value (&quot;doi&quot;).<br> <strong>Note:&nbsp;</strong>the data keep their original&nbsp;license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> </ul> <p><strong>Topic modeling</strong></p> <p>We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the&nbsp;<em><strong>topic_modeling/</strong></em>&nbsp;directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools and creating a completely customizable visual workflow [1].&nbsp;The topic modeling results for each textual feature are separated into two different folders,&nbsp;<em><strong>abstract/</strong></em>&nbsp;for the abstracts, and&nbsp;<em><strong>intext_cit/</strong></em>&nbsp;for the in-text citation contexts. Both the directories contain the datasets and visualizations generated using MITAO.&nbsp;</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>[1] Ferri, P., Heibi, I., Pareschi, L., &amp; Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135&ndash;149.&nbsp;<a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a></p>

opencc-zeroDec 2020View details →
dryad36/100

Labeled data for citation field extraction

<p>Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science. Citation field extraction is the task of parsing citations: given a citation string, extract authors, title, venue, doi etc. Since the number of citations is counted by hundreds millions, efficient computer based methods for this task are very important.</p> <p>The development of machine learning methods for citation field extraction requires ground truth: a large corpus of labeled citations. This dataset provides a very large (41M) corpus of labeled data obtained by the reverse process: we took structured citation lists and used BibTeX to generate labeled citation strings.</p>

opencc-zeroMar 2022View details →
zenodo36/100

Wikipedia Complete Citation Corpus

<p><strong>Wikipedia Complete Citation Corpus</strong> (<strong>WCCC</strong>) is a corpus of citations, references and sources mined from the English Wikipedia. WCCC was created as a knowledge base used in a machine-learning model for recommending reliable sources to support (or refute) a given textual claim, but can be used for many other purposes.</p> <p>The solution was described in the paper <em>&quot;Countering Disinformation by Finding Reliable Sources: a Citation-Based Approach&quot;</em>, presented at the 2022 International Joint Conference on Neural Networks (IJCNN 2022). Please refer to the article (in <a href="https://doi.org/10.1109/IJCNN55064.2022.9891941">conference proceedings</a> or <a href="https://home.ipipan.waw.pl/p.przybyla/bib/Learning_to_Cite_3_CR.pdf">authors&#39; version</a>) for more information on the process of mining the corpus, comparisons with similar resources and the role it plays in recommending sources.&nbsp;The research was done within the&nbsp;<a href="https://homados.ipipan.waw.pl/">HOMADOS</a>&nbsp;project at the&nbsp;<a href="https://ipipan.waw.pl/">Institute of Computer Science</a>, Polish Academy of Sciences.</p> <p>WCCC contains 4.8 million documents with 50.8 million citations of 24.3 million sources. The dataset is divided into 10 parts (WCCC-part0.zip to WCCC-part9.zip) with approximately the same size. Each of the parts contains batch archives (e.g. batch130.zip), each covering up to 1000 Wikipedia articles. An article is identified by its ID number and described by the following files:</p> <ul> <li>&lt;ID&gt;_text.txt: the textual content of the article,</li> <li>&lt;ID&gt;_citations.txt: the citations occurring in this article, saved as tab-separated values of (1) character offset in the textual content and (2) reference ID,</li> <li>&lt;ID&gt;_references.txt: the references cited in the article, saved as tab-separated values of (1) reference ID and (one or many) pairs of (2) source ID and (3) location in the source (e.g. page number),</li> <li>&lt;ID&gt;_sources.txt: the sources referenced in the article, saved as tab-separated values of (1) source ID and (2) source description (in wikicode).</li> <li>&lt;ID&gt;.txt: human-readable text, created by enriching textual content with the article title and reference IDs.</li> </ul> <p>Additionally, article metadata are included in the meta.tsv file. Each line describes a single article through the following tab-separated fields:</p> <ul> <li>article ID,</li> <li>title of the article,</li> <li>Wikipedia ID of the article, which can be used to access the article through URL, i.e. https://en.wikipedia.org/?curid=&lt;WIKI_ID&gt;</li> <li>length of the textual content of the article (number of characters),</li> <li>number of sources in the article,</li> <li>number of references in the article,</li> <li>number of citations in the article.</li> </ul> <p>Files metaTrain.tsv and metaTest.tsv contain the same information, but split into training and test set, as used in the work.</p> <p>Please refer to <a href="https://home.ipipan.waw.pl/p.przybyla/bib/Learning_to_Cite_3_CR.pdf">the paper</a> for an in-depth explanation of the data structure (citations, references, sources, etc.). WCCC was created using Wikipedia dump from 01.02.2021, but you can repeat the mining process using a different dump (or different procedure) by using the <a href="https://github.com/piotrmp/finding_reliable_sources">published source code</a>. If you intend to apply the corpus in a fact-checking use-case, you might also look at the <a href="https://dx.doi.org/10.5281/zenodo.6539087">evaluation datasets we publish separately</a>, one of which is based on WCCC with additional elements (e.g. source identifiers: URL/ISBN/DOI).</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Inputs and results of "A quantitative and qualitative citation analysis to retracted articles in the humanities domain"

<p>This repository contains the datasets and visualizations generated in our work: <strong>&quot;A quantitative and qualitative citation analysis to retracted articles in the humanities domain&quot;</strong>.</p> <p><strong>Note:</strong>&nbsp;the data are all contained inside the <strong><em>data.zip</em>&nbsp;</strong>file. You need to unzip the container to get access to all the files and directories listed below.</p> <p>The data (citations) gathered accompanied by&nbsp;their annotated characteristics&nbsp;are stored in&nbsp;<strong><em>data/</em>:</strong></p> <ul> <li><em>cits.csv:&nbsp;</em>a dataset containing&nbsp;all the entities (rows in the CSV) which have cited a retracted article in the humanities domain. Each citing entity (row) is accompanied by a set of&nbsp;features (columns) that characterizes it.<br> <strong>Note:&nbsp;</strong>this dataset is licensed under a&nbsp;<a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.</li> <li><em>content.csv: </em>a&nbsp;dataset containing&nbsp;the abstracts&nbsp;and the in-text citation&nbsp;contexts of all the citing entities gathered.<br> <strong>Note:&nbsp;</strong>the data keep their original&nbsp;license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> <li><em>excluded_hum_retractions.csv: </em>a list of the 12 humanities retracted articles with a humanities affinity score &lt; 2, therefore excluded from the analysis.&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Topic modeling</strong></p> <p>We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the&nbsp;<em><strong>topic_model/</strong></em>&nbsp;directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools and creating a completely customizable visual workflow [1].&nbsp;The directory <em><strong>workflow/ </strong></em>contains the workflows used in MITAO. The topic modeling results for each textual feature are separated into two different folders,&nbsp;<em><strong>abstract/</strong></em>&nbsp;for the abstracts, and&nbsp;<em><strong>cits_context/</strong></em>&nbsp;for the in-text citation contexts. Both the directories contain the following directories/files:&nbsp;</p> <ul> <li> <p><em><strong>datasets_and_views/:&nbsp;</strong></em>the datasets and visualizations generated using MITAO.&nbsp;&nbsp;</p> </li> <li> <p><em><strong>ldamodel_corpus_dict/:&nbsp;</strong></em>it contains the dictionary, the LDA topic&nbsp;model, and&nbsp;the tokenized and vectorized corpus.</p> </li> <li><em><strong>rawdata/: </strong></em>the textual collection, metadata, and stopwords used as input in the workflow of MITAO</li> </ul> <p>&nbsp;</p> <p><strong>References</strong></p> <p>[1] Ferri, P., Heibi, I., Pareschi, L., &amp; Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135&ndash;149.&nbsp;<a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a></p> <p>&nbsp;</p> <ol> </ol>

opencc-zeroAug 2021View details →
zenodo36/100

Citations for article https://doi.org/10.18637/jss.v036.i03

<p>Citations for the article &quot;Conducting Meta-Analyses in R with the metafor Package&quot; by Wolfgang Viechtbauer 2010-2022 in the Lens.org database.</p>

opencc-by-4.0Aug 2022View details →
dryad36/100

Forecasting the publication and citation outcomes of Covid-19 preprints

<p>The scientific community reacted quickly to the <em>Covid-19</em> pandemic in 2020, generating an unprecedented increase in publications. Many of these publications were released on preprint servers such as <em>medRxiv</em> and <em>bioRxiv</em>. It is unknown however how reliable these preprints are, and if they will eventually be published in scientific journals. In this study, we use crowdsourced human forecasts to predict publication outcomes and future citation counts for a sample of 400 preprints with high <em>Altmetric</em> scores. Most of these preprints were published within one year of upload on a preprint server (70%), and 46% of the published preprints appeared in a high-impact journal with a Journal Impact Factor of at least 10. On average, the preprints received 162 citations within the first year. We found that forecasters can predict if preprints will be published after one year and if the publishing journal has high impact. Forecasts are also informative with respect to preprints' rankings in terms of <em>Google</em> <em>Scholar</em> citations within one year of upload on a preprint server. For both types of assessment, we found statistically significant positive correlations between forecasts and observed outcomes. While the forecasts can help to provide a preliminary assessment of preprints at a faster pace than the traditional peer-review process, it remains to be investigated if such an assessment is suited to identify methodological problems in pre-prints. </p>

opencc-zeroSep 2022View details →
zenodo36/100

Citation network dataset covering the work of RP Millar and its citing literature

<p>The following describes the citation network datasets that underpins the manuscript &ldquo;A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research&rdquo; [1].</p> <p><strong>Data collection</strong></p> <p>We retrieved data from the Web of Science Core Collection under the University of Edinburgh&rsquo;s subscription in January 2024. We sought to retrieve all indexed papers of Professor Robert P. Millar (RPM). We searched the following <em>AU = (Millar, R)</em>, and then retrieved records that corresponded to his WoS profile (n=428 records) and an additional 49 paper that were authored by Robert but had not been included in his WoS record &ndash; validating the records against a CV of his published works.</p> <p>We retrieved the full citation history as record by Web of Science to these papers from other indexed records. The 477 RPM papers had been cited 21,677 times by 11,138 documents by date of retrieval, and removing self-citations left 19,256 citations by 10,719 documents. We then retrieved all metadata from WoS concerning the 477 RPM papers and the 10,719 citation papers, resulting in a dataset covering 11,196 documents.</p> <p><strong>Citation network dataset</strong></p> <p>We constructed a citation network dataset by parsing data from each paper&rsquo;s full bibliography consisting of:</p> <p>i. &lsquo;Edge-list&rsquo; that records citation links from a citing to a cited document. This is constructed by assigning unique IDs to each retrieved paper and to every unique reference string contained in their bibliographies. The edge list is composed of a &lsquo;Source&rsquo; column that contains the ID of the <em>citing</em> document and a &lsquo;Target&rsquo; column containing the IDs of its citations, with one record per row. Given that we were only interested in citations between the WoS retrieved documents, we discarded any reference string that represented a document outwith our search.</p> <p>ii. &lsquo;Node-attribute list&rsquo; that contains the ID, with relevant metadata contained in adjacent columns to identify documents, including authors, title of publication, journal, year of publication. We also parsed into this dataset the WoS full citation count for each paper and the total number of references in the bibliographies of each paper.</p> <p>This results in a dataset containing 11,196 nodes and 115,834 edges between nodes. We removed a total of 67 papers for which metadata was incomplete and/or corrupted. We further focussed on the largest interconnected component, removing nodes with no connections (isolates) or smaller components that were detached from the main network. We excluded papers &lt;10 references to remove meeting abstracts and other minor journal items, and papers not published in English. This resulted in a final dataset containing 10,901 nodes and 113,742 edges, and it is this dataset that we share as it is the basis for the analyses within the paper.</p> <p><strong>Description of dataset variables</strong></p> <p><strong>&lsquo;RPM_Edgelist.csv&rsquo;</strong> is a comma-separate values file that consists of all 113,742 citations between the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Source</em>&rsquo;, the unique identifier for the <em>citing </em>document</li> <li>&lsquo;<em>Target&rsquo;</em>, the unique identifier for the <em>cited</em> document</li> <li>&lsquo;<em>Syr</em>&rsquo;, the year of publication of the <em>citing</em> document</li> <li>&lsquo;<em>Tyr</em>&rsquo;, the year of publication of the <em>cited</em> document</li> <li>&lsquo;<em>SC</em>&rsquo;, the cluster ID of the <em>citing</em> document</li> <li>&lsquo;<em>TC</em>&rsquo;, the cluster ID of the <em>cited </em>document</li> </ul> <p><strong>&lsquo;RPM_Nodelist.csv&rsquo;</strong> is a comma-separate values file that consists of the 10,901 documents of the citation network analysed in the manuscript. The columns refer to:</p> <ul> <li>&lsquo;<em>Id</em>&rsquo;, the unique ID assigned to a document that corresponds with the edgelist</li> <li>&lsquo;<em>Reference string</em>&rsquo;, the reference string of the document</li> <li>&lsquo;<em>WoS ID</em>&rsquo;, the unique accession number assigned to a document by the Web of Science. These can be used to query WoS to find further data on all papers via the &lsquo;UT= &rsquo; field tag.</li> <li>&lsquo;<em>Authors</em>&rsquo;, all authors formatted by full last name and initials</li> <li>&lsquo;<em># of authors&rsquo;</em>, number of authors</li> <li>&lsquo;<em>Title</em>&rsquo;, title of document</li> <li>&lsquo;<em>Publication year</em>&rsquo;, publication year of document</li> <li>&lsquo;<em>Document type</em>&rsquo;, document type defined by WoS (e.g. article, review, etc.)</li> <li>&lsquo;<em>Total references</em>&rsquo;, total number of references within a documents bibliography as recorded by WoS</li> <li>&lsquo;<em>Total WoS citations</em>&rsquo;, total number of citations recorded to a document from other documents indexed in the Web of Science</li> <li>&lsquo;<em>Indegree</em>&rsquo;, total number of within network citations (i.e. counting only citations from other papers retrieved by our query)</li> <li>&lsquo;<em>Outdegree</em>&rsquo;, total number of within network references (i.e. counting only reference to other papers retrieved by our query)</li> <li>&lsquo;<em>Degree</em>&rsquo;, total number of node connections (i.e. indegree + outdegree)</li> <li>&lsquo;<em>Class</em>&rsquo;, variable used to distinguish between RPM&rsquo;s publications (&lsquo;RPM&rsquo;) and the citing documents (&lsquo;CITE&rsquo;)</li> <li>&lsquo;<em>Cluster</em>&rsquo;, provides the cluster membership number as discussed within the manuscript. This was established via modularity maximisation via the Leiden algorithm (Res 1; Q=0.67 | 25 clusters).</li> </ul> <p><strong>References</strong></p> <p>[1] Leng, R. I., Leng. G. (Under review). A career in numbers: a citation network analysis of the work of RP Millar and his contribution to GnRH research. <em>J. Neuroendocrinol</em></p> <p>All bibliographic data included in this study are derived originally from Clarivate&trade; (Web of Science&trade;) and downloaded in January 2024. &copy; Clarivate 2024. All rights reserved.&nbsp;</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

Fig. 2 in Correction of the holotype citations of three vascular plants at the herbarium of the National Institute of Biological Resources, Korea

Fig. 2. Holotype of Isoetes coreana Y.H. Chung &amp; H.K. Choi.

opencc-by-4.0Dec 2020View details →
zenodo36/100

Fig. 3 in Correction of the holotype citations of three vascular plants at the herbarium of the National Institute of Biological Resources, Korea

Fig. 3. Holotype of Huperzia jejuensis B.Y. Sun &amp; J. Lim.

opencc-by-4.0Dec 2020View details →
zenodo36/100

Methodology data of "A qualitative and quantitative citation analysis toward retracted articles: a case of study"

<p>This document contains the datasets and visualizations generated after&nbsp;the application of the&nbsp;methodology defined in&nbsp;our work: <em>&quot;A qualitative and quantitative citation analysis toward retracted articles: a case of study&quot;</em>. The methodology defines a&nbsp;citation analysis of&nbsp;the Wakefield et al. [1] retracted article from a quantitative and qualitative point of view. The data contained in this repository are&nbsp;based on the first two&nbsp;steps of the methodology. The first step of the methodology&nbsp;(i.e. &ldquo;Data gathering&rdquo;) builds&nbsp;an annotated dataset of the citing entities, this step is largely discussed also in [2]. The second step (i.e. &quot;Topic Modelling&quot;)&nbsp;runs a topic modeling analysis on the textual features contained in the dataset generated by&nbsp;the first step.&nbsp;</p> <p><strong>Note:</strong> the data are all contained inside the &quot;<strong><em>method_data.zip&quot;</em> </strong>file. You need to unzip the file to get access to all the files and directories listed below.</p> <p>&nbsp;</p> <p><strong>Data gathering</strong></p> <p>The data generated by this step are stored in&nbsp;<strong>&quot;<em>data/</em>&quot;</strong>:</p> <ol> <li><em><strong>&quot;cits_features.csv&quot;:&nbsp;</strong></em>a dataset containing&nbsp;all the entities (rows in the CSV) which have cited the&nbsp;Wakefield et al.&nbsp;retracted article, and a set of&nbsp;features characterizing each citing entity&nbsp;(columns in the CSV). The features included are:&nbsp;DOI (&quot;doi&quot;), year of publication (&quot;year&quot;), the title (&quot;title&quot;), the venue identifier (&quot;source_id&quot;), the title of the venue (&quot;source_title&quot;), yes/no value in case the entity is retracted as well (&quot;retracted&quot;), the subject area (&quot;area&quot;), the subject category (&quot;category&quot;), the sections of the in-text citations (&quot;intext_citation.section&quot;), the value of the reference pointer (&quot;intext_citation.pointer&quot;), the in-text citation function (&quot;intext_citation.intent&quot;), the in-text citation perceived sentiment (&quot;intext_citation.sentiment&quot;), and a yes/no value to denote whether the in-text citation context mentions the retraction of the cited entity&nbsp;&nbsp; &nbsp;(&quot;intext_citation.section.ret_mention&quot;).<br> <strong>Note: </strong>this dataset is licensed under a&nbsp;<a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.<br> &nbsp;</li> <li><em><strong>&quot;cits_text.csv&quot;: </strong>this dataset stores the abstract (&quot;abstract&quot;) and the in-text citations context (&quot;intext_citation.context&quot;) </em>for&nbsp;each citing entity identified using the DOI value (&quot;doi&quot;).<br> <strong>Note: </strong>the data keep their original&nbsp;license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> </ol> <p>&nbsp;</p> <p><strong>Topic modeling</strong><br> We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>&quot;topic_modeling/&quot;</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools, and creating a completely customizable visual workflow [3]. The topic modeling results for each textual feature are separated into two different folders, <em><strong>&quot;abstracts/&quot;</strong></em> for the abstracts, and <em><strong>&quot;intext_cit/&quot;</strong></em> for the in-text citation contexts. Both the directories contain the following directories/files: &nbsp; <strong>&nbsp;</strong></p> <ol> <li> <p><em><strong>&quot;mitao_workflows/&quot;</strong></em>: the workflows of MITAO. These are JSON files that could be reloaded in MITAO to reproduce the results following the same workflows.</p> </li> <li> <p><em><strong>&quot;corpus_and_dictionary/&quot;:&nbsp;</strong></em>it contains the dictionary and the vectorized corpus given as inputs for the&nbsp;LDA topic modeling.</p> </li> <li> <p><em><strong>&quot;coherence/coherence.csv&quot;:</strong></em>&nbsp;the coherence score of several&nbsp;topic models trained on a number of topics from 1 - 40.</p> </li> <li> <p><em><strong>&quot;datasets_and_views/&quot;: </strong></em>the datasets and visualizations generated using MITAO.&nbsp;&nbsp;</p> </li> </ol> <p>&nbsp;</p> <p><strong>References</strong></p> <ol> <li>Wakefield, A., Murch, S., Anthony, A., Linnell, J., Casson, D., Malik, M., Berelowitz, M., Dhillon, A., Thomson, M., Harvey, P., Valentine, A., Davies, S., &amp; Walker-Smith, J. (1998). RETRACTED: Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. <em>The Lancet</em>, <em>351</em>(9103), 637&ndash;641. <a href="https://doi.org/10.1016/S0140-6736(97)11096-0">https://doi.org/10.1016/S0140-6736(97)11096-0</a></li> <li> <p>Heibi, I., &amp; Peroni, S. (2020). A methodology for gathering and annotating the raw-data/characteristics of the documents citing a retracted article v1 (protocols.io.bdc4i2yw) [Data set]. In protocols.io. ZappyLab, Inc. <a href="https://doi.org/10.17504/protocols.io.bdc4i2yw">https://doi.org/10.17504/protocols.io.bdc4i2yw</a></p> </li> <li> <p>&nbsp;</p> Ferri, P., Heibi, I., Pareschi, L., &amp; Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135&ndash;149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a> <p>&nbsp;</p> </li> </ol>

opencc-zeroDec 2020View details →
zenodo36/100

Citation Tone of Lugang dialect in Chaonan District of Shantou City in Guangdong Province, China

<p>Citation Tone of Lugang dialect in Chaonan District of Shantou City in Guangdong Province, China&nbsp;</p>

opencc-by-4.0Nov 2018View details →
zenodo36/100

Dataset for Spence et al., "Availability of study protocols for randomized trials published in high-impact medical journals: cross-sectional analysis" (CITATION)

<p>Contains our extraction sheets (as SAS data files), code to calculate the values in the tables in our manuscript, and a supplemental file with additional notes on methods used in our study.</p>

opencc-by-4.0Aug 2018View details →
zenodo36/100

Comparing the Use of Research Resource Identifiers and Natural Language Processing for Citation of Databases, Software and Other Digital Artifacts

<p><strong>The Research Resource Identifier was introduced in biomedicine in 2014 to more precisely identify the reagents and tools used in published biomedical research and to track use of tools across the breadth of the biomedical literature. The current RRID specification covers key biological and digital resources. Authors are instructed to include an RRID after the first mention of any resource used. RRIDs are designed to be easy to find using &nbsp;a full text search search engine. </strong></p> <p><strong>The published data sets were used in our comparative study where comparing the output of our RRID curation workflow with the outputs of automated text mining systems that have been used to identify mentions of resources in the text of publications. All files in tab-separated format (tsv). </strong></p> <p><strong>Scibot.tsv: Records of the RRID curation workflow using SciBot. </strong></p> <p>Each record shows that a resource RRID was identified in paper PMID with curator tags (Tag1, Tag2, both optional)</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag1: Curator tags (optional)</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Tag2: Additional curator tags (optional)</p> <p><strong>rdwsorted.tsv: Records of the output from RDW, a text mining software. </strong></p> <p>RDW identifies mentions of research resources in papers. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p><strong>rridbyrdw05282019.tsv:&nbsp;Records of the output of the RRID-by-RDW in RDW. </strong></p> <p>RRID-by-RDW is a component in RDW that identifies mentions of research resources in papers by matching patterns of RRID specifications. Each record shows that a resource RRID was identified in paper PMID.</p> <p><strong>&nbsp;&nbsp;&nbsp; </strong>PMID: Pubmed ID</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; RRID: Research Resource Identifier</p> <p>&nbsp;&nbsp;&nbsp;&nbsp; Context: Snippet where the RRID was found</p> <p><strong>resource_metadata20190418.tsv: Metadata of RRIDs</strong></p> <p>This file contains metadata of resources and their RRIDs. See file header for column definitions.</p> <p><strong>RRIDCUR-definitions.tsv: Definitions of curator tags used in Scibot.tsv.</strong></p> <p><strong>&nbsp;&nbsp; </strong>tag: Tag name</p> <p>&nbsp;&nbsp;&nbsp; definition: Definition of the tag</p>

openbsd-3-clause-clearJun 2019View details →
zenodo36/100

CSV Citation and Metadata Sources from Dimensions on DOI: 10.2196/14434

<p>CSV Exports from Dimensions derived from&nbsp;<a href="https://doi.org/10.2196/14434">https://doi.org/10.2196/14434</a>&nbsp; - sources need to be evaluated&nbsp;</p>

opencc-by-4.0Dec 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record