Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76
datasets available to search
ShareScore release 0.7.1
Dataset results
76 results for “data citation”
Citation Location, Elements and Purpose of ICPSR Research Data Citations
<p>The dataset contains an analysis of a randomly chosen sample of 1,073 publications that cite ICPSR research datasets. The collected research data (re)use indications were analyzed according to their location in the full-text, their metadata elements, and their citation purpose.</p> <p>The data was collected and analyzed in 2020 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
Use and sharing of raw data in the Journal Citation Reports' Emergency Medicine Category: Metrics and Journals including supplementary material classification sorted by quartile of the JCR emergency medicine category.
<p>Raw data belonged to the study of use and sharing of raw research data in the Journal Citation Reports' Emergency Medicine Category.</p>
Data for "Measuring Back: Bibliodiversity and the Journal Impact Factor brand. A Case study of IF-journals included in the 2021 Journal Citations Report."
<p>This is the open data for the preprint "Measuring Back: Bibliodiversity and the Journal Impact Factor brand. A Case study of IF-journals included in the 2021 Journal Citations Report."</p>
Bibliographic data on datasets affiliated to Poznan University of Technology and indexed in Data Citation Index (retrieved by Web of Science service in January 2023))
<p>The file contains the number of datasets published by the researchers affiliated to Poznan University of Technology and indexed in Data Citation Index provided by Web of Science (database updated 10.01.2023). The Search was performed using the name of institution in the 'Affiliation' field. Dataset contains two files in two diffrent formats: plain text and xls.</p>
Highly cited tropical medicine articles in the early COVID 19 pandemic. Original Excel data on citation and subjects
<p><strong>ORIGINAL DATA SET FOR STUDY OF PUBLICATION AND CITATION TRENDS, TROPICAL MEDICINE, EARLY COVID 19 PANDEMIC. Background: </strong>An adequate response to health needs includes the identification of research patterns about the large number of people living in the tropics and subjected to tropical diseases. Studies have shown that research does not always match the real needs of those populations, and that citation reflects mostly the amount of money behind particular publications. Here we test the hypothesis that research from richer institutions is published in better-indexed journals, and thus has greater citation rates.</p> <p><strong>Methods:</strong> The data in this study was extracted from the Science Citation Index Expanded database; the 2020 journal Impact Factor (<em>IF</em><sub>2020</sub>) was updated to 30 June 2021. We considered places, subjects, institutions and journals.</p> <p><strong>Results:</strong> We identified 1 041 highly cited articles with 100 citations or more in the category of tropical medicine. About a decade is needed for an article to reach peak citation. Only two Covid-19 related were highly cited in the last three years. Most cited articles were published by the journals <em>Memorias Do Instituto Oswaldo Cruz</em> (Brazil), <em>Acta Tropica</em> (Switzerland), and <em>PLoS Neglected Tropical Diseases</em> (USA). The USA dominated five of the six publication indicators. International collaboration articles had more citations than single-country articles. The UK, South Africa, and Switzerland had high citation rates, as did the London School of Hygiene and Tropical Medicine in the UK, the Centers for Disease Control and Prevention in the USA, and the WHO in Switzerland.</p> <p><strong>Conclusions:</strong> About ten years of accumulated citations are needed to get 100 citations or more as highly cited articles in the Web of Science category of tropical medicine. Six publication and citation indicators, including authors’ publication potential and characteristics evaluated by <em>Y</em>-index, indicate that the currently available indexing system places tropical researchers at a disadvantage against their colleagues in temperate countries, and suggest that, to progress towards better control of tropical diseases, international collaboration should increase, and other tropical countries should follow the example of Brazil, which provides significant financing to its scientific community.Julián Monge-Nájera<sup>1</sup>, and Yuh-Shan Ho<sup>2</sup>*</p> <p><sup>1</sup>Laboratorio de Ecología Urbana, Vicerrectoría de Investigación, Universidad Estatal a Distancia, 2050 San José, Costa Rica; <a href="mailto:julianmonge@gmail.com"><em>julianmonge@gmail.com</em></a> (https://orcid.org/0000-0001-7764-2966)</p> <p>*Corresponding author: Trend Research Centre, Asia University, No. 500 Lioufeng Road, Wufeng, Taichung 41354, Taiwan; <a href="mailto:ysho@asia.edu.tw">ysho@asia.edu.tw</a> (<em>https://orcid.org/0000-0002-2557-8736</em>)</p>
Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals - DATA PRODUCED
<p>This zipped folders contain all the data produced for the research "Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals": the results datasets (dataset_map_disciplines, dataset_no_SSH, dataset_SSH, erih_meta_with_disciplines and erih_meta_without_disciplines).</p> <ul> <li> <p><strong>dataset_map_disciplines.zip </strong>contains CSV files with four columns ("id", "citing", "cited", "disciplines") giving information about publications stored in OpenCitations META (version 3 released on February 2023) and part of SSH journals, according to ERIH PLUS (version downloaded on 2023-04-27), specifying the disciplines associated to them and a boolean value stating if they cite or are cited, according to the OpenCitations COCI dataset (version 19 released on January 2023).</p> </li> <li> <p><strong>dataset_no_SSH.zip </strong>and <strong>dataset_SSH.zip</strong> contain CSV files with the same structure. Each dataset has four columns: "citing", "is_citing_SSH", "cited", and "is_cited_SSH". ”Citing” and “cited” columns are filled with DOIs of publications stored in OpenCitations META that according to OpenCitations COCI are involved in a citation. The "is_citing_SSH" and "is_cited_SSH" columns contain boolean values: "True" if the corresponding publication is associated with a SSH (Social Sciences and Humanities) discipline, according to ERIH PLUS, and "False" otherwise. The two datasets are built starting from the two different subsets obtained as a result of the union between OpenCitations META and ERIH PLUS: dataset_SSH comes from erih_meta_with_disciplines and dataset_no_SSH from <strong>erih_meta_without_disciplines. </strong>dataset_no_SSH comes from <strong>erih_meta_with_disciplines.zip</strong> and erih_meta_without_disciplines.zip, as explained before, contain CSV files originating from ERIH PLUS and META. erih_meta_without_disciplines has just one column “id” and contains the DOIs of all the publications in META that do not have any discipline associated, that is, have not been published on a SSH journal, while erih_meta_with_disciplines derives from all the publications in META that have at least one linked discipline and has two columns: “id” and “erih_disciplines”, containing a string with all the disciplines linked to that publication like "History, Interdisciplinary research in the Humanities, Interdisciplinary research in the Social Sciences, Sociology".</p> </li> </ul> <p>Software: https://doi.org/10.5281/zenodo.8326023</p> <p>Data preprocessed: https://doi.org/10.5281/zenodo.7973159</p> <p>Article: https://zenodo.org/record/8326044</p> <p>DMP: https://zenodo.org/record/8324973</p> <p>Protocol: https://doi.org/10.17504/protocols.io.n92ldpeenl5b/v5</p>
Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals - DATA PREPROCESSED
<p>This zipped folders contain all the data preprocessed for the research "Uncovering the Citation Landscape: Exploring OpenCitations COCI, OpenCitations Meta, and ERIH-PLUS in Social Sciences and Humanities Journals": the cleaned datasets (coci_preprocessed, meta_preprocessed, erih_preprocessed and erih_meta).</p> <ul> <li> <p><strong>coci_preprocessed.zip</strong>: this archive contains CSVs with two columns “citing” and “cited”, giving information about publications involved in citations according to the OpenCitations COCI dataset (version 19 released on January 2023), and that are entirely contained in OpenCitations META (version 3 released on February 2023). This means that the citations which have either the citing or the cited entity (or both) not contained in META are excluded from coci_preprocessed dataset.</p> </li> <li> <p><strong>meta_preprocessed.zip</strong>: all the original columns of OpenCitations META are maintained in this dataset, so the CSVs have the columns: “id”, “title”, “author”, “issue”, “volume”, “venue”, “page”, “pub_date”, “type”, “publisher” and “editor”. The only difference with the original dataset is that meta_preprocessed in the columns “id” and “venue” has respectively just the DOIs and the ISSNs, without all the other identifiers specified for each entity in META.</p> </li> <li> <p><strong>erih_preprocessed.zip</strong>: it contains a CSV file with two columns "venue_id" and "ERIH_disciplines". "venue_id" is the union of the original columns "Online ISSN" and "Print ISSN" of ERIH_PLUS (version downloaded on 2023-04-27).</p> </li> <li> <p><strong>erih_meta.zip</strong>: it contains CSV files obtained from the union of meta_preprocessed and erih_preprocessed, they have all the columns of meta_preprocessed plus a new column “erih_disciplines” containing all the disciplines linked to a venue (identified by an ISSN).</p> </li> </ul> <p> </p> <p>Software: https://doi.org/10.5281/zenodo.8326023</p> <p>Data produced: https://doi.org/10.5281/zenodo.7974816</p> <p>Article: https://zenodo.org/record/8326044</p> <p>DMP: https://zenodo.org/record/8324973</p> <p>Protocol: https://doi.org/10.17504/protocols.io.n92ldpeenl5b/v5</p>
Don't mention it: challenges to using software mentions to investigate citation and discoverability - Data and Notebooks
<p>This deposit contains the data and Jupyter notebooks used for sampling and annotation analysis of our submission to the PeerJ Computer Science special issue <em>Software Citation, Indexing, and Discoverability </em>(<a href="https://peerj.com/collections/84-software">https://peerj.com/collections/84-software</a>):</p> <blockquote> <p>Stephan Druskat, Neil P. Chue Hong, Sammie Buzzard, Olexandr Konovalov, and Patrick Kornek. Don’t mention it: challenges to using software mentions to investigate citation and discoverability.</p> </blockquote> <p>See the README in this deposit.</p> <p>The contents of this deposit can also be browsed at <a href="https://github.com/softwaresaved/habeas-corpus/tree/main/replication-package">https://github.com/softwaresaved/habeas-corpus/tree/main/replication-package</a>.</p>
Reference Manager Data Citation Analysis
<p>DESCRIPTION:</p> <p>This package contains data used to analyze citation metadata completeness and correctness for several common reference managers used in scholarly research and several common repositories in the Earth, space, and environmental sciences.</p> <p>METHODS:</p> <p>Metadata fields for import and export methods and for 8 metadata fields (authors/creators, publisher, DOI, dataset title, version, access date, publication date, and resource type) were collected from reference managers via all import methods available (app or wizard and plugin) during summer 2024 from most recent software versions of all. To encode data, citation information for each dataset as imported by Reference Manager was compared to that registered for the DOI with DataCite. Correct metadata for each of 8 fields for both import and export was encoded as 0, incorrect as 1, and missing as '' or nan. See publication and software package for more information.</p> <p>FILES:</p> <p>FOLDER 'coded-data' contains files that include information (DOIs) about the data examined in this study, preserved copies of exported data citations used in the data interpretation and processing, and the processed data itself encoded in columns.</p> <p>FOLDER 'datacite-metadata-profiles' includes the raw metadata from each dataset DOI at the time of analysis, included for reproducibility purposes. </p> <p>FOLDER 'bibtex-files' includes the downloaded .bib files, where available, for each dataset DOI examined.</p> <p>See README file for more information.</p>
Labeled data for citation field extraction
<p>Citations are an important part of scientific papers, and the proper handling of them is indispensable for the science of science. Citation field extraction is the task of parsing citations: given a citation string, extract authors, title, venue, doi etc. Since the number of citations is counted by hundreds millions, efficient computer based methods for this task are very important.</p> <p>The development of machine learning methods for citation field extraction requires ground truth: a large corpus of labeled citations. This dataset provides a very large (41M) corpus of labeled data obtained by the reverse process: we took structured citation lists and used BibTeX to generate labeled citation strings.</p>
Methodology data of "A qualitative and quantitative citation analysis toward retracted articles: a case of study"
<p>This document contains the datasets and visualizations generated after the application of the methodology defined in our work: <em>"A qualitative and quantitative citation analysis toward retracted articles: a case of study"</em>. The methodology defines a citation analysis of the Wakefield et al. [1] retracted article from a quantitative and qualitative point of view. The data contained in this repository are based on the first two steps of the methodology. The first step of the methodology (i.e. “Data gathering”) builds an annotated dataset of the citing entities, this step is largely discussed also in [2]. The second step (i.e. "Topic Modelling") runs a topic modeling analysis on the textual features contained in the dataset generated by the first step. </p> <p><strong>Note:</strong> the data are all contained inside the "<strong><em>method_data.zip"</em> </strong>file. You need to unzip the file to get access to all the files and directories listed below.</p> <p> </p> <p><strong>Data gathering</strong></p> <p>The data generated by this step are stored in <strong>"<em>data/</em>"</strong>:</p> <ol> <li><em><strong>"cits_features.csv": </strong></em>a dataset containing all the entities (rows in the CSV) which have cited the Wakefield et al. retracted article, and a set of features characterizing each citing entity (columns in the CSV). The features included are: DOI ("doi"), year of publication ("year"), the title ("title"), the venue identifier ("source_id"), the title of the venue ("source_title"), yes/no value in case the entity is retracted as well ("retracted"), the subject area ("area"), the subject category ("category"), the sections of the in-text citations ("intext_citation.section"), the value of the reference pointer ("intext_citation.pointer"), the in-text citation function ("intext_citation.intent"), the in-text citation perceived sentiment ("intext_citation.sentiment"), and a yes/no value to denote whether the in-text citation context mentions the retraction of the cited entity ("intext_citation.section.ret_mention").<br> <strong>Note: </strong>this dataset is licensed under a <a href="https://creativecommons.org/publicdomain/zero/1.0/legalcode">Creative Commons public domain dedication (CC0)</a>.<br> </li> <li><em><strong>"cits_text.csv": </strong>this dataset stores the abstract ("abstract") and the in-text citations context ("intext_citation.context") </em>for each citing entity identified using the DOI value ("doi").<br> <strong>Note: </strong>the data keep their original license (the one provided by their publisher). This dataset is provided in order to favor the reproducibility of the results obtained in our work.</li> </ol> <p> </p> <p><strong>Topic modeling</strong><br> We run a topic modeling analysis on the textual features gathered (i.e. abstracts and citation contexts). The results are stored inside the <em><strong>"topic_modeling/"</strong></em> directory. The topic modeling has been done using MITAO, a tool for mashing up automatic text analysis tools, and creating a completely customizable visual workflow [3]. The topic modeling results for each textual feature are separated into two different folders, <em><strong>"abstracts/"</strong></em> for the abstracts, and <em><strong>"intext_cit/"</strong></em> for the in-text citation contexts. Both the directories contain the following directories/files: <strong> </strong></p> <ol> <li> <p><em><strong>"mitao_workflows/"</strong></em>: the workflows of MITAO. These are JSON files that could be reloaded in MITAO to reproduce the results following the same workflows.</p> </li> <li> <p><em><strong>"corpus_and_dictionary/": </strong></em>it contains the dictionary and the vectorized corpus given as inputs for the LDA topic modeling.</p> </li> <li> <p><em><strong>"coherence/coherence.csv":</strong></em> the coherence score of several topic models trained on a number of topics from 1 - 40.</p> </li> <li> <p><em><strong>"datasets_and_views/": </strong></em>the datasets and visualizations generated using MITAO. </p> </li> </ol> <p> </p> <p><strong>References</strong></p> <ol> <li>Wakefield, A., Murch, S., Anthony, A., Linnell, J., Casson, D., Malik, M., Berelowitz, M., Dhillon, A., Thomson, M., Harvey, P., Valentine, A., Davies, S., & Walker-Smith, J. (1998). RETRACTED: Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. <em>The Lancet</em>, <em>351</em>(9103), 637–641. <a href="https://doi.org/10.1016/S0140-6736(97)11096-0">https://doi.org/10.1016/S0140-6736(97)11096-0</a></li> <li> <p>Heibi, I., & Peroni, S. (2020). A methodology for gathering and annotating the raw-data/characteristics of the documents citing a retracted article v1 (protocols.io.bdc4i2yw) [Data set]. In protocols.io. ZappyLab, Inc. <a href="https://doi.org/10.17504/protocols.io.bdc4i2yw">https://doi.org/10.17504/protocols.io.bdc4i2yw</a></p> </li> <li> <p> </p> Ferri, P., Heibi, I., Pareschi, L., & Peroni, S. (2020). MITAO: A User Friendly and Modular Software for Topic Modelling [JD]. PuntOorg International Journal, 5(2), 135–149. <a href="https://doi.org/10.19245/25.05.pij.5.2.3">https://doi.org/10.19245/25.05.pij.5.2.3</a> <p> </p> </li> </ol>
Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"
<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>
Supplemental data for: Visualization of rank-citation curves for fast detection of possible manipulations with the h-index of the university
<p>This dataset consists of papers of universities in the top 30 Scopus Ranking of Ukrainian Universities (May 2023). The data was obtained from Scopus using the search query "AF-ID (“university name”) AND PUBYEAR < 2023 AND PUBYEAR > 2002". Rank-citation curves were also generated for the publications of each university. In this analysis, the rank of publications was plotted along the horizontal axis, while the corresponding citation counts were depicted on the left axis. All types of documents were included in the dataset.</p>
The data of "A comparison of citation-based clustering and topic modeling for science mapping"
<p>These files consist of the data used in "A comparison of citation-based clustering and topic modeling for science mapping". </p> <p> </p>
High-frequency location data show that race affects citations and fines for speeding
Open the record for dataset details and reuse information.
Labeled data for citation field extraction
Open the record for dataset details and reuse information.
Data citation for a forward stratigraphic-based porosity and permeability model developed for the Volve field, Norway.
<p>The data, models, and script presented here are those used for developing a forward stratigraphic simulation. The data include: 24 suits of well logs, seismic data, forward stratigraphic simulation scenarios of the shallow marine depositional setting, synthetic wells derived from the stratigraphic model, and 3-D reservoir models in Eclipse and RMS formats. In addition, a short script from the property calculator tool in Petrel, which is was used to classify lithofacies-associations in the stratigraphic model is also provided. The Petrel software license and code used in GPM software to undertake these forward stratigraphic simulations cannot be provided, because Schlumberger, who are the developers of the software do not allow its code to be shared in any publication.</p>
Patent citation data for USPTO utility patents granted between 1976-2015 and for patents belonging to 30 technology domains
<p>These two data file contains information on patent citations for USPTO utility patents granted between 1976 and 2015 and for patents that have been classified in 30 specific technology domains.</p> <p>The file 'CITATION_INFO_no_neg_citlag.csv' is generated combining raw data <a href="https://www.patentsview.org/download/">freely dowloadable from patentsview.org</a> from which citations where the filing year of the citing patent is younger than the filing year of the cited one have been removed.</p> <p>The file 'CITATIONS_DOMAINS.csv' is a sample of the previous file that only includes citations made by patents belonging to one of 30 domains defined in the paper '<a href="https://www.sciencedirect.com/science/article/pii/S0040162520309264#ecom0001">Estimating technology performance improvement rates by mining patent data</a>' by Giorgio Triulzi, Jeff Alstott and Chris Magee.</p> <p>These two files complement another dataset <a href="http://dx.doi.org/10.17632/f4fj887y67.1">published on Mendeley Data</a>. The two datasets can be used, together with the code <a href="https://github.com/GiorgioTriulzi/TechnologyPerformanceImprovementEstimates">published on GitHub</a>, to replicate the main results from the paper.</p>
Data from: An origin of citations: Darwin's collaborators and their contributions to the Origin of Species
<p>Since the first edition of the Origin of Species (1859), Charles Darwin apologizes for not correctly referencing all the works cited in his magnum opus. More than 150 years later we catalogued these citations and analysed the resultant data. Looking for a complete selection of collaborators, a flexibilization of the term "citation" was necessary, and we define it as any reference made to a third party, independently of its form or function. Following the same idea, the last edition of the Origin, originally published in 1872 and reprinted with minor additions and corrections in 1876, was chosen for the research because it represents the end of a long debate between Darwin and his peers, naturally being the edition with the most number of citations and collaborators. Through a diverse theoric approach we hope to present a new perspective for the study of the Origin of Species: a bibliographic approach gives us the tools needed to understand the history of the book as a physical and cultural object; bibliometrics provides a theory of citations as well as a quantitative analysis; lastly, the Science Studies highlight the profound social aspects of science in the making. The analysis resulted in 639 citations to 298 collaborators, although these results are only the tip of the iceberg of all the gathered data's potential.</p>
Data from: A stochastic generative model for citation networks among academic papers
<p>We propose a stochastic generative model to represent a directed graph constructed by citations among academic papers, where nodes and directed edges represent papers with discrete publication time and citations respectively. The proposed model assumes that a citation between two papers occurs with a probability based on the type of the citing paper, the importance of cited paper, and the difference between their publication times, like the existing models. We consider the out-degrees of citing paper as its type, because, for example, survey paper cites many papers. We approximate the importance of a cited paper by its in-degrees. In our model, we adopt three functions: a logistic function for illustrating the numbers of papers published in discrete time, an inverse Gaussian probability distribution function to express the aging effect based on the difference between publication times, and an exponential distribution (or a generalized Pareto distribution) for describing the out-degree distribution. We consider that our model is a more reasonable and appropriate stochastic model than other existing models and can perform complete simulations without using original data. In this paper, we first use the Web of Science database and see the features used in our model. By using the proposed model, we can generate simulated graphs and demonstrate that they are similar to the original data concerning the in- and out-degree distributions, and node triangle participation. In addition, we analyze two other citation networks derived from physics papers in the arXiv database and verify the effectiveness of the model.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.