Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
214
datasets available to search
ShareScore release 0.7.1
Dataset results
214 results for “citation”
Dataset for Citation Network Analysis on Quality Cues for Meat Purchases
<p>Data are retrieved from two databases permissible to the Vosviewer® software: Scopus and Web of Science. Three sets of Boolean search strings are used to find articles in the databases:</p> <p>(i) [<em>(“certification” AND “meat”) OR (“certification” AND “dairy”)</em>]</p> <p>(ii) [<em>(“credence” AND meat) OR (“experience” AND meat)</em>]<em> </em></p> <p><em>(iii) </em>[<em>(“meat” AND “intrinsic”) OR (“meat” AND “extrinsic”) OR (“meat” AND “quality cue”)</em>].</p> <p>The searches targeted the title, abstract and keywords sections for the Scopus database, and the topic section for the Web of Science database. The search was limited to peer-reviewed articles published in the English language; no timeline restrictions were set for the initial search. </p>
Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"
<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>
Mediocres en la academia: el área de conocimiento de Biblioteconomía y Documentación del departamento de Historia de la Ciencia y Documentación (HCyD) de la Universitat de València como caso de estudio. Anexo II. Identificación y descripción de los plagios identificados en la tesis doctoral titulada Análisis de los artículos originales publicados en revistas específicas sobre drogodependencias incluidas en el Journal Citation Reports (2002-2006).
<p>Annex II. Identification and description of the plagiarism identified in the doctoral thesis entitled Analysis of original articles published in specific journals on drug dependence included in the Journal Citation Reports (2002-2006).</p>
Missing Citations in COCI: Publishers Analytics Result
<p>This dataset contains a JSON file containing the results retrieved through the<a href="http://doi.org/10.5281/zenodo.4735621"> software developed by the authors</a>. We opted for JSON file format to store the obtained data, since this format allows the storage of heterogeneous information in a complex and structured way. The four main structures stored in the present file are: </p> <p>1) "publishers", a list of dictionaries representing each publisher encountered;</p> <p>2) "citations", a dictionary containing two lists, the one storing the validated citational data and the other storing the still invalid citational data. Each processed citation is represented as a dictionary. </p> <p>3) "total_num_of_valid_citations", whose value is the number of citational data that could be validated throughout the process implemented by our software.</p> <p>4) "external_data_for_unrecognized_prefixes": a dictionary of dictionaries representing the publishers we didn't find on Crossref, but that were identified through other online services. </p> <p>We used as input material open data from the dataset “<a href="http://doi.org/10.5281/zenodo.4625300">Citations to invalid DOI-identified entities obtained from processing DOI-to-DOI citations to add in COCI</a>”. </p>
Quality and Trackability of Author Indicated Software Citations (NEST Case Study)
<p>This dataset contains a randomly selected minimal sample of 471 NEST software citations, which were provided by authors and published on the NEST software page as a publication list. The software citations were analyzed on their quality and trackability.</p> <p>The data was collected in 2020 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>
unarXive: All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation Network (open subset)
<h2><strong>Description</strong></h2><p>unarXive is a scholarly data set containing publications' structured full-text, annotated in-text citations, linked non-text content (mathematical notation, figure/table captions) and a citation network.</p><p>The data is generated from all LaTeX sources on <a href="https://arxiv.org/">arXiv</a> and therefore of higher quality than data generated from PDF files.</p><p>Typical uses are</p><ul><li>Training of ML models (citation recommendation, summarization, LLMs)</li><li>Citation context analysis</li><li>Bibliographic analyses</li></ul><h2><strong>Access</strong></h2><p>┏━━━━━━━━━━━━━━━━━━━━━━━━━━┓<br>┃ <a href="https://github.com/IllDepence/unarXive/raw/master/doc/unarXive_data_sample.tar.gz"><strong>D O W N L O A D S A M P L E</strong></a> ┃<br>┗━━━━━━━━━━━━━━━━━━━━━━━━━━┛</p><p>Regarding the full data set, please note the following:</p><blockquote><p><strong>Note</strong>: this Zenodo record is the "open subset" of unarXive, which contains all permissively licensed papers from arXiv.org. You can find the <a href="https://doi.org/10.5281/zenodo.7752754">full version here</a>.</p></blockquote><p>The code used for generating the data set is <a href="https://github.com/IllDepence/unarXive">publicly available</a>.</p>
Scholix dump of the OpenAIRE inferred citations
<p>This dataset contains the set of citations extracted for the large by the OpenAIRE Information Inference Service (IIS). It consists of tar archives, each containing gzip-compressed files, including new-line delimited JSON records in Scholix format.</p> <p>The dataset counts 36.430.057 citations, from citing scientific products.</p> <p>The type of PIDs among the citing and the cited products include:</p> <ul> <li>DOI</li> <li>PMC</li> <li>PMID</li> <li>ArXiv</li> <li>Handle</li> </ul>
Supplemental data for: Visualization of rank-citation curves for fast detection of possible manipulations with the h-index of the university
<p>This dataset consists of papers of universities in the top 30 Scopus Ranking of Ukrainian Universities (May 2023). The data was obtained from Scopus using the search query "AF-ID (“university name”) AND PUBYEAR < 2023 AND PUBYEAR > 2002". Rank-citation curves were also generated for the publications of each university. In this analysis, the rank of publications was plotted along the horizontal axis, while the corresponding citation counts were depicted on the left axis. All types of documents were included in the dataset.</p>
Factors Associated with Scientific Production Citations in Dentistry: Zero-inflated Negative Binomial Regression and Hurdle Modelling
<p><strong>Abstracto:</strong> La literatura científica mundial en odontología ha mostrado importantes avances en este campo, con importantes contribuciones que van desde el análisis de los aspectos epidemiológicos básicos de la prevención hasta resultados especializados en el campo de los tratamientos dentales. La presente investigación tiene como objetivo analizar el estado actual de la literatura científica sobre odontología alojada en la base de datos Web of Science. La metodología incluye dos fases en el análisis de artículos y revisiones indexadas en todas las áreas temáticas. Durante la primera fase, se analizan las siguientes variables: la producción científica por parte del editor, la evolución de la producción científica publicada por los editores, los factores asociados al impacto de la producción científica y la modelización del impacto de la producción científica en odontología. Durante la segunda fase, se analizan asociaciones, evoluciones y tendencias en el uso de palabras clave principales en la literatura científica en odontología. En conclusión, el estudio muestra que los temas más estudiados incluyen la asociación de la educación dental y el plan de estudios, la asociación de la odontología pediátrica con la salud oral y el cuidado dental. Los hallazgos muestran que también destacan temas enfatizados más recientemente, como la odontología basada en la evidencia, la pandemia, el control de infecciones y la endodoncia, así como la necesidad de futuras investigaciones para ampliar el conocimiento actual basado en temas emergentes en la literatura científica sobre odontología.</p>
The data of "A comparison of citation-based clustering and topic modeling for science mapping"
<p>These files consist of the data used in "A comparison of citation-based clustering and topic modeling for science mapping". </p> <p> </p>
Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns
<p>Questionable research practices (QRPs) have been the focus of the scientific community amid greater scrutiny and evidence highlighting issues with replicability across many fields of science. To capture the most impactful publications and the main thematic domains in the literature on QRPs, this study uses a document co-citation analysis. The analysis was conducted on a sample of 341 documents that covered the past 50 years of research in QRPs. Nine major thematic clusters emerged. Statistical reporting and statistical power emerged as key areas of research, where systemic-level factors in how research is conducted are consistently raised as the precipitating factors for QRPs. There is also an encouraging shift in the focus of research into open science practices designed to address engagement in QRPs. Such a shift is indicative of the growing momentum of the open science movement, and more research can be conducted on how these practices are employed on the ground and how their uptake by researchers can be further promoted. However, the results suggest that, while pre-registration and registered reports receive the most research interest, less attention has been paid to other open science practices (e.g., data and methods sharing).</p>
Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns
Open the record for dataset details and reuse information.
Research on the benefits of nature to people: How much overlap is there in citations and terms for ‘nature’ across disciplines?
Open the record for dataset details and reuse information.
High-frequency location data show that race affects citations and fines for speeding
Open the record for dataset details and reuse information.
Agrivoltaic grazing systems for a sustainable future: Citation database for a multi-disciplinary review & gap analysis
Open the record for dataset details and reuse information.
Labeled data for citation field extraction
Open the record for dataset details and reuse information.
Forecasting the publication and citation outcomes of Covid-19 preprints
Open the record for dataset details and reuse information.
The disruption index suffers from citation inflation and is confounded by shifts in scholarly citation practice: synthetic citation networks for bibliometric null models
Open the record for dataset details and reuse information.
CORD-19_ scite_citation_tallies+contexts
<pre>Update: As of March 27, 2020 we have now analyzed 31,527 distinct sources (articles and preprints) from the most recent CORD-19 data (<a href="https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/version/4">https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/version/4</a>. We're releasing citation tallies for these sources (covid-source-tallies 32720.csv). We're also releasing citation statements and classifications from these documents for open articles, which includes 1,682,216 out of the total 1,779,024 extracted. On March 20, 2020 we have analyzed 20,268 out of the 21,792 DOIS available from the <a href="https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge">CORD-19 data set</a>. Of these documents we found citations citing 16,775 of them, and the classifications for these citations are included in covid-source-tallies.csv. covid-citations.csv includes all citations we have from all of the ~20k documents we have processed. This file is truncated to make sure it only includes openly available documents. The tallies however are not limited by this, and it is the full set relating to all source documents scite has processed.</pre>
Data citation for a forward stratigraphic-based porosity and permeability model developed for the Volve field, Norway.
<p>The data, models, and script presented here are those used for developing a forward stratigraphic simulation. The data include: 24 suits of well logs, seismic data, forward stratigraphic simulation scenarios of the shallow marine depositional setting, synthetic wells derived from the stratigraphic model, and 3-D reservoir models in Eclipse and RMS formats. In addition, a short script from the property calculator tool in Petrel, which is was used to classify lithofacies-associations in the stratigraphic model is also provided. The Petrel software license and code used in GPM software to undertake these forward stratigraphic simulations cannot be provided, because Schlumberger, who are the developers of the software do not allow its code to be shared in any publication.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.