Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

214

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

214 results for “citation”

Learn how ShareScore rates datasets ↗
zenodo36/100

Dataset for Citation Network Analysis on Quality Cues for Meat Purchases

<p>Data are retrieved from two databases permissible to the Vosviewer&reg; software: Scopus and Web of Science. Three sets of Boolean search strings are used to find articles in the databases:</p> <p>(i) [<em>(&ldquo;certification&rdquo; AND &ldquo;meat&rdquo;) OR (&ldquo;certification&rdquo; AND &ldquo;dairy&rdquo;)</em>]</p> <p>(ii) [<em>(&ldquo;credence&rdquo;&nbsp;AND&nbsp;meat) OR&nbsp;(&ldquo;experience&rdquo;&nbsp;AND meat)</em>]<em> </em></p> <p><em>(iii) </em>[<em>(&ldquo;meat&rdquo; AND &ldquo;intrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;extrinsic&rdquo;) OR (&ldquo;meat&rdquo; AND &ldquo;quality cue&rdquo;)</em>].</p> <p>The searches targeted the title, abstract and keywords sections for the Scopus database, and the topic section for the Web of Science database. The search was limited to peer-reviewed articles published in the English language; no timeline restrictions were set for the initial search.&nbsp;</p>

opencc-by-4.0Nov 2019View details →
zenodo36/100

Dataset for "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations"

<p>This is the dataset for the paper "Are data papers cited as research data? Preliminary analysis on interdisciplinary data paper citations" submitted to iConference 2025.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Mediocres en la academia: el área de conocimiento de Biblioteconomía y Documentación del departamento de Historia de la Ciencia y Documentación (HCyD) de la Universitat de València como caso de estudio. Anexo II. Identificación y descripción de los plagios identificados en la tesis doctoral titulada Análisis de los artículos originales publicados en revistas específicas sobre drogodependencias incluidas en el Journal Citation Reports (2002-2006).

<p>Annex II. Identification and description of the plagiarism identified in the doctoral thesis entitled Analysis of original articles published in specific journals on drug dependence included in the Journal Citation Reports (2002-2006).</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Missing Citations in COCI: Publishers Analytics Result

<p>This dataset contains a JSON&nbsp;file containing the results retrieved through the<a href="http://doi.org/10.5281/zenodo.4735621"> software developed by the authors</a>. We opted for JSON file format to store the obtained data, since this format allows the storage of heterogeneous information in a complex and structured way. The four main structures stored in the present file are:&nbsp;</p> <p>1) &quot;publishers&quot;, a list of dictionaries representing each publisher encountered;</p> <p>2) &quot;citations&quot;, a dictionary containing two lists, the one storing the validated citational data and the other storing the still invalid citational data. Each processed citation&nbsp;is represented as a dictionary.&nbsp;</p> <p>3) &quot;total_num_of_valid_citations&quot;, whose value is the number of citational data that could be validated throughout the process implemented by our software.</p> <p>4)&nbsp;&quot;external_data_for_unrecognized_prefixes&quot;: a dictionary of dictionaries representing the publishers we didn&#39;t find on Crossref, but that were identified through other online services.&nbsp;</p> <p>We used as input material open data from the dataset &ldquo;<a href="http://doi.org/10.5281/zenodo.4625300">Citations to invalid DOI-identified entities obtained from processing DOI-to-DOI citations to add in COCI</a>&rdquo;.&nbsp;&nbsp;</p>

openisc-licenseMay 2021View details →
zenodo36/100

Quality and Trackability of Author Indicated Software Citations (NEST Case Study)

<p>This dataset contains a randomly selected minimal sample of 471 NEST software citations, which were provided by authors and published on the NEST software page as a publication list. The software citations were analyzed on their quality and trackability.</p> <p>The data was collected in 2020 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

unarXive: All arXiv Publications Pre-Processed for NLP, Including Structured Full-Text and Citation Network (open subset)

<h2><strong>Description</strong></h2><p>unarXive is a scholarly data set containing publications' structured full-text, annotated in-text citations, linked non-text content (mathematical notation, figure/table captions) and a citation network.</p><p>The data is generated from all LaTeX sources on <a href="https://arxiv.org/">arXiv</a> and therefore of higher quality than data generated from PDF files.</p><p>Typical uses are</p><ul><li>Training of ML models (citation recommendation, summarization, LLMs)</li><li>Citation context analysis</li><li>Bibliographic analyses</li></ul><h2><strong>Access</strong></h2><p>┏━━━━━━━━━━━━━━━━━━━━━━━━━━┓<br>┃ &nbsp;<a href="https://github.com/IllDepence/unarXive/raw/master/doc/unarXive_data_sample.tar.gz"><strong>D O W N L O A D &nbsp; S A M P L E</strong></a> &nbsp; ┃<br>┗━━━━━━━━━━━━━━━━━━━━━━━━━━┛</p><p>Regarding the full data set, please note the following:</p><blockquote><p><strong>Note</strong>: this Zenodo record is the "open subset" of unarXive, which contains all permissively licensed papers from arXiv.org. You can find the <a href="https://doi.org/10.5281/zenodo.7752754">full version here</a>.</p></blockquote><p>The code used for generating the data set is <a href="https://github.com/IllDepence/unarXive">publicly available</a>.</p>

opencc-by-sa-4.0Mar 2023View details →
zenodo36/100

Scholix dump of the OpenAIRE inferred citations

<p>This dataset contains the set of citations extracted for the large by the OpenAIRE Information Inference Service (IIS). It consists of tar archives, each containing gzip-compressed files, including new-line delimited JSON records in Scholix format.</p> <p>The dataset counts 36.430.057 citations, from&nbsp;citing scientific products.</p> <p>The type of PIDs among the citing and the cited products include:</p> <ul> <li>DOI</li> <li>PMC</li> <li>PMID</li> <li>ArXiv</li> <li>Handle</li> </ul>

opencc-zeroApr 2023View details →
zenodo36/100

Supplemental data for: Visualization of rank-citation curves for fast detection of possible manipulations with the h-index of the university

<p>This dataset consists of papers of universities in the top 30 Scopus Ranking of Ukrainian Universities (May 2023). The data was obtained from Scopus using the search query &quot;AF-ID (&ldquo;university name&rdquo;) AND PUBYEAR &lt; 2023 AND PUBYEAR &gt; 2002&quot;. Rank-citation curves were also generated for the publications of each university. In this analysis, the rank of publications was plotted along the horizontal axis, while the corresponding citation counts were depicted on the left axis. All types of documents were included in the dataset.</p>

opencc-by-4.0Jun 2023View details →
zenodo36/100

Factors Associated with Scientific Production Citations in Dentistry: Zero-inflated Negative Binomial Regression and Hurdle Modelling

<p><strong>Abstracto:</strong> La literatura cient&iacute;fica mundial en odontolog&iacute;a ha mostrado importantes avances en este campo, con importantes contribuciones que van desde el an&aacute;lisis de los aspectos epidemiol&oacute;gicos b&aacute;sicos de la prevenci&oacute;n hasta resultados especializados en el campo de los tratamientos dentales. La presente investigaci&oacute;n tiene como objetivo analizar el estado actual de la literatura cient&iacute;fica sobre odontolog&iacute;a alojada en la base de datos Web of Science. La metodolog&iacute;a incluye dos fases en el an&aacute;lisis de art&iacute;culos y revisiones indexadas en todas las &aacute;reas tem&aacute;ticas. Durante la primera fase, se analizan las siguientes variables: la producci&oacute;n cient&iacute;fica por parte del editor, la evoluci&oacute;n de la producci&oacute;n cient&iacute;fica publicada por los editores, los factores asociados al impacto de la producci&oacute;n cient&iacute;fica y la modelizaci&oacute;n del impacto de la producci&oacute;n cient&iacute;fica en odontolog&iacute;a. Durante la segunda fase, se analizan asociaciones, evoluciones y tendencias en el uso de palabras clave principales en la literatura cient&iacute;fica en odontolog&iacute;a. En conclusi&oacute;n, el estudio muestra que los temas m&aacute;s estudiados incluyen la asociaci&oacute;n de la educaci&oacute;n dental y el plan de estudios, la asociaci&oacute;n de la odontolog&iacute;a pedi&aacute;trica con la salud oral y el cuidado dental. Los hallazgos muestran que tambi&eacute;n destacan temas enfatizados m&aacute;s recientemente, como la odontolog&iacute;a basada en la evidencia, la pandemia, el control de infecciones y la endodoncia, as&iacute; como la necesidad de futuras investigaciones para ampliar el conocimiento actual basado en temas emergentes en la literatura cient&iacute;fica sobre odontolog&iacute;a.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

The data of "A comparison of citation-based clustering and topic modeling for science mapping"

<p>These files consist of the data used in &quot;A comparison of citation-based clustering and topic modeling for science mapping&quot;.&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
dryad36/100

Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns

<p>Questionable research practices (QRPs) have been the focus of the scientific community amid greater scrutiny and evidence highlighting issues with replicability across many fields of science. To capture the most impactful publications and the main thematic domains in the literature on QRPs, this study uses a document co-citation analysis. The analysis was conducted on a sample of 341 documents that covered the past 50 years of research in QRPs. Nine major thematic clusters emerged. Statistical reporting and statistical power emerged as key areas of research, where systemic-level factors in how research is conducted are consistently raised as the precipitating factors for QRPs. There is also an encouraging shift in the focus of research into open science practices designed to address engagement in QRPs. Such a shift is indicative of the growing momentum of the open science movement, and more research can be conducted on how these practices are employed on the ground and how their uptake by researchers can be further promoted. However, the results suggest that, while pre-registration and registered reports receive the most research interest, less attention has been paid to other open science practices (e.g., data and methods sharing).</p>

opencc-zeroSep 2023View details →
dryad36/100

Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns

Open the record for dataset details and reuse information.

publicSep 2023View details →
dryad36/100

Research on the benefits of nature to people: How much overlap is there in citations and terms for ‘nature’ across disciplines?

Open the record for dataset details and reuse information.

publicDec 2023View details →
dryad36/100

High-frequency location data show that race affects citations and fines for speeding

Open the record for dataset details and reuse information.

publicMar 2025View details →
dryad36/100

Agrivoltaic grazing systems for a sustainable future: Citation database for a multi-disciplinary review & gap analysis

Open the record for dataset details and reuse information.

publicJul 2025View details →
dryad36/100

Labeled data for citation field extraction

Open the record for dataset details and reuse information.

publicMar 2022View details →
dryad36/100

Forecasting the publication and citation outcomes of Covid-19 preprints

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad36/100

The disruption index suffers from citation inflation and is confounded by shifts in scholarly citation practice: synthetic citation networks for bibliometric null models

Open the record for dataset details and reuse information.

publicFeb 2025View details →
zenodo32/100

CORD-19_ scite_citation_tallies+contexts

<pre>Update: As of March 27, 2020 we have now analyzed 31,527 distinct sources (articles and preprints) from the most recent CORD-19 data (<a href="https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/version/4">https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge/version/4</a>. We&#39;re releasing citation tallies for these sources (covid-source-tallies 32720.csv). We&#39;re also releasing citation statements and classifications from these documents for open articles, which includes 1,682,216 out of the total 1,779,024 extracted. On March 20, 2020 we have analyzed 20,268 out of the 21,792 DOIS available from the <a href="https://www.kaggle.com/allen-institute-for-ai/CORD-19-research-challenge">CORD-19 data set</a>. Of these documents we found citations citing 16,775 of them, and the classifications for these citations are included in covid-source-tallies.csv. covid-citations.csv includes all citations we have from all of the ~20k documents we have processed. This file is truncated to make sure it only includes openly available documents. The tallies however are not limited by this, and it is the full set relating to all source documents scite has processed.</pre>

opencc-by-4.0Mar 2020View details →
zenodo32/100

Data citation for a forward stratigraphic-based porosity and permeability model developed for the Volve field, Norway.

<p>The data, models, and script presented here are those used for developing a forward stratigraphic simulation.&nbsp;The&nbsp;data include: 24 suits of well logs, seismic data, forward stratigraphic simulation scenarios of the shallow marine depositional setting, synthetic wells derived from the stratigraphic model, and 3-D reservoir models in Eclipse and RMS formats. In addition, a short script from the property calculator tool in Petrel, which is was used to classify lithofacies-associations&nbsp;in the stratigraphic model is also provided. The Petrel software license and code used in&nbsp;GPM&nbsp;software to undertake these forward stratigraphic simulations cannot be provided, because Schlumberger, who are the developers of the software do not allow its code to be shared in any publication.</p>

opencc-by-4.0May 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record