Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

28

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

28 results for “Dataset citation”

Learn how ShareScore rates datasets ↗
dryad36/100

Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns

<p>Questionable research practices (QRPs) have been the focus of the scientific community amid greater scrutiny and evidence highlighting issues with replicability across many fields of science. To capture the most impactful publications and the main thematic domains in the literature on QRPs, this study uses a document co-citation analysis. The analysis was conducted on a sample of 341 documents that covered the past 50 years of research in QRPs. Nine major thematic clusters emerged. Statistical reporting and statistical power emerged as key areas of research, where systemic-level factors in how research is conducted are consistently raised as the precipitating factors for QRPs. There is also an encouraging shift in the focus of research into open science practices designed to address engagement in QRPs. Such a shift is indicative of the growing momentum of the open science movement, and more research can be conducted on how these practices are employed on the ground and how their uptake by researchers can be further promoted. However, the results suggest that, while pre-registration and registered reports receive the most research interest, less attention has been paid to other open science practices (e.g., data and methods sharing).</p>

opencc-zeroSep 2023View details →
dryad36/100

Dataset for: Fifty years of research on questionable research practices in science: Quantitative analysis of co-citation patterns

Open the record for dataset details and reuse information.

publicSep 2023View details →
zenodo32/100

Wikipedia Citations: A comprehensive dataset of citations with identifiers extracted from English Wikipedia

<p>The dataset is composed of <strong>3 parts</strong>:</p> <p>1. &nbsp;The dataset of 29.276 million citations from &nbsp;35 different citation templates, &nbsp;out of which 3.92&nbsp;million citations already contained identifiers, and approximately 260,752 citations were equipped with identifiers from Crossref. This is under the filename: <strong>citations_from_wikipedia.zip</strong></p> <p>2. &nbsp;A minimal dataset containing a few of the columns from the citations from Wikipedia dataset. These columns are as follows:&nbsp;&#39;type_of_citation&#39;, &#39;page_title&#39;, &#39;Title&#39;, &#39;ID_list&#39;, metadata_file&#39;, &#39;updated_identifier&#39;.&nbsp;This is under the filename: <strong>minimal_dataset.zip. </strong>The &#39;metadata_file&#39; column can be used to refer to the metadata collected from CrossRef and page title, the title of the citation can be used to refer to the &#39;citations_from_wikipedia.zip&#39; dataset and get more information for a particular citation (such as author, periodical, chapter).</p> <p>3. &nbsp;Citations classified as a journal and their corresponding metadata/identifier extracted from Crossref to make the dataset more complete. This is under the filename: <strong>lookup_data.zip</strong>. This zip file contains a CSV file: <strong>lookup_table.gzip</strong> (a parquet file containing all citations classified as a journal) and a folder:<strong> metadata_extracted </strong>(a folder containing the metadata from CrossRef for all the citations mentioned in the table)</p> <p><br> The data was parsed from the Wikipedia XML content dumps published in May 2020.</p> <p>The source code to extract and getting used to the pipeline can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki</strong></p> <p>The taxonomy of the dataset in (1) can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki/wiki/Taxonomy-of-the-parent-dataset</strong></p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

A Comprehensive Dataset of Classified Citations with Identifiers from English Wikipedia (2024)

<p><strong>2024 (new!)</strong></p> <p>This is a dataset of 44.766.800 (+9.2%)&nbsp; citations extracted from the English Wikipedia February 2024 dump (<a href="https://dumps.wikimedia.org/enwiki/20240220/">https://dumps.wikimedia.org/enwiki/20240220/</a>).</p> <p>The same extraction and template harmonization pipeline was used as the year before. The published dataset fields are like in the previous dataset. A classification label is assigned to each citation (either 'news', 'book', 'journal' or 'other)' by the deterministic rule-based classifier that analyses available identifiers (see code documentation for details), revealing the following citation subgroups:</p> <ol> <li>The total number of news: 10.958.151 (+9.4%)</li> <li>The total number of books:* 3.277.629 (+8.6%)</li> <li>The total number of journals*: 2.248.748 (+8.7%)</li> </ol> <p>* Please note that these numbers do not represent the overall number of book and journal citations, we count only citations with DOI, PMID, PMC and ISBN identifiers assigned by authors (prior to the lookup process that augments citations with missing identifiers).&nbsp;&nbsp;</p> <p>This dataset is not equipped with identifiers located via the lookup process (no 'acquired_ID_list' field). If there is interest in such an augmented version, see the source code for instructions or contact authors for assistance with this task.&nbsp; &nbsp; &nbsp; &nbsp;</p> <p><strong>2023</strong></p> <p>This is a dataset of 40.664.485 citations extracted from the English Wikipedia February 2023 dump (<a href="https://dumps.wikimedia.org/enwiki/20230220/">https://dumps.wikimedia.org/enwiki/20230220/</a>).</p> <p>Version 1: en_citations.zip is a dataset of extracted citations&nbsp;</p> <p>Version 2: en_final.zip is the same dataset with classified citations augmented with identifiers&nbsp;&nbsp;</p> <p>The fields are as follows:</p> <ul> <li>type_of_citation - Wikipedia template type used to define the citation, e.g., 'cite journal', 'cite news', etc.</li> <li>page_title -&nbsp;title of the Wikipedia article from which the citation was extracted.</li> <li>Title - source title, e.g., title of the book, newspaper article, etc.</li> <li>URL -&nbsp;link to the source, e.g., webpage where news article was published, description of the book at the publisher's website,&nbsp;online library webpage, etc.</li> <li>tld - top link domain extracted from the URL, e.g., 'bbc' for&nbsp;https://www.bbc.co.uk/...&nbsp;</li> <li>Authors - list of article or book authors, if available.</li> <li>ID_list - list of publication identifiers mentioned in the citation, e.g.,&nbsp;DOI, ISBN, etc.</li> <li>citations - citation text as used in Wikipedia code</li> <li>actual_label - 'book', 'journal', 'news', or 'other' label assigned based on the analysis of citation identifiers or top link domain.&nbsp; &nbsp;&nbsp;</li> <li>acquired_ID_list - identifiers located via Google Books and Crossref APIs for citations which are likely to refer to books or journals, i.e., defined using 'cite book', 'cite journal', 'cite encyclopedia', and 'cite proceedings' templates.</li> </ul> <ol> <li>The total number of news: 9.926.598</li> <li>The total number of books: 2.994.601</li> <li>The total number of journals: 2.052.172</li> <li>Augmented with IDs via lookup 929.601 (out of 2.445.913&nbsp;book, journal, encyclopedia, and proceedings template citations not classified as books or journals via given identifiers).&nbsp;</li> </ol> <p>The source code to extract citations can be found here:&nbsp;<strong><a href="https://github.com/albatros13/wikicite">https://github.com/albatros13/wikicite</a>. </strong></p> <p>The code is a fork of the earlier project on Wikipedia citation extraction: <a href="https://github.com/Harshdeep1996/cite-classifications-wiki">https://github.com/Harshdeep1996/cite-classifications-wiki</a>.</p> <p>&nbsp;</p>

opencc-by-4.0Feb 2024View details →
zenodo32/100

Dataset for: Substituting confidence for competence in health literacy: A review of studies, citations, and trial registrations

<p><span>Patient health literacy is crucial for effective patient&ndash;physician communication, and interventions targeting health literacy can use measures based on either actual performance (competence) or self-ratings (confidence). This paper analyzed the development of these measures through three studies. Study 1 reviewed articles describing the development of novel measures; Study 2 examined the citations of these studies, and Study 3 evaluated data from clinical trials registries. The literature search was conducted from April 14&ndash;27, 2023. PubMed was used as the main database in which studies on health literacy measures were searched for the systematic review (Study 1). We then used Google Scholar and the OpenCitations database to describe citation patterns of the included health literacy measures (Study 2). Finally, we evaluated confidence- or competence-based health literacy measures by extracting and analyzing trial data from ClinicalTrials.gov (Study 3). Our review included 55 health literacy measures, among which 23 (42%) were competence-based, 28 (51%) confidence-based, and 4 (7%) assessed both. Recent trends show a shift toward developing more confidence-based measures and a decline in creating new competence-based measures. Confidence-based measures were increasingly cited, whereas citations for competence-based measures have plateaued. Lastly, our findings showed a steady increase in the use of confidence-based measures in recent clinical trials and a decrease in the use of competence-based measures when controlling for sample size. This shift may be detrimental to public health given our limited knowledge about patients&rsquo; ability to meet demands of shared decision-making, especially regarding new technologies like artificial intelligence in healthcare.</span></p>

opencc-by-4.0Aug 2024View details →
zenodo28/100

Datasets indexed in Data Citation Index in the Astronomy and Astrophysics category, 2010-2019

<p>The dataset comprises a single list of datasets exported from Data Citation Index (Web of Science, Clarivate Analytics) in the Astronomy and Astrophysics category, for the period 2010 - 2019, allowing to identify annual evolution, countries and institutions with higher productivity, main repositories and hosting platforms, use in publications indexed in Web of Science.</p>

opencc-by-4.0Sep 2021View details →
zenodo24/100

Pubmed citation dataset

<p>The scripts for generating the datasets are available at <a href="https://github.com/jokergoo/citation_analysis">https://github.com/jokergoo/citation_analysis</a>.</p> <p>There are three files in this dataset:</p> <p>1. citations.tab.gz: A table with two columns:</p> <ul> <li>citing: pmid (PubMed ID) of the citing paper</li> <li>cited: pmid of the cited paper</li> </ul> <pre><code>citing cited 10578099 10578100 10578100 10578099 10590126 10623754 10592169 10592170 10592169 10592175 10592169 10592200</code></pre> <p>2. pub_meta.tab.gz: A table with 8 columns:</p> <ul> <li>pmid: pmid of the paper.</li> <li>journal_uid: uid of the paper on PubMed.</li> <li>pub_year: year of the paper.</li> <li>n_authors: number of authors.</li> <li>country: identified country of the paper.</li> <li>country_type: type of the country identification. Values are `_domestic_`, `_domestic_80_`, `_international_` or `_empty_`. `_domestic_80_` means less than 20% of international authors in the middle of the author list.</li> <li>file_id: file id on PubMed FTP.</li> <li>n_references: number of references of the paper.</li> </ul> <pre><code>pmid journal_uid pub_year n_authors country country_type file_id n_references 38566917 101213162 2016 4 United States _domestic_ 1366 2 38567026 9886008 2022 3 United States _domestic_ 1366 12 38567115 101562981 2018 5 United States _domestic_ 1366 6 38567118 101668947 2018 4 United States _domestic_ 1365 0 38567245 101283276 2019 10 United States _domestic_ 1366 16</code></pre> <p>3. num_cite_country_country.tab.gz: A table with three columns:</p> <ul> <li>country_cited: country of the cited papers.</li> <li>country_citing: country of the citing papers.</li> <li>citations: total number of citations.</li> </ul> <pre><code>country_cited country_citing citations Afghanistan Afghanistan 20 Afghanistan Australia 16 Afghanistan Austria 1 Afghanistan Bangladesh 4 Afghanistan Belgium 3 Afghanistan Brazil 8</code></pre>

opencc-by-nc-4.0Apr 2024View details →
zenodo24/100

A Comprehensive Dataset of Classified Citations with Identifiers from Multilingual Wikipedia (2024)

<p>This is a collection of translated citation datasets extracted from the Multilingual Wikipedia February 2024 dumps. The same extraction and template harmonization pipeline was used as for English Wikipedia&nbsp;<a title="English Wikipedia citations" href="../records/10782978">https://zenodo.org/records/10782978</a>.&nbsp;</p> <p><strong>Note:&nbsp;</strong> Versions 2 and 3 fix issues with large Italian, French and German datasets that were corrupted (failed to upload in full) in the initial version.</p> <p>In each language, Wikipedia authors can cite sources using language-specific or English templates. Our main effort in compiling these datasets was to assemble lists of citation templates for each language and convert relevant fields into a common English template. We started with known citation templates per each language (typically covering books, journals, web pages and news), and, in some cases, augmented these lists with additional frequently used templates (films, links, webarchives, etc.) which we were able to locate via the XML reference tags vs usage frequency dictionaries. For the list of accepted templates see our source code:&nbsp;<a href="https://github.com/albatros13/wikicite/tree/multilang">https://github.com/albatros13/wikicite/tree/multilang</a> (templates are listed in __init__.py files of the wikiciteparser library).</p> <p>A classification label is assigned to each citation (either 'news', 'book', 'journal' or 'other)' by the deterministic rule-based classifier that analyses available identifiers (see code documentation for details). Please note that these numbers do not represent the overall estimation of the book and journal citation numbers. We count only citations with DOI, PMID, PMC and ISBN identifiers assigned by authors (prior to the lookup process that augments citations with missing identifiers). The number of news citations is dependent on our list of recognised 22.646 <a href="https://github.com/albatros13/wikicite/blob/master/news/domains.txt">news agency domains</a>.&nbsp;&nbsp;</p> <table> <tbody> <tr> <td>Language</td> <td>Acronym</td> <td>Link</td> <td>Dump size</td> <td>Citations</td> <td>Books</td> <td>Journals</td> <td>News</td> </tr> <tr> <td>German&nbsp;</td> <td>de</td> <td><a href="https://dumps.wikimedia.org/dewiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dewiki/20240220/</a></td> <td>6.7GB</td> <td>4.854.945</td> <td>320.179</td> <td>105.542</td> <td>901.091</td> </tr> <tr> <td>French&nbsp;</td> <td>fr</td> <td><a href="https://dumps.wikimedia.org/frwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/frwiki/20240220/</a></td> <td>5.9GB</td> <td>9.552.768</td> <td>798.525</td> <td>264.560</td> <td>1.907.183</td> </tr> <tr> <td>Russian</td> <td>ru</td> <td><a href="https://dumps.wikimedia.org/ruwiki/20240220/">https://dumps.wikimedia.org/ruwiki/20240220/</a></td> <td>5.1GB</td> <td>7.437.100</td> <td>420.828</td> <td>130.470</td> <td>1.370.665</td> </tr> <tr> <td>Spanish</td> <td>es</td> <td><a href="https://dumps.wikimedia.org/eswiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/eswiki/20240220/</a></td> <td>4.2GB</td> <td>6.918.442</td> <td>522.910</td> <td>213.767</td> <td>1.699.396</td> </tr> <tr> <td>Italian</td> <td>it</td> <td><a href="https://dumps.wikimedia.org/itwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/itwiki/20240220/</a></td> <td>3.6GB</td> <td>5.545.082</td> <td>384.816</td> <td>128.366</td> <td>917.517</td> </tr> <tr> <td>Polish</td> <td>pl</td> <td><a href="https://dumps.wikimedia.org/plwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/plwiki/20240220/</a></td> <td>2.4GB</td> <td>4.744.158</td> <td>463.783&nbsp;</td> <td>95.988</td> <td>513.006</td> </tr> <tr> <td>Portuguese</td> <td>pt</td> <td><a href="https://dumps.wikimedia.org/ptwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/ptwiki/20240220/</a></td> <td>2.2GB</td> <td>4.775.025</td> <td>243.593</td> <td>142.216&nbsp;</td> <td>1.176.140</td> </tr> <tr> <td>Dutch</td> <td>nl</td> <td><a href="https://dumps.wikimedia.org/nlwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nlwiki/20240220/</a></td> <td>1.8GB</td> <td>566.549</td> <td>27.074&nbsp;</td> <td>12.706</td> <td>114.110</td> </tr> <tr> <td>Swedish</td> <td>sv</td> <td><a href="https://dumps.wikimedia.org/svwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/svwiki/20240220/</a></td> <td>1.5GB</td> <td>3.802.416</td> <td>112.748</td> <td>155.740&nbsp;</td> <td>869.662&nbsp;</td> </tr> <tr> <td>Catalan</td> <td>ca</td> <td><a href="https://dumps.wikimedia.org/cawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/cawiki/20240220/</a></td> <td>1.2GB</td> <td>2.239.714</td> <td>261.779</td> <td>105.125</td> <td>423.241</td> </tr> <tr> <td>Finnish</td> <td>fi</td> <td><a href="https://dumps.wikimedia.org/fiwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/fiwiki/20240220/</a></td> <td>900.9MB</td> <td>1.697.731</td> <td>209.556</td> <td>12.068</td> <td>286.420</td> </tr> <tr> <td>Turkish</td> <td>tr</td> <td><a href="https://dumps.wikimedia.org/trwiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/trwiki/20240220</a></td> <td>883.9MB</td> <td>1.993.177</td> <td>85.079</td> <td>56.202</td> <td>339.122&nbsp;</td> </tr> <tr> <td>Norwegian</td> <td>no</td> <td><a href="https://dumps.wikimedia.org/nowiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/nowiki/20240220</a></td> <td>763.7MB</td> <td>796.500</td> <td>43.314</td> <td>12.373</td> <td>151.780</td> </tr> <tr> <td>Danish</td> <td>da</td> <td><a href="https://dumps.wikimedia.org/dawiki/20240220/" target="_blank" rel="noopener">https://dumps.wikimedia.org/dawiki/20240220</a></td> <td>413.3MB</td> <td>437.239</td> <td>23.303</td> <td>7.522</td> <td>70.760&nbsp;</td> </tr> </tbody> </table> <p>This datasets can be equipped with identifiers located via the lookup process (no 'acquired_ID_list' field). If there is interest in augmented versions, see the source code for instructions or contact authors for assistance with this task.&nbsp; &nbsp;</p> <p>This research was supported in part by the <a href="https://dsc.uva.nl/">University of Amsterdam Data Science Centre</a>.</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record