Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “PUBMED”

Learn how ShareScore rates datasets ↗
zenodo36/100

PubMed literature on patient narratives 1975 - 2021

<p>This dataset contains PubMed entries of papers on patient narratives published between 1975 (first mention of &#39;patient narratives&#39; in PubMed) and 2021. The original query is:&nbsp;</p> <blockquote> <p>&quot;1975/01/01&quot;[Date - Publication] : &quot;2021/12/31&quot;[Date - Publication] AND (&quot;patient s&quot;[All Fields] OR &quot;patients&quot;[MeSH Terms] OR &quot;patients&quot;[All Fields] OR &quot;patient&quot;[All Fields] OR &quot;patients s&quot;[All Fields]) AND (&quot;narration&quot;[MeSH Terms] OR &quot;narration&quot;[All Fields] OR &quot;narrative&quot;[All Fields] OR &quot;narratives&quot;[All Fields] OR &quot;narrative s&quot;[All Fields] OR &quot;narratively&quot;[All Fields])</p> </blockquote> <p>After duplicate removal the software contains&nbsp;26349 PubMed entries.</p> <p>The dataset has been generated and processed with TopicTracker (<a href="https://zenodo.org/record/6373635">https://zenodo.org/record/6373635</a>) and thus contains also results of NLP analysis in the form of tabular data, plots, and word clouds. The dataset can be further explored either with TopicTracker, or with any other software.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo36/100

2.4 million GLOVE word/phrase vectors SQLite database trained on PubMed abstracts

<p>This is a 2.4 million GLOVE word/phrase vectors SQLite database trained on PubMed 2021 abstracts that can be used as word/phrase embeddings in machine learning applications.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

PubMed inner references obtained from five freely available bibliographic data sources

<p>This dataset contains PMID-to-PMID citations of&nbsp;PubMed 2020 Baseline extracted from five freely available bibliographic data sources (COCI, Dimensions, MAG, NIH-OCC, and S2ORC).</p> <p>Each line contains one citing PubMed document and its cited references. The citing and cited documents are separated by a tab (\t) and the cited references are separated by a semicolon (;).</p>

opencc-by-4.0Aug 2021View details →
zenodo36/100

dataset of papers from PubMed related to Dendritic potential as immune contraception

<p>This is the dataset from PubMed. It consist of two files:</p> <p>1. Abstract of each previous studies used in form of txt file</p> <p>2. Details of each previous studies used in form of CSV file</p>

opencc-by-4.0Feb 2023View details →
zenodo36/100

Disambiguated Author Identity (ID) Dataset for PubMed

<p>This dataset contains disambiguated author identities (IDs) of all authors in PubMed, which are created by our method. A research paper on this method is currently in preparation.</p>

opencc-by-4.0Feb 2023View details →
dryad36/100

NLMChem a new resource for chemical entity recognition in PubMed full-text literature

Open the record for dataset details and reuse information.

publicMar 2021View details →
zenodo32/100

PubMed IDs from CENTRAL for the year 2016 - for RCT filter paper

<p>PubMed IDs from CENTRAL for the year 2016 - for RCT filter paper</p>

opencc-by-4.0Jan 2020View details →
zenodo32/100

Trialstreamer PubMed RCTs 2020-04-26

Trialstreamer annotated collection of RCTs 2020-04-26

opencc-zeroApr 2020View details →
zenodo32/100

Three Kinds of Disambiguated Author ID Systems for PubMed 2019

<p>Author identifier (ID) is essential for many downstream tasks, such as co-author network and scientist mobility analysis. As a widely used bibliometrics database, author ID of PubMed is not officially provided by National Institutes of Health (NIH), that restrict bibliometric research. This study exploited three open bibliographic databases Aminer, Microsoft Academic Graph (MAG) and Semantic Scholar (S2) to associate author ID for PubMed. For this purpose, paper linking and author linking was performed sequencely to mine paper and author links between PubMed and these databases. Performance of author name disambiguation (AND) of there available identifiers was evaluated on two AND datasets. Our findings suggested that, S2 contains full volume of PubMed regarding link completeness. With respect to correctness of author ID, S2 and MAG achieved better performance than Aminer. The best F1 score on both dataset of there available identifiers is below 90\%, indicate that AND for large scale database remain as a difficult task and the need for further improvement. We made the final dataset that contains linked paper and author of PubMed publicly available for facilitating future research.</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

Relevance assessments, bibliometrics, and altmetrics - A quantitative study on PubMed and arXiv

<p>This dataset contains the data and the script of &quot;Relevance assessments, bibliometrics, and altmetrics - A quantitative study on PubMed and arXiv&quot;.</p>

opencc-by-4.0Jan 2022View details →
zenodo28/100

Trialstreamer PubMed RCTs 2020-04-24

Trialstreamer annotated collection of RCTs 2020-04-24

opencc-zeroApr 2020View details →
zenodo28/100

Open Access of COVID-19 related publications in the first quarter of 2020: a preliminary study based in PubMed

<p>Underlying data of the article: &quot;Open Access of COVID-19 related publications in the first quarter of 2020: a preliminary study based in PubMed&quot;. Version 1 and Version 2 (after Unpaywall update).</p> <p>Data was analysed using Unpaywall and OpenRefine.</p>

opencc-by-4.0May 2020View details →
zenodo28/100

List of studies on Covid-19 identified in PubMed and excluded from our systematic review

<p>In the course of our PubMed searches we identified and screened the title and abstract of a number of studies on Covid-19 that were finally excluded from our systematic review.&nbsp;</p> <p>This file will be updated regularly.</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

Test set - 4023 PubMed abstracts (for manuscript: Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait )

<p>A .zip archive containing the set of abstracts used in the test set (4023 abstracts from PubMed) in .txt format.</p> <p>This archive contains supplementary files for the manuscript Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait.</p>

opencc-by-4.0Dec 2019View details →
zenodo28/100

LAGOS-AND-PM: A Large Gold Standard Dataset for PubMed Author Name Disambiguation

<p>LAGOS-AND-PM is an author name disambiguation (AND) dataset for the PubMed database, containing several versions, and they are built based on the ORCID database and the PubMed literature database.</p> <p>Note that we have previously created another dataset named <a href="https://zenodo.org/record/7313380">LAGOS-AND</a>, which refers to a series of AND datasets created for the MAG/OpenAlex database.</p>

opencc-by-4.0Apr 2023View details →
zenodo24/100

Pubmed citation dataset

<p>The scripts for generating the datasets are available at <a href="https://github.com/jokergoo/citation_analysis">https://github.com/jokergoo/citation_analysis</a>.</p> <p>There are three files in this dataset:</p> <p>1. citations.tab.gz: A table with two columns:</p> <ul> <li>citing: pmid (PubMed ID) of the citing paper</li> <li>cited: pmid of the cited paper</li> </ul> <pre><code>citing cited 10578099 10578100 10578100 10578099 10590126 10623754 10592169 10592170 10592169 10592175 10592169 10592200</code></pre> <p>2. pub_meta.tab.gz: A table with 8 columns:</p> <ul> <li>pmid: pmid of the paper.</li> <li>journal_uid: uid of the paper on PubMed.</li> <li>pub_year: year of the paper.</li> <li>n_authors: number of authors.</li> <li>country: identified country of the paper.</li> <li>country_type: type of the country identification. Values are `_domestic_`, `_domestic_80_`, `_international_` or `_empty_`. `_domestic_80_` means less than 20% of international authors in the middle of the author list.</li> <li>file_id: file id on PubMed FTP.</li> <li>n_references: number of references of the paper.</li> </ul> <pre><code>pmid journal_uid pub_year n_authors country country_type file_id n_references 38566917 101213162 2016 4 United States _domestic_ 1366 2 38567026 9886008 2022 3 United States _domestic_ 1366 12 38567115 101562981 2018 5 United States _domestic_ 1366 6 38567118 101668947 2018 4 United States _domestic_ 1365 0 38567245 101283276 2019 10 United States _domestic_ 1366 16</code></pre> <p>3. num_cite_country_country.tab.gz: A table with three columns:</p> <ul> <li>country_cited: country of the cited papers.</li> <li>country_citing: country of the citing papers.</li> <li>citations: total number of citations.</li> </ul> <pre><code>country_cited country_citing citations Afghanistan Afghanistan 20 Afghanistan Australia 16 Afghanistan Austria 1 Afghanistan Bangladesh 4 Afghanistan Belgium 3 Afghanistan Brazil 8</code></pre>

opencc-by-nc-4.0Apr 2024View details →
zenodo20/100

Trialstreamer PubMed RCTs 2020-04-24

Trialstreamer annotated collection of RCTs 2020-04-24

opencc-zeroApr 2020View details →
zenodo20/100

Trialstreamer PubMed RCTs 2020-04-24

Trialstreamer annotated collection of RCTs 2020-04-24

opencc-zeroApr 2020View details →
zenodo8/100

PubMed Central Open Available Subset Simialr Articles (using BM25)

<p>PubMed Central Open Available Subset Simialr Articles (using BM25)&nbsp;</p>

restrictedAug 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record