Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
39
datasets available to search
ShareScore release 0.9.0
Dataset results
39 results for “PUBMED”
PubMed literature on patient narratives 1975 - 2021
<p>This dataset contains PubMed entries of papers on patient narratives published between 1975 (first mention of 'patient narratives' in PubMed) and 2021. The original query is: </p> <blockquote> <p>"1975/01/01"[Date - Publication] : "2021/12/31"[Date - Publication] AND ("patient s"[All Fields] OR "patients"[MeSH Terms] OR "patients"[All Fields] OR "patient"[All Fields] OR "patients s"[All Fields]) AND ("narration"[MeSH Terms] OR "narration"[All Fields] OR "narrative"[All Fields] OR "narratives"[All Fields] OR "narrative s"[All Fields] OR "narratively"[All Fields])</p> </blockquote> <p>After duplicate removal the software contains 26349 PubMed entries.</p> <p>The dataset has been generated and processed with TopicTracker (<a href="https://zenodo.org/record/6373635">https://zenodo.org/record/6373635</a>) and thus contains also results of NLP analysis in the form of tabular data, plots, and word clouds. The dataset can be further explored either with TopicTracker, or with any other software.</p> <p> </p>
2.4 million GLOVE word/phrase vectors SQLite database trained on PubMed abstracts
<p>This is a 2.4 million GLOVE word/phrase vectors SQLite database trained on PubMed 2021 abstracts that can be used as word/phrase embeddings in machine learning applications.</p>
PubMed inner references obtained from five freely available bibliographic data sources
<p>This dataset contains PMID-to-PMID citations of PubMed 2020 Baseline extracted from five freely available bibliographic data sources (COCI, Dimensions, MAG, NIH-OCC, and S2ORC).</p> <p>Each line contains one citing PubMed document and its cited references. The citing and cited documents are separated by a tab (\t) and the cited references are separated by a semicolon (;).</p>
dataset of papers from PubMed related to Dendritic potential as immune contraception
<p>This is the dataset from PubMed. It consist of two files:</p> <p>1. Abstract of each previous studies used in form of txt file</p> <p>2. Details of each previous studies used in form of CSV file</p>
Disambiguated Author Identity (ID) Dataset for PubMed
<p>This dataset contains disambiguated author identities (IDs) of all authors in PubMed, which are created by our method. A research paper on this method is currently in preparation.</p>
NLMChem a new resource for chemical entity recognition in PubMed full-text literature
Open the record for dataset details and reuse information.
PubMed IDs from CENTRAL for the year 2016 - for RCT filter paper
<p>PubMed IDs from CENTRAL for the year 2016 - for RCT filter paper</p>
Trialstreamer PubMed RCTs 2020-04-26
Trialstreamer annotated collection of RCTs 2020-04-26
Three Kinds of Disambiguated Author ID Systems for PubMed 2019
<p>Author identifier (ID) is essential for many downstream tasks, such as co-author network and scientist mobility analysis. As a widely used bibliometrics database, author ID of PubMed is not officially provided by National Institutes of Health (NIH), that restrict bibliometric research. This study exploited three open bibliographic databases Aminer, Microsoft Academic Graph (MAG) and Semantic Scholar (S2) to associate author ID for PubMed. For this purpose, paper linking and author linking was performed sequencely to mine paper and author links between PubMed and these databases. Performance of author name disambiguation (AND) of there available identifiers was evaluated on two AND datasets. Our findings suggested that, S2 contains full volume of PubMed regarding link completeness. With respect to correctness of author ID, S2 and MAG achieved better performance than Aminer. The best F1 score on both dataset of there available identifiers is below 90\%, indicate that AND for large scale database remain as a difficult task and the need for further improvement. We made the final dataset that contains linked paper and author of PubMed publicly available for facilitating future research.</p>
Relevance assessments, bibliometrics, and altmetrics - A quantitative study on PubMed and arXiv
<p>This dataset contains the data and the script of "Relevance assessments, bibliometrics, and altmetrics - A quantitative study on PubMed and arXiv".</p>
Trialstreamer PubMed RCTs 2020-04-24
Trialstreamer annotated collection of RCTs 2020-04-24
Open Access of COVID-19 related publications in the first quarter of 2020: a preliminary study based in PubMed
<p>Underlying data of the article: "Open Access of COVID-19 related publications in the first quarter of 2020: a preliminary study based in PubMed". Version 1 and Version 2 (after Unpaywall update).</p> <p>Data was analysed using Unpaywall and OpenRefine.</p>
List of studies on Covid-19 identified in PubMed and excluded from our systematic review
<p>In the course of our PubMed searches we identified and screened the title and abstract of a number of studies on Covid-19 that were finally excluded from our systematic review. </p> <p>This file will be updated regularly.</p>
Test set - 4023 PubMed abstracts (for manuscript: Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait )
<p>A .zip archive containing the set of abstracts used in the test set (4023 abstracts from PubMed) in .txt format.</p> <p>This archive contains supplementary files for the manuscript Extracting knowledge networks from plant scientific literature: Potato tuber flesh color as an exemplary trait.</p>
LAGOS-AND-PM: A Large Gold Standard Dataset for PubMed Author Name Disambiguation
<p>LAGOS-AND-PM is an author name disambiguation (AND) dataset for the PubMed database, containing several versions, and they are built based on the ORCID database and the PubMed literature database.</p> <p>Note that we have previously created another dataset named <a href="https://zenodo.org/record/7313380">LAGOS-AND</a>, which refers to a series of AND datasets created for the MAG/OpenAlex database.</p>
Pubmed citation dataset
<p>The scripts for generating the datasets are available at <a href="https://github.com/jokergoo/citation_analysis">https://github.com/jokergoo/citation_analysis</a>.</p> <p>There are three files in this dataset:</p> <p>1. citations.tab.gz: A table with two columns:</p> <ul> <li>citing: pmid (PubMed ID) of the citing paper</li> <li>cited: pmid of the cited paper</li> </ul> <pre><code>citing cited 10578099 10578100 10578100 10578099 10590126 10623754 10592169 10592170 10592169 10592175 10592169 10592200</code></pre> <p>2. pub_meta.tab.gz: A table with 8 columns:</p> <ul> <li>pmid: pmid of the paper.</li> <li>journal_uid: uid of the paper on PubMed.</li> <li>pub_year: year of the paper.</li> <li>n_authors: number of authors.</li> <li>country: identified country of the paper.</li> <li>country_type: type of the country identification. Values are `_domestic_`, `_domestic_80_`, `_international_` or `_empty_`. `_domestic_80_` means less than 20% of international authors in the middle of the author list.</li> <li>file_id: file id on PubMed FTP.</li> <li>n_references: number of references of the paper.</li> </ul> <pre><code>pmid journal_uid pub_year n_authors country country_type file_id n_references 38566917 101213162 2016 4 United States _domestic_ 1366 2 38567026 9886008 2022 3 United States _domestic_ 1366 12 38567115 101562981 2018 5 United States _domestic_ 1366 6 38567118 101668947 2018 4 United States _domestic_ 1365 0 38567245 101283276 2019 10 United States _domestic_ 1366 16</code></pre> <p>3. num_cite_country_country.tab.gz: A table with three columns:</p> <ul> <li>country_cited: country of the cited papers.</li> <li>country_citing: country of the citing papers.</li> <li>citations: total number of citations.</li> </ul> <pre><code>country_cited country_citing citations Afghanistan Afghanistan 20 Afghanistan Australia 16 Afghanistan Austria 1 Afghanistan Bangladesh 4 Afghanistan Belgium 3 Afghanistan Brazil 8</code></pre>
Trialstreamer PubMed RCTs 2020-04-24
Trialstreamer annotated collection of RCTs 2020-04-24
Trialstreamer PubMed RCTs 2020-04-24
Trialstreamer annotated collection of RCTs 2020-04-24
PubMed Central Open Available Subset Simialr Articles (using BM25)
<p>PubMed Central Open Available Subset Simialr Articles (using BM25) </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.