Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “pubmed”

Learn how ShareScore rates datasets ↗
zenodo52/100

Pubmed Journal Recommendation System dataset

<p>Dataset for Journal recommendation, includes title, abstract, keywords, and journal.</p> <p>We extracted the journals and more information of:</p> <p>Jiasheng Sheng. (2022). PubMed-OA-Extraction-dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.6330817.</p> <p>Dataset Components:</p> <ul> <li> <p><strong>data_pubmed_all:</strong> This dataset encompasses all articles, each containing the following columns: 'pubmed_id', 'title', 'keywords', 'journal', 'abstract', 'conclusions', 'methods', 'results', 'copyrights', 'doi', 'publication_date', 'authors', 'AKE_pubmed_id', 'AKE_pubmed_title', 'AKE_abstract', 'AKE_keywords', 'File_Name'.</p> </li> <li> <p><strong>data_pubmed:</strong> To focus on recent and relevant publications, we have filtered this dataset to include articles published within the last five years, from January 1, 2018, to December 13, 2022&mdash;the latest date in the dataset. Additionally, we have exclusively retained journals with more than 200 published articles, resulting in 262,870 articles from 469 different journals.</p> </li> <li> <p><strong>data_pubmed_train, data_pubmed_val, and data_pubmed_test:</strong> For machine learning and model development purposes, we have partitioned the 'data_pubmed' dataset into three subsets&mdash;training, validation, and test&mdash;using a random 60/20/20 split ratio. Notably, this division was performed on a per-journal basis, ensuring that each journal's articles are proportionally represented in the training (60%), validation (20%), and test (20%) sets. The resulting partitions consist of 157,540 articles in the training set, 52,571 articles in the validation set, and 52,759 articles in the test set.</p> </li> </ul>

opencc-by-4.0Oct 2023View details →
zenodo48/100

Triangle of Biomedicine Framework to Analyze the Citations' Impact on Categories Dissemination in the PubMed Database

<p>This is the data and the most relevant script of the paper 'Triangle of Biomedicine Framework to Analyze the Citations&rsquo; Impact on Categories Dissemination in the PubMed Database'.</p>

opencc-by-4.0Jul 2024View details →
zenodo48/100

S37 | LITMINEDNEURO | Neurotoxicants from literature mining PubMed

<p>This is the collection associated with list S37 LITMINEDNEURO on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/?q=suspect-list-exchange">https://www.norman-network.com/?q=suspect-list-exchange</a></p> <p>S37</p> <p>LITMINEDNEURO</p> <p><strong>Neurotoxicants from literature mining PubMed</strong></p> <p>A list of chemicals associated with neurotoxicity compiled through systematic literature mining of PubMed using MeSH terms, compiled by Nancy Baker, Antony Williams (US EPA) and Emma Schymanski (LCSB), details in Schymanski et al (submitted) and on <a href="https://comptox.epa.gov/dashboard/chemical_lists/LITMINEDNEURO">CompTox list</a>.&nbsp; &nbsp;</p> <p>Updated 10/06/2019 to include entries that were registered but not yet publicly available when the previous version was released.</p>

opencc-by-4.0Jan 2019View details →
zenodo40/100

Spanish abstracts from PubMed (machine-translated from English)

<p>Spanish abstracts from PubMed (machine-translated from English)</p> <p>A state-of-the-art, domain-specific Neural Machine Translation system has been used to automatically translate a large number of articles (titles and abstracts) from PubMed, from English to Spanish. This dataset, in JSON format, provides MeSH terms and DeCS codes (if available), as well as other fields such as the year of publication for all articles.</p> <p>Copyright (c) 2020 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0May 2020View details →
zenodo40/100

PubMed-Temporal: A dynamic graph dataset with node-level features

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo40/100

Biomedical journals covered and uncovered in PubMed

<p>PubMed contains many biomedical journals, but is this the full range of biomedical journals?</p> <p>We compared PubMed with MAG and found that there are many journals that are not included in PubMed. In addition, we found that even for the biomedical journals covered by PubMed, there are many articles under the&nbsp;covered journals not found in PubMed. The two tables below present these pieces of information.</p> <p>intersecting-journal-with-indexing-completeness.tsv: referring to the journals whose articles are not fully included in PubMed</p> <p>uncovered-journal-with-number-biomedical-articles.tsv:&nbsp;referring to the journals that have not been touched by&nbsp;PubMed</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

PubMed-OA-Extraction-dataset

<p>This is the train-test-validation dataset for pubmed open-access articles keyphrase extraction task. The small_* file contains the all articles that have 5to 25 extractive keyphrases (keyphrase in the article that is inside the abstract of the article).</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

PGB: A PubMed Graph Benchmark for Heterogeneous Network Representation Learning

<p>PubMed Graph Benchmark (PGB)&nbsp;aggregates&nbsp;the metadata associated with the biomedical articles from PubMed into a unified source.&nbsp;The benchmark contains metadata including&nbsp;title, abstract, authors, in/out citations, MeSH terms, MeSH hierarchy, venue, publication type, and chemicals.</p>

opencc-by-4.0Mar 2022View details →
zenodo40/100

Supplemental data for: Variations in the naming of malondialdehyde (MDA) in PubMed-, Scopus-, and Web of Science-indexed literature

<p>Three scientific databases (PubMed, Scopus, and Web of Science (WoS)) were consulted (July 14, 2022) to assess the frequency of eight nomenclatural forms of malondialdehyde (MDA). Due to the peculiarities of the search interface of the selected databases, PubMed was searched in the Title and Abstract fields (search query example in PubMed: &quot;malone dialdehyde&quot;[Title/Abstract]), Scopus was searched in the Title, Abstract, and Keywords fields (search query example in Scopus: TITLE-ABS-KEY (&ldquo;malone dialdehyde&rdquo;)), and WoS Core Collection was searched in the Title, Abstract, Author keywords, and Keywords Plus fields (search query example in WoS Core Collection: TS=(&ldquo;malone dialdehyde&rdquo;)). All types of publications for the years 2002-2021 are taken into account.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Dataset: PubMatic, Inc. (PUBM) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →
zenodo40/100

Functional foods and cancer on Pinterest and PubMed: Myths and science

<p>The objective of this study was to identify&nbsp;the correlation between what Pinterest users share about food and cancer, and the scientific evidence on the relationship between cancer and consumption of specific foods. In order to do this, we used PubMed and studied 75 Pins published in Pinterest. We identified about 80,000 scientific articles in PubMed (US National Library of Medicine run by the National Institutes of Health)about cancer and those foods posted on Pinterest in Brazil. We observed that the Pins published in Pinterest have some relation to scientific information despite not being possible to establish if the correlation with the particular food and cancer involved prevention, cure or treatment.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

Molecular Biology Open Access Pubmed Word and Sentence Representations

<p><strong>Natural Language&nbsp;Embeddings about Molecular Biology</strong></p> <p>This dataset is concerned with developing a tailored training data set for word and sentence embedding based on biomedical text that has some component associated with molecular work (as opposed to the other range of work indexed in PubMed like non molecular clinical work, studies of human behavior, etc).&nbsp;&nbsp;</p> <p><strong>Raw Data</strong></p> <p>In order to develop natural language&nbsp;embeddings (for words and sentences), we queried PMC and MEDLINE for molecular papers only by using high-level MeSH terms to restrict interest to papers with a molecular focus. We used the following MeSH terms:</p> <ul> <li>Cells [A11]</li> <li>Multiprotein Complexes [D05.500]</li> <li>Protein Aggregates [D05.875]</li> <li>Hormones [D06]</li> <li>Enzymes and Coenzymes [D08]</li> <li>Carbohydrates [D08]</li> <li>Lipids [D10]</li> <li>Amino Acids, Peptides and Proteins [D12]</li> <li>Nucleic Acids, Nucleotides and Nucleosides [D13]</li> <li>Biological Factors [D23]</li> <li>Pharmaceutical Preparations [D26]</li> <li>Metabolism [G03]</li> <li>Genetic Phenomena [G06]</li> </ul> <p>Queries for these terms use the following string:</p> <blockquote> <p>&quot;cells&quot;[MeSH Terms] OR &quot;Multiprotein Complexes&quot;[mh] OR &quot;Protein Aggregates&quot;[mh] OR &quot;Hormones, Hormone Substitutes, and Hormone Antagonists&quot;[mh] OR &quot;Enzymes and Coenzymes&quot;[mh] OR &quot;Carbohydrates&quot;[mh] OR &quot;Lipids&quot;[mh] OR &quot;Amino Acids, Peptides, and Proteins&quot;[mh] OR &quot;Nucleic Acids, Nucleotides, and Nucleosides&quot;[mh] OR &quot;Biological Factors&quot;[mh] OR &quot;Pharmaceutical Preparations&quot;[mh] OR &quot;Metabolism&quot;[mh] OR &quot;Cell Physiological Phenomena&quot;[mh] OR &quot;Genetic Phenomena&quot;[mh]</p> </blockquote> <p>PubMed returns 11,447,521 abstracts. PMC returns, 1,720,266 documents, 509,722 of these are open access. We downloaded, parsed and concatenated 403,825 PMC open access documents into a single file `molecular_oa_pmc.tsv`. This is a 33GB TSV file with the following columns:</p> <ul> <li>File:Paragraph - a unique identifier for each paragraph</li> <li>SentenceId - the local number of the sentence in the document</li> <li>Sentence Text - tokenized text of the sentence (based on&nbsp;<a href="https://github.com/ClearTK/cleartk/blob/master/cleartk-token/src/main/java/org/cleartk/token/tokenizer/TokenAnnotator.java">ClearTk&#39;s TokenAnnotator.java</a>)</li> <li>Codes -&nbsp;<code>exLink</code>&nbsp;for the presence of a citation,&nbsp;<code>inLink</code>&nbsp;for the presence of link to a Figure</li> <li>Figures - Figure codes</li> <li>Headings - High level section of the paper</li> <li>Offset_Begin - offset of the start of the sentence within the paper</li> <li>Offset_End - offset of the start of the sentence within the paper</li> </ul> <p>We repeated the same process for PubMed abstracts to generate a 3.6G file (`molecular_oa_medline.tsv`) with three columns:</p> <ol> <li>Pubmed ID</li> <li>A Boolean value indicating whether the article is a review</li> <li>Text</li> </ol> <p>We concatenated the text columns of these two files into a single 30GB file (`molecular_oa.txt`) where each line is a single sentence and the text is fully tokenized.&nbsp;</p> <p>These three files are archived in&nbsp;`molecular_oa_raw_text.tar.gz`.</p> <p><strong>Fasttext Embedding</strong></p> <p>We trained a fasttext model on the raw training data (https://fasttext.cc/) using the standard `skipgram` parameter. A gzipped copy of&nbsp; the word embeddings is included in `fasttext.model.vec.gz`&nbsp;</p> <p><strong>GloVe Embedding</strong></p> <p>We trained GloVe&nbsp;models on the raw training data (https://nlp.stanford.edu/projects/glove/). A gzipped copy of&nbsp;the best performing&nbsp;word embeddings is included in `bio_GloVe_300.tar.gz`&nbsp;</p>

opencc-by-4.0Jul 2018View details →
zenodo40/100

DrugProt Complete PubMed Knowledge Graph

<p><strong>DrugProt Complete PubMed Knowledge Graph</strong></p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of&nbsp;DrugProt task at BioCreative VII: data and&nbsp;methods for&nbsp;large-scale text mining and&nbsp;knowledge graph generation of&nbsp;heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, &nbsp;title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, &nbsp;author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, &nbsp;journal={Database}, &nbsp;volume={2023}, &nbsp;pages={baad080}, &nbsp;year={2023}, &nbsp;publisher={Oxford University Press UK} }</i></p></blockquote><p><strong>Description</strong></p><p>This dataset contains a knowledge graph built from PubMed dump abstracts (December 2021). A NER system has been applied to each of the abstracts to extract mentions of type "CHEMICAL" and "GENE", as well as a RE system to detect existing relations between these mentions such as ACTIVATOR, INHIBITOR, AGONIST or PRODUCT_OF, among others (see article for a full list of relations considered).</p><p>Given the volume of the dataset, the repository is divided into 1114 folders. Each of these folders contains a chunk of PubMed abstracts, entities and relationships, divided into the following 3 files:</p><ul><li><i>abstracts.tsv</i>: Tabular file in which each line represents a pubmed document. The file has 3 columns:<ul><li>Pubmed_id: Numerical identifier of the document in PubMed</li><li>Title: Title of the document</li><li>Abstract: Abstract text.</li></ul></li><li>entities.tsv: List of the entities extracted from the abstracts. Each line represents an extracted entity, and has 5 columns:<ul><li>Pubmed_id: Numerical identifier of the document in PubMed</li><li>Mention_id: Numerical identifier of the mention in the document.</li><li>Entity_type: Type of mention. It can be CHEMICAL or GENE.</li><li>Span_ini: Index of the first character of the annotated span in the text</li><li>Span_end: Index of the first character after the annotated span.</li><li>Span: Text span of the annotation</li></ul></li><li>relations.tsv: File of existing relations between entities. Each line represents a relationship, and has the following fields:<ul><li>Pubmed_id:&nbsp;Numerical identifier of the document in PubMed</li><li>Relation_type:&nbsp;DrugProt relation type among entities/arguments.</li><li>Arg1: Mention of CHEMICAL</li><li>Arg2: Mention of GENE</li></ul></li></ul><p>&nbsp;</p><p><strong>Files:</strong></p><ul><li>drugprot-silver-standard-kg.zip : Folders with the files previously explained</li></ul><p>&nbsp;</p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://doi.org/10.5281/zenodo.4955410">DrugProt corpus</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li><li><a href="https://doi.org/10.5281/zenodo.8246229">DrugProt Complete PubMed Knowledge Graph</a><br>&nbsp;</li></ul>

opencc-by-4.0Oct 2022View details →
dryad40/100

Distribution of trial registry numbers within full-text PubMed Central - full dataset of discovered links

Open the record for dataset details and reuse information.

publicFeb 2025View details →
zenodo36/100

Europe Pubmed Central Lite metadata (JSON object stream)

<p>The raw JSON object stream associated with&nbsp;https://zenodo.org/deposit/107781/</p>

opencc-zeroApr 2016View details →
zenodo36/100

Europe PubMed Central Lite Metadata search index

<p>Metadata of all&nbsp;1,264,182 Open Access articles from the Europe PubMed Central Lite dataset, parsed from the fulltext XML (ftp://ftp.ebi.ac.uk/pub/databases/pmc/oa/) and converted to bibJSON (http://okfnlabs.org/bibjson/), then indexed with search-index (https://github.com/fergiemcdowall/search-index), and finally exported as a snapshot.</p> <p><span>An example record is included in the `additional notes` field below.</span></p> <p><span>To use this, you&#39;ll need to import the snapshot into a search-index instance:&nbsp;</span>https://github.com/fergiemcdowall/search-index/blob/master/doc/replicate.md.</p>

opencc-zeroApr 2016View details →
zenodo36/100

Biology and Computational Biology Papers in Pubmed, 1997-2014

<p>Downloaded from Pubmed on 12 February, 2016</p> <p>Bio:<br /> (&quot;Biology&quot;[Mesh]) NOT (Review[ptyp] OR Comment[ptyp] OR Editorial[ptyp] OR Letter[ptyp] OR Case Reports[ptyp] OR News[ptyp] OR &quot;Biography&quot; [Publication Type]) AND (&quot;1997/01/01&quot;[PDAT] : &quot;2014/12/31&quot;[PDAT]) AND english[language]</p> <p>Comp:<br /> (&quot;Computational Biology&quot;[Majr]) NOT (Review[ptyp] OR Comment[ptyp] OR Editorial[ptyp] OR Letter[ptyp] OR Case Reports[ptyp] OR News[ptyp] OR &quot;Biography&quot; [Publication Type]) AND (&quot;1997/01/01&quot;[PDAT] : &quot;2014/12/31&quot;[PDAT]) AND english[language]</p>

opencc-zeroJul 2016View details →
zenodo36/100

Pre-processed PubMed data for a study of coauthorship

<p>This dataset was collected from the PubMed portal to MEDLINE and other repositories of biomedical research (https://www.ncbi.nlm.nih.gov/pubmed/). Analysis of the dataset led to the paper "Effects of research complexity and competition on the incidence and growth of coauthorship in biomedicine", published in PLOS One (http://journals.plos.org/plosone/article?id=10.1371/journal.pone.0173444). The raw data were pre-processed using the script "clean.r" in the project directory on GitHub (https://github.com/corybrunson/coauthor) to obtain the file presented here.</p> <p>The dataset is formatted as a data table (https://cran.r-project.org/web/packages/data.table/index.html), a class of data frame in R, and saved as a .RData file, which can be loaded into an R session via `load("path/to/dataset/pmDat.RData")`. The fields are as follows:</p> <ul> <li>`pmid` - the unique publication identifier (PMID) used by PubMed</li> <li>`jid` - the unique journal identifier used by PubMed</li> <li>`issn` - the (print) ISSN of the journal</li> <li>`ym` - the month and year of publication</li> <li>`nau` - the number of authors credited by the publication (up to any limits imposed by PubMed, and counting each author collective as a single author)</li> <li>`cau` - whether any corporate author was credited</li> <li>`rev` - whether the publication was tagged as a review</li> <li>`trial` - whether the publication was tagged as a clinical trial</li> <li>`npmt` - the number of MeSH terms assigned to the publication that were flagged as "major" topics</li> <li>`nmh` - the number of top-level MeSH headings assigned to the publication</li> <li>`supp` - whether the publication was tagged as having received financial support</li> <li>`ng` - the number of grants acknowledged by the publication</li> <li>`co` - the country in which the journal was published</li> </ul> <p>Note that the field values for any publication can be validated by searching for the PMID in PubMed.</p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

Biotea Linked Data for Open Access Full text PubMed Central

<p>this is the dataset from a-b</p>

opencc-by-4.0Aug 2017View details →
zenodo36/100

biotea, RDf for pubmed central, c-h dataset

<p>C-H dataset of the RDF for pubmed central </p>

opencc-by-4.0Aug 2017View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record