Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

195

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

195 results for “biomedical”

Learn how ShareScore rates datasets ↗
dryad40/100

CZ Software Mentions: A large dataset of software mentions in the biomedical literature

Open the record for dataset details and reuse information.

publicSep 2022View details →
dryad40/100

CZ Software Mentions: A large dataset of software mentions in the biomedical literature - Expanded 2024

Open the record for dataset details and reuse information.

publicNov 2024View details →
zenodo36/100

"MiRoR13-P1- A scoping review on the roles and tasks of peer reviewers in the manuscript review process in biomedical journals"

<p>This dataset is related to the publication &quot;A scoping review on the roles and tasks of peer reviewers in the manuscript review process in biomedical journals&quot;. All additional data can be found under &#39;Additional files&#39; in the publication</p>

opencc-by-4.0Feb 2020View details →
zenodo36/100

Biomedical Journal Data Sharing Policies

<p>Raw data of data sharing policies in over 300 journals, supporting the article currently under review: "Reproducible and reusable research: Are journal data sharing policies meeting the mark?".  </p> <p>Raw data and analysis of data sharing policies of 318 biomedical journals. The study authors manually reviewed the author instructions and editorial policies to analyze the each journal's data sharing requirements and characteristics. The data sharing policies were ranked using a rubric to determine if data sharing was required, recommended, or not addressed at all. The data sharing method and licensing recommendations were examined, as well any mention of reproducibility or similar concepts. The data was analyzed for patterns relating to publishing volume, Journal Impact Factor, and the publishing model (open access or subscription) of each journal.</p> <p>We evaluated journals included in Thomson Reuter’s InCites 2013 Journal Citations Reports (JCR) classified within the following World of Science schema categories: Biochemistry and Molecular Biology, Biology, Cell Biology, Crystallography, Developmental Biology, Biomedical Engineering, Immunology, Medical Informatics, Microbiology, Microscopy, Multidisciplinary Sciences, and Neurosciences. These categories were selected to capture the journals publishing the majority of peer-reviewed biomedical research. The original data pull included 1,166 journals, collectively publishing 213,449 articles. We filtered this list to the journals in the top quartiles by impact factor (IF) or number of articles published 2013. Additionally, the list was manually reviewed to exclude short report and review journals, and titles determined to be outside the fields of basic medical science or clinical research. The final study set included 318 journals, which published 130,330 articles in 2013. The study set represented 27% of the original Journal Citation Report list and 61% of the original citable articles. Prior to our analysis, the 2014 Journal Citations Reports was released. After our initial analyses and first preprint submission, the 2015 Journal Citations Reports was released. While we did not use the 2014 or 2015 data to amend the journals in the study set, we did employ data from all three reports in our analyses. In our data pull from JCR, we included the journal title, International Standard Serial Number (ISSN), the total citable items for 2013, 2014, and 2015, the total citations to the journal for 2013/14/15, the impact factors for 2013/14/15, and the publisher.</p>

opencc-by-4.0Mar 2017View details →
zenodo36/100

WMT'17 Biomedical Translation Task - Scielo test and gold sets

<p>Test and gold datasets for the Biomedical Translation Task in the Second Conference on Machine Translation (WMT 17) (http://www.statmt.org/wmt17/biomedical-translation-task.html).</p> <p>It includes test files, gold files and automatic alignment files using the GMA tool.</p> <p>Documents were derived from the Scielo database (http://scielo.org).</p>

opencc-by-sa-4.0Aug 2017View details →
zenodo36/100

Divergent Characteristics of Biomedical Research across Publication Types: A Quantitative Analysis on the Aging-related Research

<p>Attached in pdf here.</p>

opencc-by-4.0Dec 2023View details →
zenodo36/100

Code & data of the paper "Data sampling via Active Learning in Cartesian Genetic Programming for Biomedical Data"

<p>Code &amp; data of the paper "Data sampling via Active Learning in Cartesian Genetic Programming for Biomedical Data"</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Pathway2Text: Dataset for Biomedical Pathway Description Generation

<p>This is the dataset of the&nbsp;NAACL 2022 paper:</p> <p>Pathway2Text: Dataset and Method for Biomedical Pathway Description Generation.</p> <p>This dataset contains 2,367 pairs of biomedical pathways and textual descriptions. It can be used for automatic pathway description generation. In our paper, we showed it is also appropriate for Text2Graph and BioNER.</p> <p>Read readme.pdf for detaild information.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

ClinSpEn-CT Data: Parallel English-Spanish Biomedical Terminology

<p><strong>UPDATE August 22nd 2022: </strong>The data in this repository has been merged with the rest of the ClinSpEn data, you may access it here: https://doi.org/10.5281/zenodo.6497350</p> <p>This repository contains the sample, test and background data for the ClinSpEn-Clinical Terms sub-track. The direction of this sub-track is ES&gt;EN.</p> <p>ClinSpEn is part of the Biomedical WMT 2022 shared task, having the aim to promote the development and evaluation of machine translation systems adapted to the medical domain with three highly relevant sub-tracks: clinical cases, medical controlled vocabularies/ontologies, and clinical terms and entities extracted from medical content.</p> <p>The terms were directly extracted from medical literature and clinical records, with particular focus on diseases, symptoms, findings, procedures and professions and translated and revised by professional medical translators.</p> <p>The sample set contains 7 000 terms as a tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms.</p> <p>The test and background data is made up of a TSV file with two columns: term number and Spanish term.</p> <p>Related Links:</p> <p><strong>- Sub-track website with more information: </strong><a href="https://temu.bsc.es/clinspen/">https://temu.bsc.es/clinspen/</a></p> <p><strong>- WMT website: </strong><a href="https://www.statmt.org/wmt22/">https://www.statmt.org/wmt22/</a></p> <p><strong>- CodaLab: </strong><a href="https://codalab.lisn.upsaclay.fr/competitions/6696">https://codalab.lisn.upsaclay.fr/competitions/6696</a></p> <p>&nbsp;</p> <p><strong>- ClinSpEn-CC (Clinical Cases):</strong> <a href="https://doi.org/10.5281/zenodo.6497350">https://doi.org/10.5281/zenodo.6497350</a></p> <p>&nbsp;</p> <p><strong>- ClinSpEn-CT (Clinical Terms): </strong><a href="https://doi.org/10.5281/zenodo.6497372">https://doi.org/10.5281/zenodo.6497372</a></p> <p><strong>- ClinSpEn-OC (Ontology Concepts): </strong><a href="https://doi.org/10.5281/zenodo.6497388">https://doi.org/10.5281/zenodo.6497388</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

Biomedical Entities and Relations on Spanish Clinical Case Corpus: BERSCCC

<p>This first version of a spanish corpus contains 200 clinical reports annotated with biomedical entities and semantic relations.&nbsp;<br> These reports belong to the Spanish Clinical Case Corpus (SPACCC) (https://doi.org/10.5281/zenodo.2560316)<br> and each of them has been annotated by three persons that work in the medicine, biomolecular or pharmaceutic area.</p> <p>The annotators had to identify the following thirteen types of entities in the spanish lenguage: Enfermedad/S&iacute;ndrome, Gen, Parte del cuerpo/&Oacute;rgano, Gl&uacute;cido, Procedimiento de Diagn&oacute;stico, Prote&iacute;na, Procedimiento Terape&uacute;tico, S&iacute;ntoma/Signo, Sustancia Farmacol&oacute;gica, L&iacute;pido, Organismo, Qu&iacute;mico Org&aacute;nico and Abreviatura/Sigla/Alias.<br> And the next eight semantic relations: Analiza, Altera, Causa, Diagnostica, Manifestaci&oacute;n de, Produce, Trata and Refiere a.&nbsp;</p> <p>Finally there were identified 6,636 biomedical entities (37,081 mentions) and 4,864 semantic relations (7,622 mentions).</p> <p>These resources are freely distributed under a Creative Commons Attribution 4.0 International License.<br> The scripts used to create this corpus can be found at: https://github.com/drugs4covid/bio-corpora</p> <p>Author: Luc&iacute;a S&aacute;nchez Gonz&aacute;lez, Ontology Engineering Group, Universidad Polit&eacute;cnica de Madrid.</p> <p>Supervisors:&nbsp;</p> <p>- Carlos Badenes Olmedo, Ontology Engineering Group, Universidad Polit&eacute;cnica de Madrid.</p> <p>- Mar&iacute;a Poveda Villal&oacute;n, Ontology Engineering Group, Universidad Polit&eacute;cnica de Madrid.</p> <p>Project Member: &Oacute;scar Corcho Garc&iacute;a,&nbsp;Ontology Engineering Group, Universidad Polit&eacute;cnica de Madrid.</p> <p>&nbsp;Contact:</p> <p>Luc&iacute;a S&aacute;nchez Gonz&aacute;lez at lu.sanchez@alumnos.upm.es or lusangonz99@gmail.com</p> <p>Acknowledgments to Project DRUGS4COVID++: Servicios de Inteligencia Artificial para la<br> creaci&oacute;n de un grafo de conocimientos sobre f&aacute;rmacos usados en el control cl&iacute;nico de la<br> enfermedad, a partir de la explotaci&oacute;n de grandes corpus de documentaci&oacute;n cient&iacute;fica sobre<br> SARS-COV-2 y COVID-19-AYUDAS FUNDACI&Oacute;N BBVA A EQUIPOS DE<br> INVESTIGACI&Oacute;N CIENT&Iacute;FICA SARS-CoV-2 y COVID-19.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo36/100

Enhanced Kinase Dictionaries associated with KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature

<p>This zip file contains Kinase dictionaries used for annotating documents with KinDER described&nbsp;in the following paper:&nbsp;</p> <p>Dopp, Daniel, Adam Morrone, and Indika Kahanda. &quot;KinDER: A Biocuration Tool for Extracting Kinase Knowledge from Biomedical Literature.&quot;&nbsp;<em>Proceedings of the BioCreative VI Workshop</em>. 2017.</p>

opencc-by-4.0Dec 2017View details →
zenodo36/100

Towards Discovery: An End-to-End System for Uncovering Novel Biomedical Relations

<p>This contains the embeddings as specified in the work: "<strong>Towards Discovery: An End-to-End System for Uncovering Novel Biomedical Relations"</strong>.</p> <p>The embeddings are created from 4 seperate knowledge bases:</p> <ul> <li><a href="https://ftp.expasy.org/databases/cellosaurus/" target="_blank" rel="noopener">Cellosaurus:&nbsp;</a>Concepts</li> <li><a href="https://ctdbase.org/downloads/" target="_blank" rel="noopener">CTD-Disease</a>: Concepts, Definitions, Synonyms</li> <li><a href="https://www.nlm.nih.gov/databases/download/mesh.html" target="_blank" rel="noopener">MeSH</a>: Concepts, Definitions, Synonyms. This includes embeddings for both the main and supplementary files, for Concepts and Definitions.&nbsp;</li> <li><a href="https://ftp.ncbi.nih.gov/gene/" target="_blank" rel="noopener">NCBI-Gene</a>: Concepts. Due ot the large size of the NCBI-cGene corpus we only provide the embeddings for the following top 7 organisms most prevalent in the corpus: <ul> <li>10090: house mouse&nbsp;</li> <li>10116: Norway Rat</li> <li>11676: human immunodeficiency virus (HIV)</li> <li>12814: Respiratory syncytial virus</li> <li>3702: thale cress</li> <li>7955: zebrafish</li> <li>9606: and human</li> </ul> </li> <li><a href="https://ftp.ncbi.nih.gov/pub/taxonomy/" target="_blank" rel="noopener">NCBI-Taxonomy</a>: Names</li> </ul> <p>The embeddings supplied were created using the <a target="_blank" rel="noopener">SapBERT</a> model, for the knowledge bases as of March 2024. The exact files used to generate the embeddings are provided as jsonl files.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo36/100

NED data for the paper Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts

<p>This data repository contains NED data from the paper,&nbsp;<em><a href="https://arxiv.org/abs/2309.01812">Into the Single Cell Multiverse: an End-to-End Dataset for Procedural Knowledge Extraction in Biomedical Texts.</a></em></p> <p>Additional data for the NER classification task can be found here: <a href="../records/10050681">zenodo</a></p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Dataset for Engineering multifunctional dynamic hydrogel for biomedical and tissue regenerative applications

<p><span>Hydrogels have emerged in various biomedical applications, including tissue engineering and medical devices, due to their ability to imitate the natural extracellular matrix (ECM) of tissues. However, conventional static hydrogels lack the ability to dynamically respond to changes in their surroundings to withstand the robust changes of the biophysical microenvironment and to trigger on-demand functionality such as drug release and mechanical change. In contrast, multifunctional dynamic hydrogels can adapt and respond to external stimuli and have drawn great attention in recent studies. It is realized that the integration of nanomaterials into dynamic hydrogels provides numerous functionalities for a great variety of biomedical applications that cannot be achieved by conventional hydrogels. This review article provides a comprehensive overview of recent advances in designing and fabricating dynamic hydrogels for biomedical applications. We describe different types of dynamic hydrogels based on breakable and reversible covalent bonds as well as noncovalent interactions. These mechanisms are described in detail as a useful reference for designing crosslinking strategies that strongly influence the mechanical properties of the hydrogels. We also discuss the use of dynamic hydrogels and their potential benefits. This review further explores different biomedical applications of dynamic nanocomposite hydrogels, including their use in drug delivery, tissue engineering, bioadhesives, wound healing, cancer treatment, and mechanistic study, as well as multiple-scale biomedical applications. Finally, we discuss the challenges and future perspectives of dynamic hydrogels in the field of biomedical engineering, including the integration of diverse technologies.</span></p>

opencc-by-4.0Mar 2024View details →
zenodo36/100

Figure 2 in Biomedical role of L-carnitine in several organ systems, cellular tissues, and COVID-19

Figure 2. The role of L-carnitine in fatty acid metabolism in mitochondria (Wang et al., 2021b).

opencc-by-4.0Dec 2022View details →
zenodo36/100

Training/evaluation data sets and databases for the operation of bio-answerfinder biomedical question answering system

<p>A zip file containing training/evaluation data sets for the&nbsp;bio-answerfinder biomedical question answering system training and evaluation. The zip file also contains SQLite databases for named entity lookups, morphology, nominalizations, acronyms, PubMED trained GLoVe word/phrase embeddings and vocabulary with document frequencies and SciCrunch ontology data for named entities such as proteins, anatomical structures,</p>

opencc-by-4.0Mar 2019View details →
zenodo36/100

Dataset for the article "Biomedical Publishing in Russia: How big is low-quality papers problem?"

<p>The dataset is used in the article "Biomedical Publishing in Russia: How big is low-quality papers problem?" submitted to the Journal of the Medical Library Association : JMLA.</p> <p><br>The dataset contains the list of Russian and international biomedical journals along with their ISSNs, and eISSNs with attached thematic categories from Web of Science and SJR. For each journal-year pair, the number of publications is calculated. The time frame of analysis is 2010-2020.</p>

opencc-by-4.0Sep 2024View details →
zenodo36/100

Two document-concept representations of the biomedical literature

<p>These two datasets represent the biomedical literature (Medline abstracts and PubMedCentral articles) in the &quot;document-concept matrix&quot; format produced by <a href="https://github.com/erwanm/tdc-tools">TDC Tools</a>.&nbsp; These datasets can be used in downstream IR applications such as Literature-Based Discovery.</p> <p>Each of the two datasets corresponds to a specific data extraction method, see details <a href="https://erwanm.github.io/tdc-tools/input-data-format/">here</a> and in the paper linked below.</p> <ul> <li>Paper: <em>pending </em></li> <li>Code: <a href="https://github.com/erwanm/tdc-tools">https://github.com/erwanm/tdc-tools</a> <ul> <li>Documentation:<a href="https://erwanm.github.io/tdc-tools/">https://erwanm.github.io/tdc-tools/</a></li> </ul> </li> </ul> <p><strong>Important:</strong> the raw data from which this data is derived was downloaded from <a href="https://www.nlm.nih.gov/medline/medline_overview.html">Medline</a>, <a href="https://www.ncbi.nlm.nih.gov/pmc/">PubMedCentral</a> and <a href="https://www.ncbi.nlm.nih.gov/research/pubtator/">PubTatorCentral</a>, provided <a href="https://www.nlm.nih.gov/databases/download/terms_and_conditions.html">courtesy of the U.S. National Library of Medicine (NLM)</a>. The data was extracted in January 2021 and do not reflect the most current/accurate data available from NLM. See the github repository above in order to generate similar datasets from up to date data.</p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

Biomedical prototype for human movement data collection and basic collection procedure demonstrations

<p>Video explaining the usage of the 1st prototype of the device produced as well as basic collection procedure example.</p> <p>The video was originally published on&nbsp;<a href="https://www.youtube.com/watch?v=tWhNt0iaYUY">https://www.youtube.com/watch?v=tWhNt0iaYUY</a></p>

opencc-by-4.0Apr 2021View details →
zenodo36/100

Table Recognition Benchmark on Biomedical Literature on Neurological Disorders

<p>The dataset contains 1650 tables from 1164 PMC OA articles in the context of neurological disorders. The tables are structured in the International Conference on Document Analysis and Recognition (ICDAR) format.</p> <p>The additional csv file contains a labeling into 3 different complexity classes in the format:</p> <p>class document_id table_id</p> <p>with the classes being:</p> <p>0 = simple<br> 1 = complicated<br> 2 = complex</p> <p>You may use the scripts from <a href="https://github.com/TimAdams84/pmc-downloader">this repository</a> to bulk download PDF sources from PMC.</p> <p>A script for evaluating results against the groundtruth data can be found <a href="https://github.com/mnamysl/benchmarking_table_recogn">here</a>.</p>

opencc-by-4.0Oct 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record