Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

195

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

195 results for “biomedical”

Learn how ShareScore rates datasets ↗
zenodo36/100

WMT'16 Biomedical Translation Task - Scielo monolingual datasets

<p>Monolingual data from Scielo for the Biomedical Translation Task in the First Conference on Machine Translation (WMT 16) (http://www.statmt.org/wmt16/biomedical-translation-task.html).</p> <p>It contains monolingual data for en, es, pt and fr.</p> <p>The documents were derived from the Scielo database (https://scielo.org/en/).</p>

opencc-by-4.0Feb 2016View details →
zenodo36/100

WMT'16 Biomedical Translation Task - Scielo parallel datasets

<p>Parallel data from Scielo for the Biomedical Translation Task in the First Conference on Machine Translation (WMT 16) (http://www.statmt.org/wmt16/biomedical-translation-task.html).</p> <p>It contains parallel data for es/en, fr/en and pt/en.</p> <p>The documents were derived from the Scielo database (https://scielo.org/en/).</p>

opencc-by-4.0Jan 2016View details →
zenodo36/100

3D-Printed Encapsulation of Thin-Film Transducers for Reliable Force Measurement in Biomedical Applications

<p>Data csv</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Extra material for Biomedical Device and Sensors book

<p>Biomedical Device and Sensors book</p> <p>Chapter 6, Personal Practice 1, data: Chap6-exercice 1.csv</p> <p>Chapter 6, Personal Practice 2, data: Chap6-exercice 2.csv</p> <p>Chapter 7, example of Adruino code (under CC BY-NC-SA): MPU6050Madwick_V2.zip</p>

opencc-by-4.0Oct 2023View details →
ClinicalTrials.gov36/100

Community Solutions to Adolescent Research Consent - Minor Consent for Biomedical HIV Research

ClinicalTrials.gov study NCT05371327. IPD Sharing: YES. Countries: 1. Publications: 2.

controlledIPD-YESFeb 2026View details →
ClinicalTrials.gov36/100

COVID-19 MP Biomedicals Rapid SARS-CoV-2 Antigen Test Usability

ClinicalTrials.gov study NCT05584189. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
ClinicalTrials.gov36/100

COVID-19 MP Biomedicals SARS-CoV-2 Ag OTC: Clinical Evaluation

ClinicalTrials.gov study NCT05584176. IPD Sharing: NO. Countries: 1. Publications: 3.

closedIPD-NOFeb 2026View details →
zenodo32/100

MiRoR14 - P2 - A survey exploring biomedical editors' perceptions of editorial interventions to improve adherence to reporting guidelines

<p>Survey questionnaire, survey answers and description of the barriers and facilitators for the interventions included in the project &quot;A survey exploring biomedical editors&rsquo; perceptions of&nbsp;editorial interventions to improve adherence to reporting&nbsp;guidelines&quot;</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Biomedical associations predictions

<p>Resulting predictions</p>

opencc-by-4.0Apr 2020View details →
zenodo32/100

Multimodal Biomedical Dataset for Evaluating Registration Methods (patches from TMA Cores)

<p>The dataset consists of 206 aligned Bright-Field (BF) and Second-Harmonic Generation (SHG) images. All images are of the same size of 834x834 pixels. The training set contains 40 image pairs. A validation set 1 of 25 image pairs for tuning the hyperparameters for the training. A validation set 2 of&nbsp;7 image pairs for tuning the registration method. The test set for evaluation: 134 image pairs (times two).</p> <p>For each image pair in the test set, a random rotation up to +/-30 degrees and random translations in x and y for up to 100px were applied. For each image pair that was transformed for the test set, we apply the transformation to both modalities and keep their corresponding reference images as well as transformed images, in order to allow for registration using each modality as a reference and the other as a floating image. The metadata.csv file includes the coordinates of the four corners of the reference image as well as the coordinates of the corners of the transformed image after displacement. The origin of this coordinate system is the upper left corner of the reference image, given by (0,0).</p> <p>Example of naming convention in the test set:</p> <ul> <li>R_1B_F8_BF.tif: reference image in Bright-Field</li> <li>R_1B_F8_SHG.tif: reference image in Second-Harmonic Generation</li> <li>T_1B_F8_BF.tif: transformed image in Bright-Field</li> <li>T_1B_F8_SHG.tif: transformed image in Second-Harmonic Generation.</li> </ul> <p>Where the transformations applied to T_1B_F8_BF.tif and T_1B_F8_SHG.tif are identical and hence the coordinates of the corners in metadata.csv are valid for both image pairs.</p> <p>The metadata.csv file contains the following columns:</p> <ul> <li>Filename: identifier for image pair</li> <li>X1_Ref: x-coordinate of upper left corner of reference image</li> <li>Y1_Ref: y-coordinate of upper left corner of reference image</li> <li>X2_Ref: x-coordinate of lower left corner of reference image</li> <li>Y2_Ref: y-coordinate of lower left corner of reference image</li> <li>X3_Ref: x-coordinate of upper right corner of reference image</li> <li>Y3_Ref: y-coordinate of upper right corner of reference image</li> <li>X4_Ref: x-coordinate of lower right corner of reference image</li> <li>Y4_Ref: y-coordinate of lower right corner of reference image</li> <li>X1_Trans: x-coordinate of upper left corner of transformed image</li> <li>Y1_Trans: y-coordinate of upper left corner of transformed image</li> <li>X2_Trans: x-coordinate of lower left corner of transformed image</li> <li>Y2_Trans: y-coordinate of lower left corner of transformed image</li> <li>X3_Trans: x-coordinate of upper right corner of transformed image</li> <li>Y3_Trans: y-coordinate of upper right corner of transformed image</li> <li>X4_Trans: x-coordinate of lower right corner of transformed image</li> <li>Y4_Trans: y-coordinate of lower right corner of transformed image</li> <li>Displacement: Mean Euclidean distance between reference corner points and transformed corner points</li> </ul> <p>The data set was originally produced by the authors of&nbsp;<em>Aligned Collagen Is a Prognostic Signature for Survival in Human Breast Carcinoma</em> (<a href="https://www.sciencedirect.com/science/article/pii/S0002944010002336">https://www.sciencedirect.com/science/article/pii/S0002944010002336</a>). The registered and non-registered sub-image pairs in this data set were created by Johan &Ouml;fverstedt and Elisabeth Wetzer.</p>

opencc-by-4.0Jun 2020View details →
zenodo32/100

Biomedical preprints per month, by source and as a fraction of total literature

<p>This is a snapshot of a <a href="https://docs.google.com/spreadsheets/d/1c1WzYIoLhzY09AkwyhSVDtan3oYV8qrHa3WgPHAeVVM/edit?usp=sharing">Google sheet</a>&nbsp;containing counts of biomedical preprints by source and as a fraction of the total biomedical literature <strong>through 2020-06.</strong></p> <p>Note: this does not yet include &gt;6,000 preprints on relevant OSF platforms.</p> <p><strong>Data sources</strong></p> <p>arXiv q-bio, PeerJ Preprints, and bioRxiv counts through&nbsp;2018-12 were sourced from Jordan Anaya&#39;s <a href="https://github.com/OmnesRes/prepub/tree/master/analyses">PrePubMed</a>. Following that, arXiv q-bio counts were sourced from <a href="https://arxiv.org/year/q-bio/19">arXiv statistics</a>, and PeerJ Preprints and bioRxiv counts were taken from searches of <a href="https://europepmc.org/">EuropePMC</a>.</p> <p>All counts from F1000 &amp; Open Research platforms, preprints.org, ChemRxiv, and medRxiv were taken from searches of <a href="https://europepmc.org/">EuropePMC</a>.</p> <p>Counts for Research Square are derived from Crossref through 2020-03, thereafter from EuropePMC.&nbsp;Counts for the Lancet and Sneak Peek were taken from web searches.</p> <p>Total biomedical literature is from PubMed.</p>

opencc-by-4.0Feb 2020View details →
zenodo32/100

Biomedical ELECTRA based deep language representation models for biomedical text mining.

<p>The gzipped tar file contains two biomedical language representation models based on ELECTRA&nbsp; (Clark et al., 2020)&nbsp;&nbsp;deep transformers architecture to be used for down-stream biomedical text mining tasks.&nbsp;</p> <p>Bio-ELECTRA is pre-trained from scratch on PubMed abstracts for 1.8 million steps. Bio-ELECTRA++ is the further pre-trained version of Bio-ELECTRA trained on a corpus of open access full papers from PubMed.</p>

opencc-by-4.0Aug 2020View details →
dryad32/100

Data from: How can we get close to zero?: the potential contribution of biomedical prevention and the investment framework towards an effective response to HIV

Background: In 2011 an Investment Framework was proposed that described how the scale-up of key HIV interventions could dramatically reduce new HIV infections and deaths in low and middle income countries by 2015. This framework included ambitious coverage goals for prevention and treatment services resulting in a reduction of new HIV infections by more than half. However, it also estimated a leveling in the number of new infections at about 1 million annually after 2015. Methods: We modeled how the response to AIDS can be further expanded by scaling up antiretroviral treatment (ART) within the framework provided by the 2013 WHO treatment guidelines. We further explored the potential contributions of new prevention technologies: 'Test and Treat', pre-exposure prophylaxis and an HIV vaccine. Findings: Immediate aggressive scale up of existing approaches including the 2013 WHO guidelines could reduce new infections by 80%. A 'Test and Treat' approach could further reduce new infections. This could be further enhanced by a future highly effective pre-exposure prophylaxis and an HIV vaccine, so that a combination of all four approaches could reduce new infections to as low as 80,000 per year by 2050 and annual AIDS deaths to 260,000. Interpretation: In a set of ambitious scenarios, we find that immediate implementation of the 2013 WHO antiretroviral therapy guidelines could reduce new HIV infections by 80%. Further reductions may be achieved by moving to a 'Test and Treat' approach, and eventually by adding a highly effective pre-exposure prophylaxis and an HIV vaccine, if they become available.

opencc-zeroDec 2013View details →
zenodo32/100

Learning the structure of biomedical relationships from unstructured text (Part II)

<p>These files contain the seed sets and corresponding test sets of drug-gene pairs reflecting pharmacogenomic (PGx) and drug-target relationships that were used to evaluate EBC in the PLoS Comp Bio paper. Unfortunately, they didn&#39;t make it into the first upload for this paper. There are four zipped directories in this upload:</p> <p>data-drug-gene (PGx relationships, dense matrix)</p> <p>data-drug-target (drug-target relationships, dense matrix)</p> <p>data-drug-gene-fullmatrix (PGx relationships, sparse matrix)</p> <p>data-drug-target-fullmatrix (drug-target relationships, sparse matrix)</p>

opencc-zeroJul 2015View details →
zenodo32/100

Large-scale semantic indexing of Spanish biomedical literature using contrastive transfer learning

Open the record for dataset details and reuse information.

opencc-by-4.0Sep 2024View details →
zenodo32/100

A Framework for Automated Construction of Heterogeneous Large-Scale Biomedical Knowledge Graphs (Recorded Talk)

<p>This entry contains the&nbsp;recording of the presentation that was presented at the 2020 Intelligent Systems for Molecular Biology as part of the Bio-Ontologies COSI (https://www.iscb.org/ismb2020).</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Nanoprobes for Biomedical Imaging with Tunable Near-Infrared Optical Properties Obtained via Green Synthesis

<p>Dataset of:&nbsp;10.1002/adpr.202100260</p>

opencc-by-4.0Nov 2021View details →
zenodo32/100

Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications.

<p>The datasets used in the publications titled &quot;Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications&quot;</p>

opencc-by-4.0Feb 2022View details →
zenodo32/100

Dataser from paper "Proliferation of osteoblast precursor cells on the surface of TiO2 nanowires anodically grown on a -type biomedical titanium alloy"

<p>Dataser from paper &quot;Proliferation of osteoblast precursor cells on the surface of TiO2 nanowires anodically grown on a -type biomedical titanium alloy&quot;:</p> <p>-&nbsp;<strong>Contact Angle: </strong>Images and measurements.</p> <p><strong>- Fluorescence Microscopy Images:</strong>&nbsp;Images.</p> <p><strong>- MTT and pixel counting:</strong>&nbsp;Measurements.</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo32/100

Data for "RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature"

<div> <p><strong>RegulaTome corpus</strong>: this <a href="../api/records/10808330/files/RegulaTome-corpus.tar.gz/content" target="_blank" rel="noopener">file</a> contains the RegulaTome corpus in&nbsp;<a href="https://brat.nlplab.org/">BRAT</a> format. The directory&nbsp;<strong>"splits" </strong>has the corpus split based on the train/dev/test used for the training of the relation extraction system</p> <p><strong>RegulaTome annodoc</strong>: The annotation guidelines along with the annotation configuration files for BRAT are provided in <a href="../api/records/10808330/files/annodoc+config.tar.gz/content" target="_blank" rel="noopener">annodoc+config.tar.gz</a>. The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/regulatome-annodoc/&nbsp;</a></p> <p>The tagger software can be found here:&nbsp;<a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a>. The command used to run tagger before large-scale execution of the RE system is:</p> <p><code>gzip -cd `ls -1 pmc/*.en.merged.filtered.tsv.gz` `ls -1r pubmed/*.tsv.gz` | cat dictionary/excluded_documents.txt - | tagger/tagcorpus --threads=16 --autodetect --types=dictionary/curated_types.tsv --entities=dictionary/all_entities.tsv --names=dictionary/all_names_textmining.tsv --groups=dictionary/all_groups.tsv --stopwords=dictionary/all_global.tsv --local-stopwords=dictionary/all_local.tsv --type-pairs=dictionary/all_type_pairs.tsv --out-matches=all_matches.tsv</code></p> <p><strong>Input documents </strong>for large-scale execution, which is done on entire <a href="https://a3s.fi/March-2024-PubMed/PubMed_20230314.tar.gz" target="_blank" rel="noopener">PubMed</a> (as of March 2024) and <a href="https://a3s.fi/Jan-2024-documents/PMC_Nov_23.tar.gz" target="_blank" rel="noopener">PMC Open Access</a> (as of November 2023) articles in BioC format. The files are converted to a <a href="https://a3s.fi/March-2024-PubMed/all_documents.tsv" target="_blank" rel="noopener">tab-delimited format&nbsp;</a>to be compatible with the RE system input (see below).</p> <p><strong>Input dictionary files</strong>: all the files necessary to execute the command above are available in&nbsp;<a href="../api/records/10808330/files/tagger_dictionary_files.tar.gz/content" target="_blank" rel="noopener">tagger_dictionary_files.tar.gz&nbsp;</a></p> <p><strong>Tagger output</strong>: we filter the results of the tagger run down to gene/protein hits, and documents with more than 1 hit (since we are doing relation extraction) before feeding it to our RE system. The filtered output is available in <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a></p> <p><strong>Relation extraction system input</strong>:&nbsp;<a href="../api/records/10808330/files/combined_input_for_re.tar.gz/content" target="_blank" rel="noopener">combined_input_for_re.tar.gz</a>: these are the directories with all the .ann and .txt files used as input for the large-scale execution of the relation extraction pipeline. The files are generated from the tagger tsv output (see above, <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a>) using the&nbsp;<a href="https://github.com/spyysalo/string-db-tools/blob/main/tagger2standoff.py">tagger2standoff.py</a> script from the <a href="https://github.com/spyysalo/string-db-tools/">string-db-tools</a> repository.</p> <p><strong>Relation extraction models</strong>. The Transformer-based model used for large-scale relation extraction and prediction on the test set is at&nbsp;<a href="../api/records/10808330/files/relation_extraction_multi-label-best_model.tar.gz/content" target="_blank" rel="noopener">relation_extraction_multi-label-best_model.tar.gz</a></p> <p>The pre-trained RoBERTa model on PubMed and PMC and MIMIC-III with a BPE Vocab learned from PubMed (RoBERTa-large-PM-M3-Voc), which is used by our system is available <a href="https://github.com/facebookresearch/bio-lm/blob/main/README.md">here</a>.</p> <p><strong>Relation extraction system output</strong>: the tab-delimited outputs of the relation extraction system are found at&nbsp;<a href="https://a3s.fi/regulatome-ls/large_scale_relation_extraction_results.tar.gz" target="_blank" rel="noopener">large_scale_relation_extraction_results.tar.gz </a><strong>!!!ATTENTION this file is approximately 1TB in size, so make sure you have enough space to download it on your machine!!!</strong></p> <p>The relation extraction system output files have 86 columns: PMID, Entity BRAT ID1, Entity BRAT ID2, and scores per class produced by the relation extraction model. Each file has a header to denote which score is in which column.</p> </div>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record