Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
195
datasets available to search
ShareScore release 0.9.0
Dataset results
195 results for “Biomedical”
WMT'16 Biomedical Translation Task - Scielo monolingual datasets
<p>Monolingual data from Scielo for the Biomedical Translation Task in the First Conference on Machine Translation (WMT 16) (http://www.statmt.org/wmt16/biomedical-translation-task.html).</p> <p>It contains monolingual data for en, es, pt and fr.</p> <p>The documents were derived from the Scielo database (https://scielo.org/en/).</p>
WMT'16 Biomedical Translation Task - Scielo parallel datasets
<p>Parallel data from Scielo for the Biomedical Translation Task in the First Conference on Machine Translation (WMT 16) (http://www.statmt.org/wmt16/biomedical-translation-task.html).</p> <p>It contains parallel data for es/en, fr/en and pt/en.</p> <p>The documents were derived from the Scielo database (https://scielo.org/en/).</p>
3D-Printed Encapsulation of Thin-Film Transducers for Reliable Force Measurement in Biomedical Applications
<p>Data csv</p>
Extra material for Biomedical Device and Sensors book
<p>Biomedical Device and Sensors book</p> <p>Chapter 6, Personal Practice 1, data: Chap6-exercice 1.csv</p> <p>Chapter 6, Personal Practice 2, data: Chap6-exercice 2.csv</p> <p>Chapter 7, example of Adruino code (under CC BY-NC-SA): MPU6050Madwick_V2.zip</p>
Community Solutions to Adolescent Research Consent - Minor Consent for Biomedical HIV Research
ClinicalTrials.gov study NCT05371327. IPD Sharing: YES. Countries: 1. Publications: 2.
COVID-19 MP Biomedicals Rapid SARS-CoV-2 Antigen Test Usability
ClinicalTrials.gov study NCT05584189. IPD Sharing: NO. Countries: 1. Publications: 3.
COVID-19 MP Biomedicals SARS-CoV-2 Ag OTC: Clinical Evaluation
ClinicalTrials.gov study NCT05584176. IPD Sharing: NO. Countries: 1. Publications: 3.
MiRoR14 - P2 - A survey exploring biomedical editors' perceptions of editorial interventions to improve adherence to reporting guidelines
<p>Survey questionnaire, survey answers and description of the barriers and facilitators for the interventions included in the project "A survey exploring biomedical editors’ perceptions of editorial interventions to improve adherence to reporting guidelines"</p>
Biomedical associations predictions
<p>Resulting predictions</p>
Multimodal Biomedical Dataset for Evaluating Registration Methods (patches from TMA Cores)
<p>The dataset consists of 206 aligned Bright-Field (BF) and Second-Harmonic Generation (SHG) images. All images are of the same size of 834x834 pixels. The training set contains 40 image pairs. A validation set 1 of 25 image pairs for tuning the hyperparameters for the training. A validation set 2 of 7 image pairs for tuning the registration method. The test set for evaluation: 134 image pairs (times two).</p> <p>For each image pair in the test set, a random rotation up to +/-30 degrees and random translations in x and y for up to 100px were applied. For each image pair that was transformed for the test set, we apply the transformation to both modalities and keep their corresponding reference images as well as transformed images, in order to allow for registration using each modality as a reference and the other as a floating image. The metadata.csv file includes the coordinates of the four corners of the reference image as well as the coordinates of the corners of the transformed image after displacement. The origin of this coordinate system is the upper left corner of the reference image, given by (0,0).</p> <p>Example of naming convention in the test set:</p> <ul> <li>R_1B_F8_BF.tif: reference image in Bright-Field</li> <li>R_1B_F8_SHG.tif: reference image in Second-Harmonic Generation</li> <li>T_1B_F8_BF.tif: transformed image in Bright-Field</li> <li>T_1B_F8_SHG.tif: transformed image in Second-Harmonic Generation.</li> </ul> <p>Where the transformations applied to T_1B_F8_BF.tif and T_1B_F8_SHG.tif are identical and hence the coordinates of the corners in metadata.csv are valid for both image pairs.</p> <p>The metadata.csv file contains the following columns:</p> <ul> <li>Filename: identifier for image pair</li> <li>X1_Ref: x-coordinate of upper left corner of reference image</li> <li>Y1_Ref: y-coordinate of upper left corner of reference image</li> <li>X2_Ref: x-coordinate of lower left corner of reference image</li> <li>Y2_Ref: y-coordinate of lower left corner of reference image</li> <li>X3_Ref: x-coordinate of upper right corner of reference image</li> <li>Y3_Ref: y-coordinate of upper right corner of reference image</li> <li>X4_Ref: x-coordinate of lower right corner of reference image</li> <li>Y4_Ref: y-coordinate of lower right corner of reference image</li> <li>X1_Trans: x-coordinate of upper left corner of transformed image</li> <li>Y1_Trans: y-coordinate of upper left corner of transformed image</li> <li>X2_Trans: x-coordinate of lower left corner of transformed image</li> <li>Y2_Trans: y-coordinate of lower left corner of transformed image</li> <li>X3_Trans: x-coordinate of upper right corner of transformed image</li> <li>Y3_Trans: y-coordinate of upper right corner of transformed image</li> <li>X4_Trans: x-coordinate of lower right corner of transformed image</li> <li>Y4_Trans: y-coordinate of lower right corner of transformed image</li> <li>Displacement: Mean Euclidean distance between reference corner points and transformed corner points</li> </ul> <p>The data set was originally produced by the authors of <em>Aligned Collagen Is a Prognostic Signature for Survival in Human Breast Carcinoma</em> (<a href="https://www.sciencedirect.com/science/article/pii/S0002944010002336">https://www.sciencedirect.com/science/article/pii/S0002944010002336</a>). The registered and non-registered sub-image pairs in this data set were created by Johan Öfverstedt and Elisabeth Wetzer.</p>
Biomedical preprints per month, by source and as a fraction of total literature
<p>This is a snapshot of a <a href="https://docs.google.com/spreadsheets/d/1c1WzYIoLhzY09AkwyhSVDtan3oYV8qrHa3WgPHAeVVM/edit?usp=sharing">Google sheet</a> containing counts of biomedical preprints by source and as a fraction of the total biomedical literature <strong>through 2020-06.</strong></p> <p>Note: this does not yet include >6,000 preprints on relevant OSF platforms.</p> <p><strong>Data sources</strong></p> <p>arXiv q-bio, PeerJ Preprints, and bioRxiv counts through 2018-12 were sourced from Jordan Anaya's <a href="https://github.com/OmnesRes/prepub/tree/master/analyses">PrePubMed</a>. Following that, arXiv q-bio counts were sourced from <a href="https://arxiv.org/year/q-bio/19">arXiv statistics</a>, and PeerJ Preprints and bioRxiv counts were taken from searches of <a href="https://europepmc.org/">EuropePMC</a>.</p> <p>All counts from F1000 & Open Research platforms, preprints.org, ChemRxiv, and medRxiv were taken from searches of <a href="https://europepmc.org/">EuropePMC</a>.</p> <p>Counts for Research Square are derived from Crossref through 2020-03, thereafter from EuropePMC. Counts for the Lancet and Sneak Peek were taken from web searches.</p> <p>Total biomedical literature is from PubMed.</p>
Biomedical ELECTRA based deep language representation models for biomedical text mining.
<p>The gzipped tar file contains two biomedical language representation models based on ELECTRA (Clark et al., 2020) deep transformers architecture to be used for down-stream biomedical text mining tasks. </p> <p>Bio-ELECTRA is pre-trained from scratch on PubMed abstracts for 1.8 million steps. Bio-ELECTRA++ is the further pre-trained version of Bio-ELECTRA trained on a corpus of open access full papers from PubMed.</p>
Data from: How can we get close to zero?: the potential contribution of biomedical prevention and the investment framework towards an effective response to HIV
Background: In 2011 an Investment Framework was proposed that described how the scale-up of key HIV interventions could dramatically reduce new HIV infections and deaths in low and middle income countries by 2015. This framework included ambitious coverage goals for prevention and treatment services resulting in a reduction of new HIV infections by more than half. However, it also estimated a leveling in the number of new infections at about 1 million annually after 2015. Methods: We modeled how the response to AIDS can be further expanded by scaling up antiretroviral treatment (ART) within the framework provided by the 2013 WHO treatment guidelines. We further explored the potential contributions of new prevention technologies: 'Test and Treat', pre-exposure prophylaxis and an HIV vaccine. Findings: Immediate aggressive scale up of existing approaches including the 2013 WHO guidelines could reduce new infections by 80%. A 'Test and Treat' approach could further reduce new infections. This could be further enhanced by a future highly effective pre-exposure prophylaxis and an HIV vaccine, so that a combination of all four approaches could reduce new infections to as low as 80,000 per year by 2050 and annual AIDS deaths to 260,000. Interpretation: In a set of ambitious scenarios, we find that immediate implementation of the 2013 WHO antiretroviral therapy guidelines could reduce new HIV infections by 80%. Further reductions may be achieved by moving to a 'Test and Treat' approach, and eventually by adding a highly effective pre-exposure prophylaxis and an HIV vaccine, if they become available.
Learning the structure of biomedical relationships from unstructured text (Part II)
<p>These files contain the seed sets and corresponding test sets of drug-gene pairs reflecting pharmacogenomic (PGx) and drug-target relationships that were used to evaluate EBC in the PLoS Comp Bio paper. Unfortunately, they didn't make it into the first upload for this paper. There are four zipped directories in this upload:</p> <p>data-drug-gene (PGx relationships, dense matrix)</p> <p>data-drug-target (drug-target relationships, dense matrix)</p> <p>data-drug-gene-fullmatrix (PGx relationships, sparse matrix)</p> <p>data-drug-target-fullmatrix (drug-target relationships, sparse matrix)</p>
Large-scale semantic indexing of Spanish biomedical literature using contrastive transfer learning
Open the record for dataset details and reuse information.
A Framework for Automated Construction of Heterogeneous Large-Scale Biomedical Knowledge Graphs (Recorded Talk)
<p>This entry contains the recording of the presentation that was presented at the 2020 Intelligent Systems for Molecular Biology as part of the Bio-Ontologies COSI (https://www.iscb.org/ismb2020).</p>
Nanoprobes for Biomedical Imaging with Tunable Near-Infrared Optical Properties Obtained via Green Synthesis
<p>Dataset of: 10.1002/adpr.202100260</p>
Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications.
<p>The datasets used in the publications titled "Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications"</p>
Dataser from paper "Proliferation of osteoblast precursor cells on the surface of TiO2 nanowires anodically grown on a -type biomedical titanium alloy"
<p>Dataser from paper "Proliferation of osteoblast precursor cells on the surface of TiO2 nanowires anodically grown on a -type biomedical titanium alloy":</p> <p>- <strong>Contact Angle: </strong>Images and measurements.</p> <p><strong>- Fluorescence Microscopy Images:</strong> Images.</p> <p><strong>- MTT and pixel counting:</strong> Measurements.</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
Data for "RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature"
<div> <p><strong>RegulaTome corpus</strong>: this <a href="../api/records/10808330/files/RegulaTome-corpus.tar.gz/content" target="_blank" rel="noopener">file</a> contains the RegulaTome corpus in <a href="https://brat.nlplab.org/">BRAT</a> format. The directory <strong>"splits" </strong>has the corpus split based on the train/dev/test used for the training of the relation extraction system</p> <p><strong>RegulaTome annodoc</strong>: The annotation guidelines along with the annotation configuration files for BRAT are provided in <a href="../api/records/10808330/files/annodoc+config.tar.gz/content" target="_blank" rel="noopener">annodoc+config.tar.gz</a>. The online version of the annotation documentation can be found here: <a href="https://katnastou.github.io/s1000-corpus-annotation-guidelines/">https://katnastou.github.io/regulatome-annodoc/ </a></p> <p>The tagger software can be found here: <a href="https://github.com/larsjuhljensen/tagger">https://github.com/larsjuhljensen/tagger</a>. The command used to run tagger before large-scale execution of the RE system is:</p> <p><code>gzip -cd `ls -1 pmc/*.en.merged.filtered.tsv.gz` `ls -1r pubmed/*.tsv.gz` | cat dictionary/excluded_documents.txt - | tagger/tagcorpus --threads=16 --autodetect --types=dictionary/curated_types.tsv --entities=dictionary/all_entities.tsv --names=dictionary/all_names_textmining.tsv --groups=dictionary/all_groups.tsv --stopwords=dictionary/all_global.tsv --local-stopwords=dictionary/all_local.tsv --type-pairs=dictionary/all_type_pairs.tsv --out-matches=all_matches.tsv</code></p> <p><strong>Input documents </strong>for large-scale execution, which is done on entire <a href="https://a3s.fi/March-2024-PubMed/PubMed_20230314.tar.gz" target="_blank" rel="noopener">PubMed</a> (as of March 2024) and <a href="https://a3s.fi/Jan-2024-documents/PMC_Nov_23.tar.gz" target="_blank" rel="noopener">PMC Open Access</a> (as of November 2023) articles in BioC format. The files are converted to a <a href="https://a3s.fi/March-2024-PubMed/all_documents.tsv" target="_blank" rel="noopener">tab-delimited format </a>to be compatible with the RE system input (see below).</p> <p><strong>Input dictionary files</strong>: all the files necessary to execute the command above are available in <a href="../api/records/10808330/files/tagger_dictionary_files.tar.gz/content" target="_blank" rel="noopener">tagger_dictionary_files.tar.gz </a></p> <p><strong>Tagger output</strong>: we filter the results of the tagger run down to gene/protein hits, and documents with more than 1 hit (since we are doing relation extraction) before feeding it to our RE system. The filtered output is available in <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a></p> <p><strong>Relation extraction system input</strong>: <a href="../api/records/10808330/files/combined_input_for_re.tar.gz/content" target="_blank" rel="noopener">combined_input_for_re.tar.gz</a>: these are the directories with all the .ann and .txt files used as input for the large-scale execution of the relation extraction pipeline. The files are generated from the tagger tsv output (see above, <a href="../api/records/10808330/files/tagger_matches_ggp_only_gt_1_hit.tsv.gz/content" target="_blank" rel="noopener">tagger_matches_ggp_only_gt_1_hit.tsv.gz</a>) using the <a href="https://github.com/spyysalo/string-db-tools/blob/main/tagger2standoff.py">tagger2standoff.py</a> script from the <a href="https://github.com/spyysalo/string-db-tools/">string-db-tools</a> repository.</p> <p><strong>Relation extraction models</strong>. The Transformer-based model used for large-scale relation extraction and prediction on the test set is at <a href="../api/records/10808330/files/relation_extraction_multi-label-best_model.tar.gz/content" target="_blank" rel="noopener">relation_extraction_multi-label-best_model.tar.gz</a></p> <p>The pre-trained RoBERTa model on PubMed and PMC and MIMIC-III with a BPE Vocab learned from PubMed (RoBERTa-large-PM-M3-Voc), which is used by our system is available <a href="https://github.com/facebookresearch/bio-lm/blob/main/README.md">here</a>.</p> <p><strong>Relation extraction system output</strong>: the tab-delimited outputs of the relation extraction system are found at <a href="https://a3s.fi/regulatome-ls/large_scale_relation_extraction_results.tar.gz" target="_blank" rel="noopener">large_scale_relation_extraction_results.tar.gz </a><strong>!!!ATTENTION this file is approximately 1TB in size, so make sure you have enough space to download it on your machine!!!</strong></p> <p>The relation extraction system output files have 86 columns: PMID, Entity BRAT ID1, Entity BRAT ID2, and scores per class produced by the relation extraction model. Each file has a header to denote which score is in which column.</p> </div>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.