Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

195

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

195 results for “biomedical”

Learn how ShareScore rates datasets ↗
zenodo44/100

A study on biomedical researchers' perspectives on public engagement in Southeast Asia

<p>Qualitative data on researcher&#39;s perceptions of public and community engagement in South and Southeast Asia</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

A high-throughput 3D X-ray histology facility for biomedical research and preclinical applications - Underlying Data

<p><strong>Video files and logs</strong></p> <p>Single-slice and thick-slice roll* source videos are included. Each video is accompanied by a .txt log that contains information about the source file, slice thickness, and a brief description of the visualization mode.</p> <p>List of files:</p> <ul> <li>20211019-23h59m_20xAvgInt.mp4</li> <li>20211019-23h59m_20xAvgInt.txt</li> <li>20211019-23h59m_20xMaxInt.mp4</li> <li>20211019-23h59m_20xMaxInt.txt</li> <li>20211019-23h59m_20xStDev.mp4</li> <li>20211019-23h59m_20xStDev.txt</li> <li>20211019-23h59m_XYSliceRoll.mp4</li> <li>20211019-23h59m_XYSliceRoll.txt</li> <li>20211019-23h59m_XZSliceRoll.mp4</li> <li>20211019-23h59m_XZSliceRoll.txt</li> <li>20211019-23h59m_YZSliceRoll.mp4</li> <li>20211019-23h59m_YZSliceRoll.txt</li> </ul> <p>*&nbsp;<em>Thick-slice rolling is a 2D thick-slice viewing that allows rolling of a pre-selected number of slices (n) along the z-axis of the 3D data. A single thick-slice roll forwards is accomplished by translating the thick-slice by one single slice forwards; that is moving forward by one (+1) slice from the first and nth element and reapplying the criteria or operations to the new slice sub-stack.</em></p> <p><strong>Volume XRH data</strong><br> These are processed raw volume file saved in .raw and/or .tiff format, which are resliced to a histology-relevant orientation and/or have been enhanced using noise reduction (3D median filter) and/or ct-artefact removal techniques (e.g. cBC identifies a bandpass filter used to remove intensity variations originating from the histology cassette).</p> <p>List of volume files:</p> <ul> <li><strong>32220_20200703_XRH_2504_OLK_DEMO02019-FFPE_1620x1959x164x16bit.raw</strong> <ul> <li>sample: Human lung adenocarcinoma</li> <li>histology-relevant resliced volume (2x2x2 3D medial filter applied)</li> <li>import as 1620 x 1959 x 164 x 16-bit, big-endian; voxel edge size (mm): 0.0160042 isotropic</li> </ul> </li> <li><strong>cBC_32220_20200703_XRH_2504_OLK_DEMO02019-FFPE_1588x1674x164x16bit.raw</strong> <ul> <li>sample: Human lung adenocarcinoma</li> <li>cassette artefacts background correction (bandpass) of volume 32220_20200703_XRH_2504_OLK_DEMO02019-FFPE_1620x1959x164x16bit.raw</li> <li>import as 1620 x 1959 x 164 x 16-bit, big-endian; voxel edge size (mm): 0.0160042 isotropic</li> </ul> </li> <li><strong>Med3D_HPass_2111_20190606_MEDX_2234_EH_HN2_recon_2000x1952x501x32bit.raw</strong> <ul> <li>sample: Human head and neck tumour</li> <li>histology-relevant resliced volume (1x1x1 3D medial filter applied)</li> <li>import as 2000 x 1952 x 501 x 32-bit, big-endian; voxel edge size (mm): 0.00999782 isotropic</li> </ul> </li> </ul> <p><strong>Conventional Histology and correlative imaging</strong></p> <ul> <li><strong>HN2_Level001_MEDX080_Manual_BW_Series4.tif</strong> <ul> <li>H&amp;E histology slice of the human head and neck tumour sample shown in &quot;Med3D_HPass_2111_20190606_MEDX_2234_EH_HN2_recon_2000x1952x501x32bit.raw&quot;</li> </ul> </li> <li><strong>HN2_Level001_MEDX080_Manual_BW</strong> <ul> <li>manual landmark selection used for registering the conventional histology slice onto the &mu;CT slice</li> </ul> </li> <li><strong>HN2_MEDX_rotated_0080.tif</strong> <ul> <li>Slice 80 from volume &quot;Med3D_HPass_2111_20190606_MEDX_2234_EH_HN2_recon_2000x1952x501x32bit.raw&quot; that corresponds to histological slice &quot;HN2_Level001_MEDX080_Manual_BW&quot;</li> </ul> </li> </ul>

opencc-by-4.0Jun 2023View details →
zenodo40/100

Database of Parenthetic Biomedical Abbreviations

<p>This dataset includes the biomedical abbreviations stated between parentheses in the titles of the scholarly publications indexed by PubMed between 1947 and 2019. Each&nbsp;abbreviation is extracted thanks to the parenthetic level count algorithm and is assigned to the title, PMID and year of publication of each corresponding research paper. Then, every acronym is allocated&nbsp;its length and the number of upper and lower case letters it involves. Finally, the entities including one or no upper case letter, less than three&nbsp;characters, eight characters or more, or a high rate of&nbsp;non-alphanumeric characters are semi-automatically eliminated&nbsp;to ensure the consistency of the research database.</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

Data for: Wikipedia as a gateway to biomedical research

<p>Wikipedia has been described as a gateway to knowledge. However, the extent to which this gateway ends at Wikipedia or continues via supporting citations is unknown. This dataset was used&nbsp;to establish benchmarks for the relative distribution and referral (click) rate of citations, as indicated by presence of a Digital Object Identifier (DOI), from Wikipedia with a focus on medical citations.</p> <p>This data set includes for each day in August 2016 a listing of all DOI present in the English language version of Wikipedia and whether or not the DOI are biomedical in nature. Source Code for these data are available at: Ryan Steinberg. (2017, July 9). Lane-Library/wiki-extract: initial Zenodo/DOI release. Zenodo. http://doi.org/10.5281/zenodo.824813</p> <p>This dataset also includes a listing from Crossref DOIs that were referred from Wikipedia in August 2016 (Wikipedia_referred_DOI). Source code for these data sets is available at:&nbsp;Joe Wass. (2017, July 4). CrossRef/logppj: Initial DOI registered release. Zenodo. http://doi.org/10.5281/zenodo.822636&nbsp;</p> <p>An&nbsp;article based on this data was published in PLOS One:</p> <p>Maggio LA, Willinsky JM, Steinberg RM, Mietchen D, Wass JL, Dong T. Wikipedia as a gateway to biomedical research: The relative distribution and use of citations in the English Wikipedia. PloS one. 2017 Dec 21;12(12):e0190046.&nbsp;</p> <p>https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0190046&nbsp;</p>

opencc-zeroJul 2017View details →
zenodo40/100

Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19

<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities.&nbsp;</span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764&nbsp;398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p>&nbsp;</p> <p>&nbsp;</p>

openJun 2022View details →
zenodo40/100

tBiomedL: Larger Semantic Table Annotations Benchmark for Biomedical Domain

<p><strong>tBiomedL </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p><p><strong>tBiomedL&nbsp;</strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using five levels of a recursive hierarchy of related concepts in Wikidata. It is the successor work of <a href="https://doi.org/10.5281/zenodo.10283103">tBiomed</a></p><p><strong>tBiomedL&nbsp;</strong>contains <strong>860,479</strong> entity and horizontal tables, while this repository contains only <strong>a sample of 1%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>27</strong> <strong>GB</strong>. We will update this repository with the full dataset, including the test fold with its ground truth data in the Future.</p><p>Please get in touch if you are interested in the full dataset,&nbsp;</p><p>The supported tasks for semantic table annotations are:&nbsp;</p><ol><li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li><li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li><li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li><li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li><li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li></ol>

opencc-by-4.0Dec 2023View details →
zenodo40/100

Ontology Enrichment from Texts (OET): A Biomedical Dataset for Concept Discovery and Placement

<p>A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic (CPP) product.</p> <p>The dataset is documented in the work,&nbsp;<em>Ontology Enrichment from Texts: A Biomedical Dataset for Concept Discovery and Placement</em>, on arXiv: <a href="https://arxiv.org/abs/2306.14704">https://arxiv.org/abs/2306.14704</a> (CIKM 2023). The companion code is available at https://github.com/KRR-Oxford/OET.</p> <p>Out-of-KB mention discovery (including the settings of mention-level data) is further partly documented in the work, <em>Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking</em>, on arXiv: <a href="https://arxiv.org/abs/2302.07189">https://arxiv.org/abs/2302.07189</a> (CIKM 2023).</p> <p>ver4: we made a version of mention-level data for out-of-KB discovery and concept placement separately: the former (for out-of-KB discovery) has out-of-KB mentions in training data, while the latter (for concept placement) has only out-of-KB mentions during the evaluation (validation and test) and not in the training data. Also, we split the original "test-NIL.jsonl" (now "test-NIL-all.jsonl") into "valid-NIL.jsonl" and "test-NIL.jsonl" for a better evaluation.</p> <p>ver3: we revised and updated mention-level data (syn_full, synonym augmentation setting) and the folder structure, and also updated the edge catalogues with complex edges.</p> <p>ver2: we revised the mention-level data by only keeping out-of-KB mentions (or "NIL" mentions) associated with one-hop edges (including leaf nodes, as &lt;leaf node, NULL&gt;) and two-hop edges in the ontology (SNOMED CT 20140901).</p> <p>Acknowledgement of data sources and tools below:</p> <p>* SNOMED CT https://www.nlm.nih.gov/healthit/snomedct/archive.html (and use snomed-owl-toolkit to form .owl files)<br>* UMLS https://www.nlm.nih.gov/research/umls/licensedcontent/umlsarchives04.html (and mainly use MRCONSO for mapping UMLS to SNOMED CT)<br>* MedMentions https://github.com/chanzuckerberg/MedMentions (source of entity linking)</p> <p>* Prot&eacute;g&eacute; http://protegeproject.github.io/protege/<br>* snomed-owl-toolkit https://github.com/IHTSDO/snomed-owl-toolkit<br>* DeepOnto https://github.com/KRR-Oxford/DeepOnto (based on OWLAPI https://owlapi.sourceforge.net/) for ontology processing and complex concept verbalisation</p>

opencc-by-4.0Jun 2023View details →
zenodo40/100

3D printed biomedical devices and their applications: A review on state-of-the-art technologies, existing challenges, and future perspectives

<p>This repository consists of the data that have been utilized to imagine and subsequently construct Fig. 1, Fig. 3 , Fig. 5 , Fig. 6 and Fig. 7 of the article " H. B. Mamo, M. Adamiak and A. Kunwar. 3D printed biomedical devices and their applications: A review on state-of-the-art technologies, existing challenges, and future perspectives, Journal of the Mechanical Behavior of Biomedical Materials, 143 (2023) 105930. doi: 10.1016/j.jmbbm.2023.105930 ".</p> <p>&nbsp;Brief introduction of the files contained in this repository</p> <ol> <li>&nbsp;acronyms.csv: This file consists of&nbsp; the list of acronyms associated with materials and techniques used in 3D printing of biomedical devices.</li> <li>fig1a-data.csv: The csv file provides a relative ranking of different types of 3D printing techniques for biomedical applications based upon &nbsp;merits and limitations (wherever applicable). The technique listed as Rank 1 is considered as the most commonly used 3D printing procedure.</li> <li>&nbsp;fig1b-data.csv: The csv file enlists the benfits of the 3D printing techniques in biomedical applications as compared to subtractive manufacturing methods.</li> <li>&nbsp;fig3-data.csv: The file contains a comparison between conventional and customized tablet printing &nbsp;methods. The illustration is made through the manufacturing or printing of pharmaceutical tablets or pills.&nbsp; This illustation also applies to the production of pharmaceutical capsules.&nbsp; It thus illustrates that customized medicine is enabled using 3D printing technology.</li> <li>&nbsp;fig5-data.csv: This file enumerates the major challenges associated with 3D printing of biomedical devices.</li> <li>&nbsp;fig6-data.csv: The file enlists the roles of wearable smarts, cloud-based platforms and physicians in context of the hospitals implementing IoMT.</li> <li>&nbsp;fig7-data.csv: The aspects of IoMT Sensors,IoMT Platforms,3D Printers,Design and Prototypes within the integrated 3D printing-IoMT ecosystem are listed in the file.</li> </ol>

opencc-zeroMay 2023View details →
zenodo40/100

Extracting Biomedical Entities from Noisy Audio Transcripts--Dataset

<p><strong>SUMMARY</strong>:</p> <p>This repo contains the CADEC and Synthetic BTACT datasets that were used for the paper titled "<em>Extracting Biomedical Entities from Noisy Audio Transcripts</em>."</p> <p>The dataset includes two sets: i) CADEC (Karimi et al., 2015) and ii) Synthetic BTACT. CADEC is a well-known NER dataset used to identify adverse drug reactions based on what patients have written about their experiences. Synthetic BTACT is the data that we have made up. It is created based on questions similar to those in the Brief Test of Adult Cognition by Telephone (BTACT)(Tun et al., 2006).</p> <p>CADEC includes two sets of audio files; one is read from the original CADEC, and the other one is with additional audio noise. It also includes the original CADEC scripts, annotations, and the transcripts of the noisy audio. The transcripts are generated using Whisper. The annotations encompass named entities, their types, and string indexes of their occurrence in the text. Annotations also include "AnnotatorNotes" which explains some of the annotations.</p> <p>The synthetic BTACT data include two types: i) animals and ii) fruits. Similar to CADEC, it includes two sets of audio files: one that is read from the original scripts and another one with additional audio noise.&nbsp; The text files include the original scripts, annotations, and the Whisper-transcribed of the noisy audio files. The annotations include indexes of named entities, their string indices and types.</p> <p><strong>REFERENCES</strong>:</p> <p>Karimi, S., Metke-Jimenez, A., Kemp, M., &amp; Wang, C. (2015). Cadec: A corpus of adverse drug event annotations. Journal of biomedical informatics, 55, 73-81.</p> <p>Tun, P. A., &amp; Lachman, M. E. (2006). Telephone assessment of cognitive function in adulthood: the Brief Test of Adult Cognition by Telephone. Age and Ageing, 35(6), 629-632.</p> <p><strong>DETAILS</strong>:</p> <p>Data_1: CADEC (1250 TextFiles, 1000 Audio, types=5):</p> <p>&nbsp; &nbsp; General Categories and Counts<br>&nbsp; &nbsp; ADR (Adverse Drug Reactions): 5316<br>&nbsp; &nbsp; DRUG: 1797<br>&nbsp; &nbsp; FINDING: 397<br>&nbsp; &nbsp; DISEASE: 280<br>&nbsp; &nbsp; SYMPTOM: 255<br>&nbsp; &nbsp; Specific Items (Drugs) and Counts<br>&nbsp; &nbsp; Arthrotec: 145<br>&nbsp; &nbsp; cambia: 4<br>&nbsp; &nbsp; cataflam: 10<br>&nbsp; &nbsp; diclofenac-potassium: 3<br>&nbsp; &nbsp; diclofenac-sodium: 7<br>&nbsp; &nbsp; flector: 1<br>&nbsp; &nbsp; Lipitor: 997<br>&nbsp; &nbsp; Pennsaid: 4<br>&nbsp; &nbsp; solarez: 3<br>&nbsp; &nbsp; voltaren: 46<br>&nbsp; &nbsp; voltaren-rx: 22<br>&nbsp; &nbsp; zipsor: 5</p> <p>Data_2: Synthetic BTACT (500 Fruits, 500 Animals, types=2)</p> <p>&gt;&gt; Audios can be matched with annotations, scripts and transcripts using their filenames.&nbsp;</p> <p>---<br>audio [original&amp;noisy]:<br>&nbsp; &nbsp; 1. cadec<br>&nbsp; &nbsp; &nbsp; &nbsp; 1.1 cadec original<br>&nbsp; &nbsp; &nbsp; &nbsp; 1.2 cadec noisy<br>&nbsp; &nbsp; 2. synthetic btact<br>&nbsp; &nbsp; &nbsp; &nbsp; 2.1 btact original<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.1.1 fruits<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; fruit-script[0:500].mp3<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.1.2 animals<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].mp3<br>&nbsp; &nbsp; &nbsp; &nbsp; 2.2 btact noisy<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.2.1 fruits<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; fruit-script[0:500].mp3<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.2.2 animals<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].mp3<br>text[scripts, annotations, transcripts]:<br>&nbsp; &nbsp; 1. cadec<br>&nbsp; &nbsp; &nbsp; &nbsp; 1.1 scripts [1,250]<br>&nbsp; &nbsp; &nbsp; &nbsp; 1.2 annotations [1,250] (index/AnnotatorsNote, type, indices, named-entities)<br>&nbsp; &nbsp; &nbsp; &nbsp; 1.3 transcripts [1,000]<br>&nbsp; &nbsp; 2. synthetic btact<br>&nbsp; &nbsp; &nbsp; &nbsp; 2.1 animals<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.1 scripts (original scripts)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].txt<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.2 annotations<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].ann (index, type, start/end indices, named entity)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.3 transcripts<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].txt<br>&nbsp; &nbsp; &nbsp; &nbsp; 2.2. fruits<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.1 scripts (original scripts)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].txt<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.2 annotations<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; script-[0:500].ann (index, type, start/end indices, named entity)<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 2.3 transcripts<br>&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; fruit-script-[0:500].txt</p> <p>&nbsp;</p> <p><strong>CITATION</strong>:</p> <p>Ebadi, N., Morgan, K., Tan, A., Linares, B., Osborn, S., Majors, E., Davis, J., &amp; Rios, A. (2024). Extracting biomedical entities from noisy audio transcripts. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024).</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

tBiomed: Semantic Table Annotations Benchmark for Biomedical Domain

<p><strong>tBiomed </strong>is a dataset for tabular data to knowledge graph matching. It is derived for the Biodiversity domain and has&nbsp;two types of tables.&nbsp;On the one hand, <strong>Horizontal Relational Tables</strong>&nbsp;are where&nbsp;each table&nbsp;represents a collection of entities. On the other&nbsp;hand, <strong>Entity Tables</strong> represent a single entity. We supported ground truth data from Wikidata as a target knowledge graph (KG).</p> <p><strong>tBiomed&nbsp;</strong>is generated by <a href="https://github.com/fusion-jena/KG2Tables">KG2Tables </a>using two levels of a recursive hierarchy of related concepts in Wikidata.</p> <p><strong>tBiomed&nbsp;</strong>contains <strong>26,778</strong> entity and horizontal tables, while this repository contains only a <strong>validation fold</strong> of the original data representing <strong>20%</strong> of the total of the entire benchmark with its ground truth data (gt). The Full size of this dataset is <strong>1</strong> <strong>GB</strong>.</p> <p>We included the full version of the dataset. We will update this repository ground truth data of the test set in the Future.</p> <p>The supported tasks for semantic table annotations are:&nbsp;</p> <ol> <li>Topic Detection (<strong>TD</strong>) links the entire table to an entity or a class from the target KG.</li> <li>Cell Entity Annotation (<strong>CEA</strong>) maps individual table cells to entities from the target KG.</li> <li>Column Type Annotation (<strong>CTA</strong>) links individual table columns to classes from the target KG.</li> <li>Column Property Annotation (<strong>CPA</strong>) detects the relations between column pairs from the target knowledge graph.</li> <li>Row Annotation (<strong>RA) </strong>annotates the entire row to a KG entity or property.</li> </ol>

opencc-by-4.0Dec 2023View details →
zenodo40/100

PrimeKGQA, the dataset from paper: Bridging the Gap: Generating a Comprehensive Biomedical Knowledge Graph Question Answering Dataset

<p>Despite the plethora of resources such as large-scale&nbsp;corpora and manually curated Knowledge Graphs (KGs), the ability to perform reasoning with natural language inputs over biomedical graphs remains challenging due to insufficient training data. We&nbsp;propose a novel method for automatically constructing a Biomedical&nbsp;Knowledge Graph Question Answering (BioKGQA) dataset sourced&nbsp;from PrimeKG, the largest precision medicine-oriented KG. In total,<br>we create 83999 question-answer pairs along with their respective&nbsp;SPARQL queries. Our approach generates a diverse array of contextually relevant questions covering a wide spectrum of biomedical&nbsp;concepts and levels of complexity. We evaluate our method based on&nbsp;automatic metrics alongside manual annotations. We establish novel&nbsp;standards tailored for KGQA systems to highlight the linguistic correctness and semantical faithfulness of the generated questions based&nbsp;on extracted KG facts. The compiled dataset &ndash; PrimeKGQA &ndash; serves&nbsp;as a valuable benchmarking resource for advancing knowledge-driven biomedical research and evaluating KGQA system.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences in Biomedical Literature and Its Integration with PubChemRDF

<p>This Zenodo record contains the co-occurrence RDF data generated in the work described in the paper &ldquo;<strong>A resource description framework (RDF) model of named entity co-occurrences in biomedical literature and its integration with PubChemRDF</strong>&rdquo; by Li et al., published in the Journal of Cheminformatics (<a href="https://doi.org/10.1186/s13321-025-01017-0" target="_blank" rel="noopener">https://doi.org/10.1186/s13321-025-01017-0</a>).&nbsp; It also contains the SPARQL query examples, the RDF schema in SHACL and ShEx, and the validation scripts.</p> <p>All content in this Zenodo record is for archival purposes.&nbsp;The latest version of the co-occurrence RDF data and other PubChemRDF data can be accessed via the PubChem FTP site (<a href="https://ftp.ncbi.nlm.nih.gov/pubchem/RDF/" target="_blank" rel="noopener">https://ftp.ncbi.nlm.nih.gov/pubchem/RDF/</a>). The up-to-date RDF schema in various formats is available on the PubChemRDF Schema page (<a href="https://pubchem.ncbi.nlm.nih.gov/docs/rdf-schema" target="_blank" rel="noopener">https://pubchem.ncbi.nlm.nih.gov/docs/rdf-schema</a>). A set of SPARQL query examples can be found on the PubChemRDF use case pages (<a href="https://pubchem.ncbi.nlm.nih.gov/docs/rdf-use-cases" target="_blank" rel="noopener">https://pubchem.ncbi.nlm.nih.gov/docs/rdf-use-cases</a>).</p>

opencc-by-4.0Nov 2023View details →
zenodo40/100

Determinant Factors in Systematic Review of Biomedical Engineering Topics

<p>Students and researchers around the world are increasingly turning to literature reviews. Review articles give you a broad picture of the field and help synthesize published research that is expanding at a rapid pace.</p> <p>Professionally crafted literature reviews, whether written by a student in class or by an experienced researcher for publication, should aim to add to the literature rather than detract from it. This is not an easy feat, but it is a necessary one.<br> Systematic reviews aim to identify, evaluate, and summarize the findings of all relevant individual studies on a related topic; which makes the available evidence more accessible to decision makers.</p> <p>Research Question<br> &bull; Provide enough detail so that the audience can easily understand your purpose without the need for additional explanation.</p> <p>&bull; It is not answered with a simple &ldquo;yes&rdquo; or &ldquo;no&rdquo;, but requires synthesis and analysis of ideas and sources before writing a response.</p> <p>Search strategy<br> &bull; Organized structure of key terms used to search databases.</p> <p>&bull; Combine key concepts to get accurate results.</p> <p>Inclusion and exclusion criteria<br> They are determined after the research question is established, usually before the search is performed, however, scoping searches may be necessary to determine the appropriate criteria</p> <p>Synthesis<br> &bull; A review matrix and summary tables are convenient ways to quickly summarize and reorganize qualitative information and observations to achieve a preliminary outline.</p> <p>Systematic reviews serve several critical functions</p> <p>Provide synthesis of the state of knowledge in a field, from which future research priorities can be identified;</p> <p>Addressing questions that could not otherwise be answered by individual studies;</p> <p>Identify problems in primary research that should be rectified in future studies;Generate or evaluate theories about how or why phenomena occur.</p>

opencc-by-4.0Nov 2021View details →
zenodo40/100

BioVAE: a pre-trained latent variable language model for biomedical text mining

<p>We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.</p> <p>This version contains&nbsp;the pre-trained models for text mining tasks such as named entity recognition or&nbsp;relation extraction, and text generation task.</p> <p>Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)</p> <ul> <li>pm-full-lt32-beta00</li> <li>pm-full-lt32-beta05</li> <li>pm-full-lt768-beta00</li> <li>pm-full-lt768-beta05</li> <li>pm-full-generation</li> </ul>

openapache2.0Nov 2021View details →
dryad40/100

The new normal? Redaction bias in biomedical science

<p>A concerning amount of biomedical research is not reproducible. Unreliable results impede empirical progress in medical science, ultimately putting patients at risk. Many proximal causes of this irreproducibility have been identified, a major one being inappropriate statistical methods and analytical choices by investigators. Within this, we formally quantify the impact inappropriate redaction beyond a threshold value in biomedical science. This is effectively truncation of a data-set by removing extreme data points, and we elucidate its potential to accidentally or deliberately engineer a spurious result in significance testing. We demonstrate that the removal of a surprisingly small number of data points can be used to dramatically alter a result. It is unknown how often redaction bias occurs in the broader literature, but given the risk of distortion to the literature involved, we suggest that it must be studiously avoided, and mitigated with approaches to counteract any potential malign effects to the research quality of medical science.</p>

opencc-zeroDec 2021View details →
zenodo40/100

Biomedical journals covered and uncovered in PubMed

<p>PubMed contains many biomedical journals, but is this the full range of biomedical journals?</p> <p>We compared PubMed with MAG and found that there are many journals that are not included in PubMed. In addition, we found that even for the biomedical journals covered by PubMed, there are many articles under the&nbsp;covered journals not found in PubMed. The two tables below present these pieces of information.</p> <p>intersecting-journal-with-indexing-completeness.tsv: referring to the journals whose articles are not fully included in PubMed</p> <p>uncovered-journal-with-number-biomedical-articles.tsv:&nbsp;referring to the journals that have not been touched by&nbsp;PubMed</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Spanish SciELO Crawled Biomedical Corpus

<p>We present a corpus of Spanish medical articles extracted from the SciELO website (https://scielo.cl/). The corpus was constructed using web scraping extraction techniques and consists of 5694 articles published between 2002 and 2020 across 34 journals specialized in health and biology (specified below).</p> <p>Two&nbsp;variables of the corpus are presented here: the first contains the text without pre-processing, which makes it possible to adapt the corpus for different purposes. The second and third have&nbsp;been subjected to pre-processing, which is carried out with the NLTK library in Python and removes capital letters, accents, punctuation, and any non-alphanumeric symbols. In addition, the corpus was tokenized by sentences, obtaining a total of 500828 sentences and 13M tokens.&nbsp; In both cases, the articles are grouped by journal, in order of publication from most recent to oldest. These journals and their respective total of issues correspond to:</p> <ol> <li> <p>Acta bioethica&nbsp; (43 issues)</p> </li> <li> <p>Anales del Instituto de la Patagonia (31&nbsp;issues)</p> </li> <li> <p>Andes pediatrica (3&nbsp;issues)</p> </li> <li> <p>Biological Research (63&nbsp;issues)</p> </li> <li> <p>Ciencia y enfermer&iacute;a - Revista iberoamericana de investigaci&oacute;n&nbsp; (45&nbsp;issues)</p> </li> <li> <p>Gayana (Concepci&oacute;n) - International Journal of Biodiversity, Oceanology and Conservation (45&nbsp;issues)</p> </li> <li> <p>Gayana. Bot&aacute;nica (41&nbsp;issues)</p> </li> <li> <p>International Journal of Morphology (79&nbsp;issues)</p> </li> <li> <p>International journal of interdisciplinary dentistry (5 issues)</p> </li> <li> <p>International journal of odontostomatology (39 issues)</p> </li> <li> <p>Latin american journal of aquatic research (58&nbsp;issues)</p> </li> <li> <p>Revista chilena de cardiolog&iacute;a (37 issues)</p> </li> <li> <p>Revista chilena de enfermedades respiratorias (76 issues)</p> </li> <li> <p>Revista chilena de entomolog&iacute;a (5&nbsp;issues)</p> </li> <li> <p>Revista chilena de historia natural (64&nbsp;issues)</p> </li> <li> <p>Revista chilena de infectolog&iacute;a (131 issues)</p> </li> <li> <p>Revista chilena de neuro-psiquiatr&iacute;a (89&nbsp;issues)</p> </li> <li> <p>Revista chilena de nutrici&oacute;n (86&nbsp;issues)</p> </li> <li> <p>Revista chilena de obstetricia y ginecolog&iacute;a (117 issues)</p> </li> <li> <p>Revista chilena de radiolog&iacute;a (79&nbsp;issues)</p> </li> <li> <p>Revista de biolog&iacute;a marina y oceanograf&iacute;a (58&nbsp;issues)</p> </li> <li> <p>Revista de cirug&iacute;a (16&nbsp;issues)</p> </li> <li> <p>Revista de otorrinolaringolog&iacute;a y cirug&iacute;a de cabeza y cuello (52&nbsp;issues)</p> </li> <li> <p>Revista m&eacute;dica de Chile (264 issues)</p> </li> </ol> <p>&nbsp;</p> <p><strong>Non-current titles</strong></p> <ol> <li> <p>&nbsp;Bolet&iacute;n chileno de parasitolog&iacute;a (5&nbsp;issues) - Jan&nbsp; 2002: Completed; Continued as Parasitolog&iacute;a latinoamericana</p> </li> <li> <p>Ciencia &amp; trabajo (18&nbsp;issues) - Feb&nbsp; 2020: Indexing interrupted</p> </li> <li> <p>&nbsp;Electronic Journal of Biotechnology (84&nbsp;issues) - April&nbsp; 2017: Indexing interrupted</p> </li> <li> <p>&nbsp;Investigaciones marinas (22&nbsp;issues) - Nov&nbsp; 2007: Completed; Continued as Latin american journal of aquatic research</p> </li> <li> <p>&nbsp;&nbsp;Parasitolog&iacute;a al d&iacute;a (9&nbsp;issues) - Jul&nbsp; 2001: Completed; Continued as Parasitolog&iacute;a latinoamericana</p> </li> <li> <p>Parasitolog&iacute;a latinoamericana (13 issues) - Dec&nbsp; 2008: Completed</p> </li> <li> <p>Revista chilena de anatom&iacute;a (14&nbsp;issues) - 2002: Completed ; Continued as International Journal of Morphology</p> </li> <li> <p>Revista chilena de cirug&iacute;a (78 issues) - May&nbsp; 2019: Completed ; Continued as Revista de cirug&iacute;a</p> </li> <li> <p>Revista chilena de pediatr&iacute;a (476&nbsp;n&uacute;meros) - March&nbsp; 2021: Completed ; Continued as Andes pediatrica</p> </li> <li> <p>Revista cl&iacute;nica de periodoncia, implantolog&iacute;a y rehabilitaci&oacute;n oral (30&nbsp;n&uacute;meros) - April&nbsp; 2020: Completed ; Continued as International journal of interdisciplinary dentistry</p> </li> </ol>

opencc-by-4.0Jan 2022View details →
dryad40/100

CZ Software Mentions: A large dataset of software mentions in the biomedical literature

<p>We describe the CZ Software Mentions dataset, a new dataset of software mentions in biomedical papers. Plain-text software mentions are extracted with a trained SciBERT model from several sources: the NIH PubMed Central collection and from papers provided by various publishers to the Chan Zuckerberg Initiative. The dataset provides sources, context and metadata, and, for a number of mentions, the disambiguated software entities and links. We extract 1.12 million unique string software mentions from 2.4 million papers in the NIH PMC-OA Commercial subset, 481k unique mentions from the NIH PMC-OA Non-Commercial subset (both gathered in October 2021) and 934k unique mentions from 4 million papers in the Publishers' collection. There is variation in how software is mentioned in papers and extracted by the NER algorithm. We propose a clustering-based disambiguation algorithm to map plain-text software mentions into distinct software entities and apply it on the NIH PubMed Central Commercial collection. Through this methodology, we disambiguate 1.12 million unique strings extracted by the NER model into ~97000 unique software entities, covering 78% of all links. We link 185 000 of the mentions to a repository, covering about 55% of all software-paper links. We make all data and code publicly available as a new resource to help assess the impact of software (in particular scientific open source projects) on science.</p>

opencc-zeroSep 2022View details →
zenodo40/100

FAIRness Assessment of Biomedical Data Using Automated Tools (Dataset)

<p>The data were collected as part of a Master's thesis project aimed at evaluating various automated FAIR assessment tools, applying them to biomedical data. The data sets identifiers were gathered as part of the Open Data LoM and IoM incentivization at Charit&eacute; Universit&auml;tsmedizin Berlin, available at&nbsp;<a title="Dataset of the results of data validation for articles from 2021" href="https://doi.org/10.5281/zenodo.8249758">https://doi.org/10.5281/zenodo.8249758</a>, and reused in this project.</p> <p>The data represents cleaned, aggregated, and transformed results obtained from the API services of the following FAIR assessment tools: F-UJI, FAIR Enough, FAIR-Checker, and FAIR EVA.</p> <p>The raw data in .Rdata format will be shared on GitHub repository at <a title="FAIR Tools Analysis" href="https://github.com/anastasiabright/fair-tools-analysis">https://github.com/anastasiabright/fair-tools-analysis</a>.</p>

opencc-by-4.0Jun 2024View details →
zenodo40/100

Dataset: Clearside Biomedical, Inc. (CLSD) Stock Performance

This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.

opencc-zeroJun 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record