Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5,565
datasets available to search
ShareScore release 0.9.0
Dataset results
5,565 results for “medical”
Dataset for the IntoValue 1 + 2 studies on results dissemination from clinical trials conducted at German university medical centers completed between 2009 and 2017
<p>The IntoValue dataset contains clinical trials conducted at one of 35 German UMCs and registered on ClinicalTrials.gov or the German Clinical Trials Registry (DRKS). All trials were reported as complete between 2009 and 2017 on the trial registry at the time of data collection. The dataset also includes a results publication found via manual searches; if multiple results publications were found, the earliest was included.</p> <p>Trials were associated with a German UMC by searching for trials with a UMC listed as responsible party or lead sponsor, or with a principle investigator (PI) from a UMC ('lead_city'). Version 1 additionally includes trials with a UMC only as a facility (`facility_city`). A lookup table of regular expressions used to identify German UMCs is available at <a href="https://github.com/quest-bih/IntoValue2/blob/master/data/1_sample_generation/city_search_terms.csv">https://github.com/quest-bih/IntoValue2/blob/master/data/1_sample_generation/city_search_terms.csv</a>.</p> <p>Trials include all interventional studies and are not limited to investigational medical product trials, as regulated by the EU's Clinical Trials Directive or Germany's Arzneimittelgesetz (AMG) or Novelle des Medizinproduktegesetzes (MPG).</p> <p>DRKS data were searched (pre-filtered for completion years and study status as well as Germany as 'Country of recruitment') and downloaded as CSVs from the DRKS website (<a href="https://www.drks.de/">https://www.drks.de/</a>). ClinicalTrials.gov data were downloaded downloaded as pipe files from Clinical Trials Transformation Initiative (CTTI) Aggregate Content of ClinicalTrials.gov (AACT) (<a href="https://aact.ctti-clinicaltrials.org/pipe_files">https://aact.ctti-clinicaltrials.org/pipe_files</a>). DRKS and ClinicalTrials.gov use different terminology for various trial aspects, such as phase and masking; these different levels are captured in the data dictionary as `levels_drks` and `levels_ctgov`. For later analyses requiring parity across registries, levels for some variables were collapsed and a lookup table is provided in `iv_data_lookup_registries.csv`.</p> <p>These data were generated and used for two publications (Wieschowski et al., 2019; Riedel et al. 2021) and therefore comprises two versions (indicated as `iv_version`).</p> <p>For version 1, registry data was collected on April 17, 2017 from ClinicalTrials.gov and on July 27, 2017 for DRKS and was limited to trials with a completion date on DRKS and primary completion date on ClinicalTrials.gov between 2009 and 2013. Version 1 manual searches for results publications were conducted from 2017-07-01 to 2017-12-01.<br> For version 2, registry data was collected on June 3, 2020 and was limited to trials with a completion date on DRKS and ClinicalTrials.gov between 2014 and 2017. Version 2 manual searches for results publications were conducted from 2020-07-01 to 2020-09-01.</p> <p>Raw registry data for versions 1 and 2 is available in `raw-registries.zip`.</p> <p>Publication identifiers (DOI, PMID, URL) were manually entered during the publication search and then further enhanced using the API of Internet Archive's open-source Fatcat catalog of research publications, to add PMIDs based on DOIs, and vice versa.</p> <p>Manual search steps differed slightly in the two versions and are indicated and described in `identification_step`.<br> Version 1 includes trials with a German UMC as either a `lead_city` or a `facility_city`, whereas version 2 is limited to trials a German UMC as a `lead_city`.</p> <p>Each row indicates a single trial registration. Due to changes in completion dates, some trials are duplicated between versions as indicated in `is_dupe`. Cross-registered trials were manually deduplicated, and some cross-registered duplicates remain (e.g., DRKS00004156 and NCT00215683) and are not indicated in the dataset.</p> <p>All dates are provided as `yyyy-mm-dd`.</p> <p>Additional documentation on each variable (type, description, levels) is provided in `iv_data_dictionary.csv`.</p> <p>Additional information on the project and methods for generating the dataset is available in associated publications and at the project's OSF page (<a href="https://osf.io/98j7u/">https://osf.io/98j7u/</a>). Code for the project is available at <a href="https://github.com/quest-bih/IntoValue2">https://github.com/quest-bih/IntoValue2</a>.</p> <p><strong>References:</strong></p> <p>Wieschowski, S., Riedel, N., Wollmann, K., Kahrass, H., Müller-Ohlraun, S., Schürmann, C., Kelley, S., Kszuk, U., Siegerink, B., Dirnagl, U., Meerpohl, J., & Strech, D. (2019). Result dissemination from clinical trials conducted at German university medical centers was delayed and incomplete. Journal of Clinical Epidemiology, 115, 37–45. <a href="https://doi.org/10.1016/j.jclinepi.2019.06.002">https://doi.org/10.1016/j.jclinepi.2019.06.002</a></p> <p>Riedel, N., Wieschowski, S., Bruckner, T., Holst, M. R., Kahrass, H., Nury, E., Meerpohl, J. J., Salholz-Hillel, M., & Strech, D. (2021). Results dissemination from completed clinical trials conducted at German university medical centers remained delayed and incomplete. The 2014-2017 cohort. Journal of Clinical Epidemiology, 0(0). <a href="http://doi.org/10.1016/j.jclinepi.2021.12.012">https://doi.org/10.1016/j.jclinepi.2021.12.012</a><br> </p>
Titanium Alloys Database for Medical Applications
<p>The new 2.0 version (12.7.2023) includes the following modifications: 247 biocompatible Ti alloys; the table shows only literature data; a Jupyter notebook provides the calculated data.</p> <p>In this database, 238 titanium alloys were collected, almost entirely of biocompatible alloying elements. The primary motivation behind creating such a database is to establish a foundation for designing new alloys using machine learning methods. The database can assist researchers, engineers, and biomedical professionals in developing titanium alloys for various medical applications, thereby improving health outcomes and driving advancements in biomaterials and biomedical engineering.</p> <p>For more information read the paper at: <a href="https://doi.org/10.30544/MMD5"> https://doi.org/10.30544/MMD5 </a></p> <p>NOTE: To avoid misunderstandings, please cite both the database and the published article when citing this database.</p> <p>We invite other authors to contribute to the updating of this database (send at least 20 new alloys to appear as co-author)</p>
Data set supplementing "Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study"
<p>This is the de-identified data set used to conduct the analyses in the study published as Original Research in the JMIR under the title "Benchmarking triage capability of symptom checkers against that of medical laypersons: Survey study" (https://doi.org/10.2196/24475)</p> <p>The data set contains the assessments of the urgency of symptoms to 45 fictitious clinical case vignettes by 91 US participants, and the participants' age, gender and level of education. Data for the symptom checker apps is needed to fully reproduce our study and can be found in the appendix of the paper "Evaluation of symptom checkers for self diagnosis and triage: audit study" by Semigran et al. (2015) (https://doi.org/10.1136/bmj.h3480).</p>
A living catalogue of artificial intelligence datasets and benchmarks for medical decision making
<p>We provide a comprehensive curated catalogue of <strong>artificial intelligence datasets</strong> and <strong>benchmarks for medical decision making</strong>. At the time of first release (April 2021), the dataset contains more than 400 biomedical and clinical datasets of which 252 are publicly available or available upon request.</p> <p>The dataset was compiled based on a systematic literature review covering both biomedical and computer science literature and grey literature data sources. All datasets were manually systematized and annotated for meta-information, such as:</p> <ul> <li>Availability and licensing information</li> <li>Type of source data</li> <li>Links to source publications, main references or dataset repositories</li> </ul> <p>Benchmark dataset were additionally annotated for the following information:</p> <ul> <li>Associated task</li> <li>Performance metrics commonly used for evaluation</li> <li>Clinical relevance</li> <li>The availability of data splits</li> </ul> <p>In addition to the versioned TSV file on Zenodo, the dataset can also be explored live via <a href="https://docs.google.com/spreadsheets/d/1QjUxxnZ3tuyW5dj6nkt_o5yJcWUZec4ttfJxO8Zlty4/edit?usp=sharing">this Google Spreadsheet</a>. The dataset is intended as a living, extendable resource. Edit suggestions and additions are encouraged and can be submitted via the comment function of the Google sheet.</p> <p> </p> <p><strong>File descriptions</strong></p> <p><em>annotated-datasets.tsv</em> -- contains the annotated datasets</p> <p><em>arXiv-literature-export.tsv</em> -- contains the original literature record export from arXiv</p> <p><em>pubmed-literature-export.tsv</em> -- contains the original literature record export from PubMed</p> <p><em>README.md</em> -- contains a detailed description of all annotation fields</p>
Impact of medical radionuclide discharges on people and the environment: scenario data used in the non-human biota impact assessment
<p>This dataset contains the input data for the D-DAT model: activity concentrations in water for the simulated Molse Nete scenario. It also contains the dynamic model-calculated activity concentrations in sediment and the non-human biota. These are the primary data upon which the dose calculations werte performed, and they can be used to reproduce these calculations. The related preprint article is also given in this repository: https://zenodo.org/records/10488393.</p> <p> </p> <p> </p> <p> </p>
MESINESP: Medical Semantic Indexing in Spanish - Development dataset
<p><em><strong>Please use the <a href="https://doi.org/10.5281/zenodo.4612274">MESINESP2 corpus (the second edition of the shared-task)</a> since it has a higher level of curation, quality and is organized by document type (scientific articles, patents and clinical trials).</strong></em></p> <p> </p> <p> </p> <p><strong>Introduction</strong></p> <p>The Mesinesp (Spanish BioASQ track, see https://temu.bsc.es/mesinesp) development set has a total of 750 records indexed manually by seven experienced medical literature indexers. Indexing is done using <em>DeCS codes, a sort of Spanish equivalent to MeSH terms</em>. Records were distributed in a way that each article was annotated, at least, by two different human indexers.</p> <p>The data annotation process consisted in two steps:</p> <ol> <li>Manual indexing step. DeCS codes were manually assigned to each record following the DeCS manual indexing guidelines.</li> <li>Manual validation and consensus. The joined set of manually indexed DeCS codes generated by both indexers were manually revised and corrections were done.</li> </ol> <p>These annotations were analyzed, resulting in an agreement using the Jaccard index.</p> <p>Records consisted basically in medical literature abstracts and titles from the IBECS and LILACS databases.</p> <p><strong>Zip structure</strong><br> The zip file contains two different development sets:</p> <ul> <li><em>Official development set</em>, which has the union of the annotations, with an agreement of macro = 0.6568 and micro = 0.6819. This set is composed by all the different (unique) DeCS codes that have been added by any annotator for each document; and</li> <li><em>Core-descriptors development set</em>, which has the intersection of the annotations, with an agreement of macro = 1.0 and micro = 1.0. This set is composed of the common DeCS codes that have been added by two or more annotators for each document.</li> </ul> <p><strong>Corpus format</strong></p> <p>Each dataset is a JSON object with one single key named "articles", which contains a list of documents. So, the raw format of the file is one line per document plus two additional lines (the first and the last) to enclose that list of documents and the expected type of data is as follows:</p> <pre><code class="language-json">{"articles":[ {"abstractText":str,"db":str,"decsCodes":list,"id":str,"journal":str,"title":str,"year":int}, ... ]}</code></pre> <p>To clarify, the order of appearance of the fields in each document is as follows (note that this example it is pretty printed for readability purposes):</p> <pre><code class="language-json">{ "articles": [ { "abstractText": "Content of the abstract", "db": "Name of the source database", "decsCodes": [ "code1", "code2", "code3" ], "id": "Id of the document", "journal": "Name of the journal", "title": "Title of the document", "year": 2019 } ] }</code></pre> <p>Note: The fields "db", "journal" and "year" might be null.</p> <p>Copyright (c) 2020 Secretaría de Estado de Digitalización e Inteligencia Artificial</p>
Dataset for Automated Medical Transcription
<p>We generated this dataset to train a machine learning model for automatically generating psychiatric case notes from doctor-patient conversations. Since, we didn't have access to real doctor-patient conversations, we used transcripts from two different sources to generate audio recordings of enacted conversations between a doctor and a patient. We employed eight students who worked in pairs to generate these recordings. Six of the transcripts that we used to produce this recordings were hand-written by Cheryl Bristow and rest of the transcripts were adapted from Alexander Street which were generated from real doctor-patient conversations. Our study requires recording the doctor and the patient(s) in seperate channels which is the primary reason behind generating our own audio recordings of the conversations. </p> <p>We used Google Cloud Speech-To-Text API to transcribe the enacted recordings. These newly generated transcripts are auto-generated entirely using AI powered automatic speech recognition whereas the source transcripts are either hand-written or fine-tuned by human transcribers (transcripts from Alexander Street). </p> <p>We provided the generated transcripts back to the students and asked them to write case notes. The students worked independently using a software that we developed earlier for this purpose. The students had past experience of writing case notes and we let the students write case notes as they practiced without any training or instructions from us.</p> <p><strong>NOTE:</strong> Audio recordings are not included in Zenodo due to large file size but they are available in the <a href="https://github.com/nazmulkazi/dataset_automated_medical_transcription">GitHub</a> repository.</p>
Data relating to Chiedozie et al. How many medications do doctors in primary care use? An observational study of the DU90% indicator in primary care in England.
<p>Data in .csv format relating to the paper Chiedozie et al. 2020 "How many medications do doctors in primary care use? An observational study of the DU90% indicator in primary care in England." Also contains eTables 5-8 in Excel format, and Stata do file for deriving the DU90% indicator.</p>
Raw data acquired necessary to produce the plots introduced in the scientific paper: "Upper-limb kinematic reconstruction during stroke robot-aided therapy" (Medical & Biological Engineering & Computing)
<p>These files contain the raw data acquired necessary to produce the plots introduced the Figure 6 of the scientific paper: “Upper-limb kinematic reconstruction during stroke robot-aided therapy” (Medical & Biological Engineering & Computing).</p> <p>Fig. 6 shows the data recorded from two patients performing five forward/backward movements at InMotion2 robot before and after rehabilitation treatment. Mean values of the five execution have been reported in Fig. 6.</p>
MultiCaRe: An open-source clinical case dataset for medical image classification and multimodal AI applications
<p>The dataset contains multi-modal data from over 70,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.</p> <p>More than 90,000 patients and 280,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.</p> <p>Refer to the examples showcased in this <a href="https://github.com/mauro-nievoff/MultiCaRe_Dataset">GitHub repository</a> to understand how to optimize the use of this dataset.<br><br>The license of the dataset as a whole is CC BY-NC-SA. However, its individual contents may have less restrictive license types (CC BY, CC BY-NC, CC0). For instance, regarding image filess, 66K of them are CC BY, 32K are CC BY-NC-SA, 32K are CC BY-NC, and 20 of them are CC0.</p>
Clinical evidence for high-risk CE-marked medical devices for glucose management: a systematic review and meta-analysis
<p><strong><span>Aims: </span></strong><span>High-risk medical devices are increasingly used in diabetes management, but<span><span> there are no specific European recommendations on how they should be evaluated. Within</span></span> the Coordinating Research and Evidence for Medical Devices (CORE-MD) project, we conducted a systematic review and meta-analysis evaluating CE-marked high-risk devices for glucose management. </span></p> <p><strong><span>Materials and Methods: </span></strong><span>We<span> identified interventional and observational studies evaluating the </span>efficacy and safety of 8 automated insulin delivery (AID) systems, 2 implantable insulin pumps, and 3 implantable continuous glucose monitoring (CGM) devices.<span> </span></span><span>We meta-analysed randomized controlled trials (RCTs) comparing AID systems with other treatments.</span></p> <p><strong><span>Results:</span></strong><span> 99 studies published from 2009–2022 were included, comprising 83 on AID systems, 6 on insulin pumps, and 10 on CGM; 43% reported industry funding;30% were pre-market; 45% had a comparator group. 33% were RCTs, 25% non-randomized trials, and 41% observational studies. Median sample size was 52 (interquartile range 25–111), age 37.8 years (17–45.5), and study duration 13 weeks (4.5–26). AID systems lowered HbA1c by 0.3 percentage points (absolute mean difference [MD]=-0.3; 9 RCTs; I<sup>2</sup>=85%) and increased time in target range for sensor glucose level by 10.5 percentage points (MD=10.5; 14 RCTs; I<sup>2</sup>=89%). 69% of studies reported on at least one safety outcome.</span></p> <p><strong><span>Conclusions:</span></strong><span> High-risk devices for glucose monitoring or insulin dosing, in particular AID systems, improve glucose control safely but evidence on diabetes-related end organ damage is lacking due to short study durations. Methodological heterogeneity highlights the need for d<span><span>eveloping standards for future pre- and post-market investigations of diabetes-specific high-risk medical devices.</span></span></span></p>
MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish
<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets). </p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p> </p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud, the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others. </p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then, for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L – Scientific Literature: </strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials </strong></strong>contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. </li> <li><strong>[Subtrack 3 corpus] MESINESP-P – Patents: </strong>This corpus includes patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants' predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p> </p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized (folder <em>separated</em>). This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>. </p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train: <ul> <li>training_set_track1_all.json: Full training set for subtrack 1. </li> <li>training_set_track1_only_articles.json: Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json: </li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1. </li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json: Manually annotated development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json: Manually annotated development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip </strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020 set)</li> <li>List of synonyms (the descriptors and synonyms from Latin Spanish DeCS 2020 set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo </strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p> </p> <p><strong>Data format description</strong></p> <p>The <strong>input text files</strong> for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial >140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial >130/80 y <140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP <strong>entity mention files</strong> contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p> </p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L – Scientific Literature </strong>: <ul> <li><em><strong>Training set: </strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish. We have filtered out empty abstracts and non-Spanish abstracts. We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database. We distribute two different datasets: <ul> <li><strong>Articles training set: </strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set: </strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: <ul> <li><strong>Training set: </strong>The training dataset contains records from <a href="https://reec.aemps.es/reec/public/web.html">Registro Español de Estudios Clínicos (REEC)</a>. REEC doesn't provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC <a href="https://github.com/luisgasco/REECapi">API</a>. Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a <a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team. </li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set: </strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P – Patents: </strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code “A61P” and “A61K31”. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set: </strong>We provide a <strong>test set</strong> containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes “A61P” and “A61K31”.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li> We provide this information to the participants as additional data in the “Additional Data” folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p> </p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td> </td> <td> </td> <td> </td> <td> </td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p> </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2 Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p> </p> <p>For further information, please email us at luis.gasco@bsc.es</p>
WikiMed and PubMedDS: Two large-scale datasets for medical concept extraction and normalization research
<p>Two large-scale, automatically-created datasets of medical concept mentions, linked to the <a href="https://uts.nlm.nih.gov/uts/umls/home">Unified Medical Language System (UMLS)</a>.</p> <p><strong>WikiMed</strong></p> <p>Derived from Wikipedia data. Mappings of Wikipedia page identifiers to UMLS Concept Unique Identifiers (CUIs) was extracted by crosswalking Wikipedia, Wikidata, Freebase, and the NCBI Taxonomy to reach existing mappings to UMLS CUIs. This created a 1:1 mapping of approximately 60,500 Wikipedia pages to UMLS CUIs. Links to these pages were then extracted as mentions of the corresponding UMLS CUIs.</p> <p>WikiMed contains:</p> <ul> <li>393,618 Wikipedia page texts</li> <li>1,067,083 mentions of medical concepts</li> <li>57,739 unique UMLS CUIs</li> </ul> <p>Manual evaluation of 100 random samples of WikiMed found 91% accuracy in the automatic annotations at the level of UMLS CUIs, and 95% accuracy in terms of semantic type.</p> <p><strong>PubMedDS</strong></p> <p>Derived from biomedical literature abstracts from <a href="https://pubmed.ncbi.nlm.nih.gov/">PubMed</a>. Mentions were automatically identified using distant supervision based on Medical Subject Heading (MeSH) headers assigned to the papers in PubMed, and recognition of medical concept mentions using the high-performance <a href="https://allenai.github.io/scispacy/">scispaCy</a> model. MeSH header codes are included as well as their mappings to UMLS CUIs.</p> <p>PubMedDS contains:</p> <ul> <li>13,197,430 abstract texts</li> <li>57,943,354 medical concept mentions</li> <li>44,881 unique UMLS CUIs</li> </ul> <p>Comparison with existing manually-annotated datasets (NCBI Disease Corpus, BioCDR, and MedMentions) found 75-90% precision in automatic annotations. Please note this dataset is <em>not </em>a comprehensive annotation of medical concept mentions in these abstracts (only mentions located through distant supervision from MeSH headers were included), but is intended as data for <em>concept n</em><em>ormalization</em> research.</p> <p>Due to its size, PubMedDS is distributed as 30 individual files of approximately 1.5 million mentions each.</p> <p><strong>Data format</strong></p> <p>Both datasets use JSON format with one document per line. Each document has the following structure:</p> <pre><code class="language-json">{ "_id": "A unique identifier of each document", "text": "Contains text over which mentions are ", "title": "Title of Wikipedia/PubMed Article", "split": "[Not in PubMedDS] Dataset split: <train/test/valid>", "mentions": [ { "mention": "Surface form of the mention", "start_offset": "Character offset indicating start of the mention", "end_offset": "Character offset indicating end of the mention", "link_id": "UMLS CUI. In case of multiple CUIs, they are concatenated using '|', i.e., CUI1|CUI2|..." }, {} ] }</code></pre> <p><strong>Version history</strong></p> <table align="left"> <thead> <tr> <th scope="col">Version</th> <th scope="col">Notes</th> </tr> </thead> <tbody> <tr> <td>1.0.0</td> <td>Initial release</td> </tr> <tr> <td>1.0.1</td> <td>Corrected duplication error in WikiMed.zip file</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p>
Perspectives on Medical Education Journal Data and Supplemental Files (2012 - 2019)
<p>This is the supplemental data, figures, and tables for <em>Joining the meta-research movement: A bibliometric case study of Perspectives on Medical Education</em>. </p> <p>For Figures 2-4 from the manuscript, to open the network maps, use both the network and map file for each figure in VoS viewer - https://www.vosviewer.com/</p>
Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review - Metadata
<p>This file is the metadata related to the publication "Quantification of ADHD Medication in Biological Fluids with Liquid Chromatography: A Comprehensive Review".</p>
Data and materials for Wallace et al (2018) Self-report versus electronic medical record recorded healthcare utilisation in older community-dwelling adults: comparison of two prospective cohort studies v1.2
<p>This comprises the data and materials for the study: Wallace E, Moriarty F, McGarrigle C, Smith SM, Kenny RA, Fahey T. (2018) Self-report versus electronic medical record recorded healthcare utilisation in older community-dwelling adults: Comparison of two prospective cohort studies. PLOS ONE 13(10): e0206201. <a href="https://doi.org/10.1371/journal.pone.0206201">https://doi.org/10.1371/journal.pone.0206201</a></p> <p>The anonymised TILDA dataset is publicly available to researchers who meet the criteria for access, at no monetary cost, from the Irish Social Science Data Archive (ISSDA) at University College Dublin (<a href="https://emea01.safelinks.protection.outlook.com/?url=http%3A%2F%2Fwww.ucd.ie%2Fissda%2Fdata%2Ftilda%2F&data=02%7C01%7C%7Ccc2345c4f5c543bbcddc08d5fd28200c%7C607041e7a8124670bd3030f9db210f06%7C0%7C0%7C636693271128219875&sdata=%2Fcochi1RuRtYSUa5sF9uA%2BjOOoNYIg7DPpk0mZl5D2s%3D&reserved=0">http://www.ucd.ie/issda/data/tilda/</a>) and the Interuniversity Consortium for Political and Social Research (ICPSR) at the University of Michigan (<a href="https://emea01.safelinks.protection.outlook.com/?url=http%3A%2F%2Fwww.icpsr.umich.edu%2Ficpsrweb%2FICPSR%2Fstudies%2F34315&data=02%7C01%7C%7Ccc2345c4f5c543bbcddc08d5fd28200c%7C607041e7a8124670bd3030f9db210f06%7C0%7C0%7C636693271128219875&sdata=7LHSSqU8xotACMsalpAjVrV5m95DlapgQViyr4P%2FsXY%3D&reserved=0">http://www.icpsr.umich.edu/icpsrweb/ICPSR/studies/34315</a>). For the CPCR cohort, no provision for data sharing was included in the original ethical approval and participant consent form. As a minimal data set necessary to replicate the present study could not be deidentified due to the large number of demographic variables considered, a synthetic version of the study dataset was produced using the synthpop package in R: <a href="https://emea01.safelinks.protection.outlook.com/?url=https%3A%2F%2Fcran.r-project.org%2Fweb%2Fpackages%2Fsynthpop%2Findex.html&data=02%7C01%7C%7Ccc2345c4f5c543bbcddc08d5fd28200c%7C607041e7a8124670bd3030f9db210f06%7C0%7C0%7C636693271128229884&sdata=j3If%2FNe%2F1eGsGAt9hyg3ICMqmLec4aOrjRVKppaRSFU%3D&reserved=0">https://cran.r-project.org/web/packages/synthpop/index.html</a>. This dataset and the analytical code for the present study are presented here. Code developed on the synthetic data can be sent to frankmoriarty@rcsi.ie or <a href="mailto:enquiries.cpcr@rcsi.ie">enquiries.cpcr@rcsi.ie</a> to be run on the original data.</p> <p>v1.2 includes a more detailed description of how the dataset was synthesised.</p>
Data for Weighted Manifold Alignment using Wave Kernel Signatures for Aligning Medical image Datasets
<p>Data used in MRI experiments in paper 'Weighted Manifold Alignment using Wave Kernel Signatures for Aligning Medical image Datasets'. For each volunteer, breath-hold data (folder bhs) and dynamic free-breathing (folder dyn) data is provided in NIFTI format.</p>
Analysis of the type of license for medical dissertation (Lublin 2011-2019)
<p>Analysis of the type of license for medical dissertation (Lublin 2011-2019).</p> <p>The data comes from the Digital Library of the Medical University of Lublin and the Internal Digital Library (2011-2019 to number 90/2019).</p> <p>Percentage of authors who both shared their dissertation and consented to its copying by users - (authors potentially open to sharing work under CC or similar licenses) - "extended" permitted use; percentage of authors who shared their dissertation, but without permission to copy it - only limited use complying with fair use doctrine; percentage of authors who chose access to their doctoral dissertation only in the Library (maximum access limitation, no online access)</p>
MeSDiCon - Medical Spanish Disease and symptom name Collection lexicon (unfiltered initial version)
<p>The MeSDiCon - (Medical Spanish Disease and symptom name Collection lexicon) consists of a list or gazetteer of candidate names of diseases and symptoms mentioned in Spanish clinical texts. Thus MeSDiCon serves as a lexical resource or dictionary for automatic detection of disease/symptom mentions, as well as indexing or classification of medical texts with such concept types.</p> <p>This collection was generated in a five step procedure:</p> <ol> <li>Automatic detection of mentions of disease/symptom terms in biomedical texts in English (including mapping/normalization to MeSH terms or OMIM identifiers).</li> <li>Generation of a unique name list from the detected concept mentions.</li> <li>Basic filtering of non- disease/symptom names or highly ambiguous mentions-abbreviations using basic characteristics like name morphology and length criteria.</li> <li>Automatic translation of name lists form English to Spanish using a medical machine translation system (see Soares, F. and Krallinger, M. BSC Participation in the WMT Translation of Biomedical Abstracts. In <em>Proceedings of the Fourth Conference on Machine Translation, Volume 3: Shared Task Papers, </em>pp. 175-178 2019; https://zenodo.org/record/3346802)</li> <li>Automatic mention lookup of translated names in a collection of 20 million Spanish clinical notes (primary care and pediatrics).</li> </ol> <p>Every term in MeSDiCon is identified by a text span (in Spanish), a target terminology namespace to which it was automatically mapped (MeSH or OMIM) and its corresponding concept identifier in that target terminology. Moreover, we provide for every text span the absolute term frequency, i.e. the number of matches in the corpus of 20 million clinical notes and the number of documents or notes in which it was automatically.</p> <p>Important note: no manual filtering of the MeSDiCon was carried out, implying that some entries might comprise errors, either due to the initial name recognition and concept mapping in English or due to wrong automatic translations into Spanish.</p> <p>The MeSDiCon resource is provided in two formats:</p> <ul> <li>TSV. Data is separated by tabs (\t). Every row of the file has the following fields:</li> </ul> <pre><code>terminology identifier translatedTerm termCount documentCount</code></pre> <ul> <li>JSON. Records are stored as a list of JSON objects. They have the following fields:</li> </ul> <pre><code>{ "terminology":"MESH", "identifier":"D025861", "translatedTerm":"Trastornos de la coagulación", "termFrequency":9, "documentFrequency":9 }</code></pre> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
MeSCCon - Medical Spanish Chemical compound, drug and medication Name Lexicon (unfiltered version)
<p>The MeSCCon (Medical Spanish Chemical compound, drug and medication Name Lexicon) consists of a list or gazetteer of candidate names of chemicals, drugs, and medications mentioned in Spanish clinical texts. Thus MeSCCon serves as a lexical resource or dictionary for automatic detection of chemical/drug mentions, as well as indexing or classification of medical texts with such concept types.</p> <p>This collection was generated in a five step procedure:</p> <ol> <li>Automatic detection of mentions of chemicals and drugs in biomedical texts in English (including mapping/normalization to MeSH terms or ChEBI identifiers).</li> <li>Generation of a unique name list from the detected concept mentions.</li> <li>Basic filtering of non-chemical names or highly ambiguous mentions-abbreviations using basic characteristics like name morphology and length criteria.</li> <li>Automatic translation of name lists from English to Spanish using a medical machine translation system (see Soares, F. and Krallinger, M. BSC Participation in the WMT Translation of Biomedical Abstracts. In <em>Proceedings of the Fourth Conference on Machine Translation, Volume 3: Shared Task Papers, </em>pp. 175-178 2019; https://zenodo.org/record/3346802)</li> <li>Automatic mention lookup of translated names in a collection of 20 million Spanish clinical notes (primary care and pedriatrics).</li> </ol> <p>Every term in MeSCCon is identified by a text span (in Spanish), a target terminology namespace to which it was automatically mapped (MeSH or ChEBI) and the corresponding concept identifier in that terminology.</p> <p>Moreover, we provide for every text span the absolute term frequency, i.e. the number of matches in the corpus of 20 million clinical notes and the number of documents or notes in which it was found.</p> <p>Important note: no manual filtering of the MeSCCon was carried out, implying that some entries might comprise errors, either due to the initial name recognition and concept mapping in English or due to wrong automatic translations into Spanish.</p> <p>The MeSCCon resource is provided in two formats:</p> <ul> <li>TSV. Data is separated by tabs (\t). Every row of the file has the following fields:</li> </ul> <pre><code>terminology identifier translatedTerm termCount documentCount</code></pre> <ul> <li>JSON. Records are stored as a list of JSON objects. They have the following fields:</li> </ul> <pre><code class="language-javascript">{ "terminology":"MESH", "identifier":"D009020", "translatedTerm":"clorhidrato de morfina", "termFrequency":1, "documentFrequency":1 }</code></pre> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.