Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Medical Corpus”
MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports
<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98% </p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>. </p> <p> </p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p> </p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations given only the plain text files. </p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation: </strong>Montserrat Marimon et al. “Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.” In: IberLEF@ SEPLN. 2019, pp. 618–638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a> or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital (SEAD)</p>
The Chilean Waiting List sub-Corpus with medical entities normalized to UMLS terminology
<p>A collection of 2000 medical referrals from the Chilean Waiting List Corpus, manually annotated with six entity types (Finding, Procedure, Disease, Family Member, Body Part, and Medication) and manually normalized to the Unified Medical Language System (UMLS).</p>
MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction
<p><strong>MEDDOPLACE</strong> stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.</p> <p>This repository includes the corpus' <strong>train and test sets</strong> in multiple formats, as well as the <strong>SNOMED gazetteer</strong>, <strong>cross-mapping</strong> between SNOMED and MeSH and the <strong>multilingual silver standard in 8 languages </strong>(Catalan, English, French, Italian, Dutch, Portuguese, Romanian and Swedish). For more information, please check the attached README file.</p> <p>MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a>.</p> <p> </p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-López, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoplace, title={MEDDOPLACE Shared Task overview: recognition, normalization and classification of locations and patient movement in clinical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Brivá-Iglesias, Vicent and Gasco-Sanchez, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {71}, year={2023}, issn = {1135-5948},<br>DOI = {10.26342/2023-71-23}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561/3961}, pages = {301--311} }</code></pre> <p><strong>Related Links:</strong></p> <p>- MEDDOPLACE website: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a></p> <p>- MEDDOPLACE overview paper: <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561">http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561</a></p> <p>- Annotation Guidelines (Spanish): <a href="https://doi.org/10.5281/zenodo.7775234">https://doi.org/10.5281/zenodo.7775234</a></p> <p>- Annotation Guidelines (English): <a href="https://doi.org/10.5281/zenodo.7928145">https://doi.org/10.5281/zenodo.7928145</a></p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
MEDDOPROF corpus: complete gold standard annotations for occupation detection in medical documents in Spanish
<p><strong>UPDATE 27/09/2022: </strong>A complete normalization of all mentions in the corpus to SNOMED CT has been added to the 'meddoprof-norm.tsv' file.</p> <p><strong>Description</strong></p> <p>This repository contains the complete MEDDOPROF Gold Standard, a collection of 1,844 clinical cases in Spanish with annotations for occupations, working statuses and activities. MEDDOPROF is a Shared Task celebrated in 2021 that explores the application of natural language processing to occupational health. If you'd like to learn more, please visit: <a href="https://temu.bsc.es/meddoprof">https://temu.bsc.es/meddoprof</a>.</p> <p><strong>Folder and File Structure</strong></p> <p>The corpus' files are presented in the format used by the annotation tool brat. That is, for each clinical case there is a .txt file with the text and a .ann file with its corresponding annotations.</p> <p><em>- meddoprof-ner/</em></p> <p>Clinical cases annotated with these labels: PROFESION (PROFESSION), SITUACION_LABORAL (WORKING_STATUS) or ACTIVIDAD (ACTIVIDAD).</p> <p><em>- meddoprof-class/</em></p> <p>Clinical cases with the same annotations as 'meddoprof-ner' but with these labels instead: PACIENTE (patient), FAMILIAR (family member), SANITARIO (health professional) or OTRO (other).</p> <p><em>- ner_class_joint/</em></p> <p>Clinical cases with both levels of annotation (ner and class) joint (that is, a mention classified as as PROFESOR in meddoprof-ner and as PACIENTE in meddoprof-class would be PROFESION-PACIENTE here).</p> <p><em>- meddoprof-norm.tsv</em></p> <p>Tab-separated file (.tsv) with the mapping of each mention in the corpus to ESCO and SNOMED CT. The file has five columns: filename, mention text, span, ESCO code and SNOMED code.</p> <p>Additionally, two files with the filenames of the train and test partitions are included.</p> <p> </p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-López, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoprof, title={NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Miranda-Escalada, Antonio and Brivá-Iglesias, Vicent and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {67}, year={2021}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6393}, pages = {243--256} }</code></pre> <p><strong>Related Resources:</strong></p> <p>- <a href="http://temu.bsc.es/meddoprof">Web</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4694768">Training Data</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4889776">Test set</a></p> <p>- <a href="https://zenodo.org/record/4722741">Codes Reference List</a> (for MEDDOPROF-NORM)</p> <p>- <a href="https://zenodo.org/record/4720833">Annotation Guidelines</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4524658">Occupations Gazetteer</a></p> <p> </p> <blockquote> <p>MEDDOPROF is part of the IberLEF 2021 workshop, which is co-located with the SEPLN 2021 conference. For further information, please visit <a href="https://temu.bsc.es/meddoprof/">https://temu.bsc.es/meddoprof/</a> or email us at encargo-pln-life@bsc.es</p> <p>MEDDOPROF is promoted by the Plan de Impulso de las Tecnologías del Lenguaje de la Agenda Digital (Plan TL) and the Spanish government's 2020 Proyectos de I+D+i RTI Tipo A (AI4PROFHEALTH - DESCIFRANDO EL PAPEL DE LAS PROFESIONES EN LA SALUD DE LOS PACIENTES A TRAVES DE LA MINERIA DE TEXTOS (PID2020-119266RA-I00)).</p> </blockquote>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.