Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

4

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

4 results for “Medical Corpus”

Learn how ShareScore rates datasets ↗
zenodo40/100

MEDDOCAN corpus: gold standard annotations for Medical Document Anonymization on Spanish clinical case reports

<p><strong>Intro:</strong></p> <p>Meddocan shared task dataset (divided in train, dev and test). In addition, we include here the Meddocan background set.</p> <p>It contains the training, development and test sets of the Meddocan shared task with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p>&nbsp;</p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 98%&nbsp;</p> <p>For more information, see the <a href="http://ceur-ws.org/Vol-2421/MEDDOCAN_overview.pdf">paper</a>.&nbsp;</p> <p>&nbsp;</p> <p><strong>Format:</strong></p> <p>Annotations are distributed in Brat format. See&nbsp;<a href="https://brat.nlplab.org/standoff.html">Brat webpage</a>&nbsp;for more information.</p> <p>In addition, annotations are also distributed in XML format (based on i2b2 XML format).</p> <p>In the <a href="https://temu.bsc.es/meddocan/index.php/resources/">Meddocan webpage</a>, there is a script to convert&nbsp;between MEDDOCAN-Brat, MEDDOCAN-XML, and i2b2 formats.</p> <p>&nbsp;</p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations&nbsp;given only the plain text files.&nbsp;</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/meddocan/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>Montserrat Marimon et al. &ldquo;Automatic De-identification of Medical Texts in Spanish: the MEDDOCAN Track, Corpus, Guidelines, Methods and Evaluation of Results.&rdquo; In: IberLEF@ SEPLN. 2019, pp. 618&ndash;638.</li> <li><strong>Silver Standard corpus</strong></li> <li><a href="https://doi.org/10.5281/zenodo.4279337"><strong>Annotation guidelines</strong></a></li> </ul> <p>&nbsp;</p> <p>For further information, please visit&nbsp;<a href="https://temu.bsc.es/meddocan/">https://temu.bsc.es/meddocan/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital (SEAD)</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

The Chilean Waiting List sub-Corpus with medical entities normalized to UMLS terminology

<p>A collection of 2000 medical referrals from the Chilean Waiting List Corpus, manually annotated with six entity types (Finding, Procedure, Disease, Family Member, Body Part, and Medication) and manually normalized to the Unified Medical Language System (UMLS).</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction

<p><strong>MEDDOPLACE</strong>&nbsp;stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.</p> <p>This repository includes the corpus' <strong>train and test sets</strong> in multiple formats, as well as the <strong>SNOMED gazetteer</strong>, <strong>cross-mapping</strong> between SNOMED and MeSH and the <strong>multilingual silver standard in 8 languages&nbsp;</strong>(Catalan, English, French, Italian, Dutch, Portuguese, Romanian and Swedish). For more information, please check the attached README file.</p> <p>MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a>.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-L&oacute;pez, Eul&agrave;lia Farr&eacute;-Maduell, Antonio Miranda-Escalada, Vicent Briv&aacute;-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoplace, title={MEDDOPLACE Shared Task overview: recognition, normalization and classification of locations and patient movement in clinical texts}, author={Lima-L&oacute;pez, Salvador and Farr&eacute;-Maduell, Eul&agrave;lia and Briv&aacute;-Iglesias, Vicent and Gasco-Sanchez, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {71}, year={2023}, issn = {1135-5948},<br>DOI = {10.26342/2023-71-23}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561/3961}, pages = {301--311} }</code></pre> <p><strong>Related Links:</strong></p> <p>- MEDDOPLACE website: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a></p> <p>- MEDDOPLACE overview paper: <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561">http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561</a></p> <p>- Annotation Guidelines (Spanish): <a href="https://doi.org/10.5281/zenodo.7775234">https://doi.org/10.5281/zenodo.7775234</a></p> <p>- Annotation Guidelines (English): <a href="https://doi.org/10.5281/zenodo.7928145">https://doi.org/10.5281/zenodo.7928145</a></p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-L&oacute;pez (&lt;salvador [dot] limalopez [at] gmail [dot] com&gt;)<br>- Martin Krallinger (&lt;krallinger [dot] martin [at] gmail [dot] com&gt;)</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

MEDDOPROF corpus: complete gold standard annotations for occupation detection in medical documents in Spanish

<p><strong>UPDATE 27/09/2022: </strong>A complete normalization of all mentions in the corpus to SNOMED CT has been added to the &#39;meddoprof-norm.tsv&#39; file.</p> <p><strong>Description</strong></p> <p>This repository contains the complete MEDDOPROF Gold Standard, a collection of 1,844 clinical cases in Spanish with annotations for occupations, working statuses and activities. MEDDOPROF is a Shared Task celebrated in 2021 that explores the application of natural language processing to occupational health. If you&#39;d like to learn more, please visit: <a href="https://temu.bsc.es/meddoprof">https://temu.bsc.es/meddoprof</a>.</p> <p><strong>Folder and File Structure</strong></p> <p>The corpus&#39; files are presented in the format used by the annotation tool brat. That is, for each clinical case there is a .txt file with the text and a .ann file with its corresponding annotations.</p> <p><em>- meddoprof-ner/</em></p> <p>Clinical cases annotated with these labels: PROFESION (PROFESSION), SITUACION_LABORAL (WORKING_STATUS) or ACTIVIDAD (ACTIVIDAD).</p> <p><em>- meddoprof-class/</em></p> <p>Clinical cases with the same annotations as &#39;meddoprof-ner&#39; but with these labels instead: PACIENTE (patient), FAMILIAR (family member), SANITARIO (health professional) or OTRO (other).</p> <p><em>- ner_class_joint/</em></p> <p>Clinical cases with both levels of annotation (ner and class) joint (that is, a mention classified as as PROFESOR in meddoprof-ner and as PACIENTE in meddoprof-class would be PROFESION-PACIENTE here).</p> <p><em>- meddoprof-norm.tsv</em></p> <p>Tab-separated file (.tsv) with the mapping of each mention in the corpus to ESCO and SNOMED CT. The file has five columns: filename, mention text, span, ESCO code and SNOMED code.</p> <p>Additionally, two files with the filenames of the train and test partitions are included.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-L&oacute;pez, Eul&agrave;lia Farr&eacute;-Maduell, Antonio Miranda-Escalada, Vicent Briv&aacute;-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoprof, title={NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Miranda-Escalada, Antonio and Brivá-Iglesias, Vicent and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {67}, year={2021}, issn = {1989-7553}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6393}, pages = {243--256} }</code></pre> <p><strong>Related Resources:</strong></p> <p>- <a href="http://temu.bsc.es/meddoprof">Web</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4694768">Training Data</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4889776">Test set</a></p> <p>- <a href="https://zenodo.org/record/4722741">Codes Reference List</a> (for MEDDOPROF-NORM)</p> <p>- <a href="https://zenodo.org/record/4720833">Annotation Guidelines</a></p> <p>- <a href="https://doi.org/10.5281/zenodo.4524658">Occupations Gazetteer</a></p> <p>&nbsp;</p> <blockquote> <p>MEDDOPROF is part of the IberLEF 2021 workshop, which is co-located with the SEPLN 2021 conference. For further information, please visit&nbsp;<a href="https://temu.bsc.es/meddoprof/">https://temu.bsc.es/meddoprof/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>MEDDOPROF is promoted by the Plan de Impulso de las Tecnolog&iacute;as del Lenguaje de la Agenda Digital (Plan TL) and the Spanish government&#39;s 2020 Proyectos de I+D+i RTI Tipo A (AI4PROFHEALTH - DESCIFRANDO EL PAPEL DE LAS PROFESIONES EN LA SALUD DE LOS PACIENTES A TRAVES DE LA MINERIA DE TEXTOS (PID2020-119266RA-I00)).</p> </blockquote>

opencc-by-4.0Jul 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record