Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “CodiEsp”
CodiEsp corpus: gold standard Spanish clinical cases coded in ICD10 (CIE10) - eHealth CLEF2020
<p><strong>Introduction</strong></p> <p>These are the train, development and test sets of the CodiEsp corpus. Train, development and test have gold standard annotations. In addition, the unannotated background set is also distributed. All documents are released in the context of the CodiEsp track for CLEF ehealth 2020 (<a href="http://temu.bsc.es/codiesp/">http://temu.bsc.es/codiesp/</a>).</p> <p>The CodiEsp corpus contains manually coded clinical cases. All documents are in Spanish language and CIE10 is the coding terminology (it is the Spanish version of ICD10-CM and ICD10-PCS). The CodiEsp corpus has been randomly sampled into three subsets: the train, the development, and the test set. The train set contains 500 clinical cases, and the development and test set 250 clinical cases each. CodiEsp participants must submit predictions for the test and background set, but they will only be evaluated on the test set.</p> <p> </p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p><strong>Annotation quality</strong></p> <p>Inter-annotator agreement: 88.6% for diagnosis coding, 88.9% for procedure coding and 80.5% for the textual reference annotation. For more information, see the <a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">paper</a>.</p> <p><br> <strong>Zip structure</strong><br> Four folders: train, dev, test and background. Each one of them contains the files for the train, development, test and background corpora, respectively.</p> <ul> <li><strong>train, dev and test</strong> folders have: <ul> <li>3 tab-separated files with the annotation information relevant for each of the 3 sub-tracks of CodiEsp. </li> <li>A subfolder named <em>text_files</em> with the plain text files of the clinical cases.</li> <li>A subfolder named <em>text_files_en</em> with the plain text files machine-translated to English. Due to the translation process, the text files are sentence-splitted.</li> </ul> </li> <li>The <strong>background</strong> folder has only <em>text_files</em> and <em>text_files_en</em> subfolders with the plain text files.</li> </ul> <p><br> <strong>Format</strong><br> The CodiEsp corpus is distributed in plain text in UTF8 encoding, where each clinical case is stored as a single file whose name is the clinical case identifier. Annotations are released in a tab-separated file. Since the CodiEsp track has 3 sub-tracks, every set of documents (train and test) has 3 tab-separated files associated with it. </p> <p>For the sub-tracks CodiEsp-D and CodiEsp-P, the file has the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p><br> <strong>Corpus summary statistics</strong><br> The final collection of 1000 clinical cases that make up the corpus had a total of 16504 sentences, with an average of 16.5 sentences per clinical case. It contains a total of 396,988 words, with an average of 396.2 words per clinical case.</p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong><a href="http://ceur-ws.org/Vol-2696/paper_263.pdf">Citation</a>: </strong>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</li> <li><strong><a href="https://doi.org/10.5281/zenodo.3859869">Silver Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3730566">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhA0crlSVCYMPqMUWd4mXc4x"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/codiesp/index.php/participants-systems/"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>For more information, visit the track webpage: http://temu.bsc.es/codiesp/ or email us at encargo-pln-life@bsc.es</p> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
CodiEsp-abstracs: Abstracts from Lilacs and Ibecs with ICD10 codes
<p>JSON file with abstracts from Lilacs and Ibecs with ICD10 codes (ICD10-CM and ICD10-PCS) associated to them (CIE10 in Spanish).</p> <p> </p> <p><strong>Please, cite us:</strong></p> <p>Miranda-Escalada, A., Gonzalez-Agirre, A., Armengol-Estapé, J., Krallinger, M.: Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of eHealth CLEF 2020.<em> In: CLEF (Working Notes) (2020)</em></p> <pre><code>@inproceedings{miranda2020overview, </code> <code>title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, </code> <code>author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, </code> <code>booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, </code> <code>year={2020} }</code></pre> <p> </p> <p>Lilacs and Ibecs databases have MeSH terms describing some of their documents. Then, using UMLS Metathesaurus, those MeSH terms have been translated into ICD10 codes (ICD10-CM and ICD10-PCS). Every abstract have at least one ICD10 code. </p> <p>In addition, MeSH codes given by the databases (Lilacs and Ibecs) have a "word" describing them. These "words" have been used to add further ICD10 codes. We have done strict string matching to find whether those "words" were a descriptor of any ICD10 code (in the Spanish version, CIE10).</p> <p>The format of the JSON file is the following:</p> <pre>{'articles': [{'title': 'title', 'pmid': 'pmid', 'abstractText': 'abtract (in Spanish)', 'Mesh': [{'Code': 'MeSHCode', 'Word': 'reference', 'CIE': [CIE10_1, CIE10_2, ...]}, ...] }, ...] }</pre> <p> </p> <p>Additionally, the compressed file includes a folder with all the abstracts extracted in individual UTF-8 encoded text files and a tab-separated file with 4 fields:</p> <blockquote> <p>pmid label cie10-code word</p> </blockquote> <p>Summary statistics:</p> <ul> <li>number of abstracts: 355 840</li> <li>number abstracts with at least one ICD10 code: 176 294</li> <li>Percentage of MeSH codes mapped to ICD10: 10.6% (there were 2 526 772 MeSH codes and 266 949 mapped to ICD10)</li> <li>average number of MeSH codes per article: 7.1</li> <li>average number of ICD10 codes per article: 2.5</li> <li>number of ICD10 codes that have an associated MeSH code in UMLS: 3293</li> <li>number of ICD10 codes that have an associated MeSH code in UMLS and appear in this dataset: 3082</li> </ul>
MeSDiCon subset for CodiEsp: MESH terms in MeSDiCon mapped to ICD10 CM and ICD10 PCS
<p>The MeSDiCon consists of a list or gazetteer of candidate names of diseases and symptoms mentioned in Spanish clinical texts. Thus MeSDiCon serves as a lexical resource or dictionary for automatic detection of disease/symptom mentions, as well as indexing or classification of medical texts with such concept types. Terms in MeSDiCon were mapped to MESH terminology.</p> <p>In this subset, we have mapped MESH codes to ICD10-CM and ICD10-PCS through UMLS Metathesaurus. Then, this resource contains diseases and symptoms terms from Spanish clinical texts mapped to MESH and ICD10.</p> <p> </p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p><strong>File structure</strong></p> <p>TSV. Data is separated by tabs (\t). Every row of the file has the following fields:</p> <pre><code>terminology identifier translatedTerm termCount documentCount ICD10CM-code ICD10PCS-code</code></pre> <p>In case one MESH term is mapped to more than one ICD10 code, they are separated by commas.</p>
CodiEsp codes: list of valid CIE10 codes for the CodiEsp task
<p><strong>Please cite if you use this dataset:</strong></p> <p>Antonio Miranda-Escalada, Aitor Gonzalez-Agirre, Jordi Armengol-Estapé and Martin Krallinger. Overview of automatic clinical coding: annotations, guidelines, and solutions for non-English clinical cases at CodiEsp track of CLEF eHealth 2020. In CLEF (Working Notes). 2020</p> <pre><code>@inproceedings{miranda2020overview, title={Overview of automatic clinical coding: annotations, guidelines, and solutions for non-english clinical cases at codiesp track of CLEF eHealth 2020}, author={Miranda-Escalada, Antonio and Gonzalez-Agirre, Aitor and Armengol-Estap{\'e}, Jordi and Krallinger, Martin}, booktitle={Working Notes of Conference and Labs of the Evaluation (CLEF) Forum. CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p>This compressed folder contains two files:</p> <p> + codiesp-D_codes.tsv: list of CIE10-Diagnósticos terms (2018 version) with their description in Spanish and in English.<br> + codiesp-P_codes.tsv: list of CIE10-Procedimiento terms (2018 version) with their description in Spanish and in English. In addition, the list also contains the codes until the 4th axis, which are also used in the CodiEsp-P track due to annotation reasons.</p> <p>A limited number of codes do not have an English description because they were removed from the English version but maintained in the Spanish version of the terminology.</p> <p> </p> <p>Format: <br> Tab-separated files with 3 columns<br> code es-description en-description</p> <p> </p> <p>Spanish to English description mapping:<br> For CodiEsp-D, the mapping to the English description was done through the files in the National Center for Health Statistics webpage: https://www.cdc.gov/nchs/icd/icd10cm.htm<br> Specifically, the file used was: ftp://ftp.cdc.gov/pub/Health_Statistics/NCHS/Publications/ICD10CM/2018/2018-ICD-10-CM-Codes-File.zip/icd10cm_codes_2018.txt</p> <p>For CodiEsp-P, the mapping to the English description was done through the files in the Centers for Medicare Services webpage: https://www.cms.gov/Medicare/Coding/ICD10/2018-ICD-10-PCS-and-GEMs<br> Specifically, the file used was: 2018_icd10pcs_codes_file.zip/icd10pcs_codes_2018.txt</p>
CodiEsp Silver Standard: Participant predictions in eHealth CLEF2020 - Spanish clinical cases coded in ICD10 (CIE10)
<p><strong>Introduction</strong></p> <p>Predictions in the background set of <a href="https://temu.bsc.es/codiesp/">eHealth CLEF 2020 Task 1</a> participants.</p> <p> </p> <p><strong>Zip structure</strong></p> <p>One directory per CodiEsp subtask. Within each CodiEsp subtask directory, there is one directory per team that contains the prediction runs.</p> <p> </p> <p><strong>Format</strong><br> The text documents are distributed in plain text files, UTF-8 encoding.<br> The CodiEsp Silver Standard annotations have the following format:</p> <p>For the sub-tracks CodiEsp-Diagnostic and CodiEsp-Procedure, the file files have the following fields:</p> <pre>articleID ICD10-code </pre> <p>Tab-separated files for the sub-track CodiEsp-X (explainability) contain extra fields that provide the text-reference and its position:</p> <pre>articleID label ICD10-code text-reference reference-position</pre> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/codiesp/">Web</a></strong></li> <li><strong>Citation: </strong>Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><strong><a href="https://zenodo.org/record/3837305#.X7T9KVlKg5k">Gold Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>All credit to CodiEsp participants</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.