Skip to main content
zenodoopen

Cantemist corpus: gold standard of oncology clinical cases annotated with CIE-O 3 terminology

<p><strong>Intro:</strong></p> <p>Cantemist shared task dataset (divided in train, dev1, dev2 and test). In addition, we include here the Cantemist background set.</p> <p>It contains the train, development and test sets of the three subtasks: cantemist-ner, cantemist-norm and cantemist-coding with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p>&nbsp;</p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Miranda-Escalada, A., Farr&eacute;, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In&nbsp;<em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</p> <pre><code>@inproceedings{miranda2020named, title={Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results}, author={Miranda-Escalada, A and Farr{\'e}, E and Krallinger, M}, booktitle={Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings}, year={2020} }</code></pre> <p>&nbsp;</p> <p><strong>Format:</strong></p> <p>For subtasks cantemist-norm and cantemist-ner, annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information</p> <p>For subtask cantemist-coding, codes are grouped in a TSV file with the following columns (this follows the format used in <a href="https://temu.bsc.es/codiesp/">CodiEsp shared task</a>):&nbsp;</p> <blockquote> <p>filename&nbsp;&nbsp; &nbsp;code</p> </blockquote> <p>&nbsp;</p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations (either the ANN files or the TSV with the codes) given only the plain text files.&nbsp;</p> <p>&nbsp;</p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/cantemist/">Web</a></strong></li> <li><strong><a href="http://ceur-ws.org/Vol-2664/cantemist_overview.pdf">Citation</a>:&nbsp;</strong>Miranda-Escalada, A., Farr&eacute;, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In&nbsp;<em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><strong><a href="https://doi.org/10.5281/zenodo.4010899">Silver Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p>&nbsp;</p> <p>For further information, please visit <a href="https://temu.bsc.es/cantemist/">https://temu.bsc.es/cantemist/</a> or email us at encargo-pln-life@bsc.es</p>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
0
Engagement
4

Topics