Cantemist corpus: gold standard of oncology clinical cases annotated with CIE-O 3 terminology
<p><strong>Intro:</strong></p> <p>Cantemist shared task dataset (divided in train, dev1, dev2 and test). In addition, we include here the Cantemist background set.</p> <p>It contains the train, development and test sets of the three subtasks: cantemist-ner, cantemist-norm and cantemist-coding with Gold Standard annotations.</p> <p>In addition, it contains the documents of the background set, without annotations.</p> <p> </p> <p><strong>Please cite if you use this dataset:</strong></p> <p>Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</p> <pre><code>@inproceedings{miranda2020named, title={Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results}, author={Miranda-Escalada, A and Farr{\'e}, E and Krallinger, M}, booktitle={Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings}, year={2020} }</code></pre> <p> </p> <p><strong>Format:</strong></p> <p>For subtasks cantemist-norm and cantemist-ner, annotations are distributed in Brat format. See <a href="https://brat.nlplab.org/standoff.html">Brat webpage</a> for more information</p> <p>For subtask cantemist-coding, codes are grouped in a TSV file with the following columns (this follows the format used in <a href="https://temu.bsc.es/codiesp/">CodiEsp shared task</a>): </p> <blockquote> <p>filename code</p> </blockquote> <p> </p> <p><strong>Shared task goal:</strong></p> <p>In the three subtasks, the goal will be to predict the annotations (either the ANN files or the TSV with the codes) given only the plain text files. </p> <p> </p> <p><strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/cantemist/">Web</a></strong></li> <li><strong><a href="http://ceur-ws.org/Vol-2664/cantemist_overview.pdf">Citation</a>: </strong>Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>.</li> <li><strong><a href="https://doi.org/10.5281/zenodo.4010899">Silver Standard corpus</a></strong></li> <li><strong><a href="https://doi.org/10.5281/zenodo.3878178">Annotation guidelines</a></strong></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhC24g5dsp5eVMp8BZFWCraX"><strong>YouTube presentations</strong></a></li> <li><a href="https://temu.bsc.es/cantemist/?p=4606"><strong>Participant codes</strong></a></li> </ul> <p> </p> <p>For further information, please visit <a href="https://temu.bsc.es/cantemist/">https://temu.bsc.es/cantemist/</a> or email us at encargo-pln-life@bsc.es</p>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4