Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
42
datasets available to search
ShareScore release 0.9.0
Dataset results
42 results for “Named Entities”
Named Entity Recognition Dataset for Dutch Biographical Texts
<p>A dataset for Named Entity Recognition for Dutch biographies. The original data is available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/). The annotations are for 6 types of entities: PERSON, LOCATION, ORGANIZATION, DATE, ARTWORK, MISC. Additionally, the CoNLL formatted files were manually checked for tokenization and sentence splitting.</p>
Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes
<p>This contains the merged dataset as described in the work "<strong>Multi-head CRF classifier for biomedical multi-class Named Entity Recognition on Spanish clinical notes"</strong>.</p> <p>This dataset consists of 4 seperate datasets:</p> <ul> <li><a href="../records/8224056" target="_blank" rel="noopener">MedProcNer</a></li> <li><a href="../records/7614764" target="_blank" rel="noopener">DisTEMIST</a></li> <li><a href="../records/4270158" target="_blank" rel="noopener">PharmaCoNER</a></li> <li><a href="../records/10635215" target="_blank" rel="noopener">SympTEMIST</a></li> </ul> <p>The dataset contains two tasks:</p> <p><strong>Task 1:</strong> This task is related to multi-class Named Entity Recognition. This dataset contains 5 possible classes: SYMPTOM, PROCEDURE, DISEASE, CHEMICAL and PROTEIN.</p> <p><strong>Task 2:</strong> This task is related to Named Entity Linking, where each code corresponds to a code within the SNOMED-CT corpus. The exact corpus used can be obtained <a href="https://download.nlm.nih.gov/umls/kss/IHTSDO20190131/SnomedCT_SpanishRelease-es_PRODUCTION_20190430T120000Z.zip" target="_blank" rel="noopener">here</a>. Further for the MedProcNER, SympTEMIST and DisTEMIST datasets, a gazetteer is provided in the original datasets. </p> <p>For more information on the construction of the dataset, aswell as dataloaders, we refer you to our <a href="https://github.com/ieeta-pt/Multi-Head-CRF" target="_blank" rel="noopener">GitHub repository</a>.<br><br>Further this also contains the embeddings from the <a href="https://huggingface.co/cambridgeltl/SapBERT-UMLS-2020AB-all-lang-from-XLMR-large" target="_blank" rel="noopener">SapBERT</a> model.</p> <p><strong>Please, cite:</strong></p> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <div> <blockquote> <div>@article{jonker2024a, title = {Multi-head {{CRF}} classifier for biomedical multi-class named entity recognition on {{Spanish}} clinical notes}, author = {Jonker, Richard A. A. and Almeida, Tiago and Antunes, Rui and Almeida, Jo{\~a}o R. and Matos, S{\'e}rgio}, year = {2024}, journal = {Database}, publisher = {Oxford University Press} }</div> </blockquote> <div>Jonker, R. A. A., Almeida, T., Antunes, R., Almeida, J. R., & Matos, S. (2024). Multi-head CRF classifier for biomedical multi-class named entity recognition on Spanish clinical notes. (Submitted.) </div> <div> </div> <div> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div> </div>
Named Entity Corpus for Occupational Substance Exposure Assessment
<p>This is a corpus consisting of selected sections (i.e., <em>Abstract, Methods</em> and <em>Results</em>) of scientific research articles concerning occupational exposures to two different types of substance, i.e., diesel exhaust (51 articles) and respirable crystalline silica (RCS) (50 articles). The article sections have been annotated by experts in the field with 6 categories of named entities (NEs) relevant to the assessment of occupational substance exposures, particularly in the context of Job Exposure Matrices (JEMs).</p> <p>The corpus is available in two different formats, <a href="https://brat.nlplab.org/standoff.html">brat standoff format</a> and JSON. </p> <p>The corpus and associated NER models are described in more detail in the following atricle, which should be cited if you use the corpus: </p> <p>Thompson, P., Ananiadou, S., Basinas I., Brinchmann, B. C., Cramer, C., Galea, K. S., Ge, C., Georgiadis, P., Kirkeleit, J., Kuijpers, E., Nguyen, N., Nuñez, R., Schlünssen, V., Stokholm, Z. A., Taher, E. A., Tinnerberg, H., Van Tongeren, M. and Xie, Q. (2024).<a href="https://doi.org/10.1371/journal.pone.0307844"> </a><a href="https://doi.org/10.1371/journal.pone.0307844">Supporting the working life exposome: annotating occupational exposure for enhanced literature search</a>. PLoS ONE 19(8): e0307844</p>
FloraNER: a Named Entity Recognition Dataset for Botanical French Text
<p>FloraNER is a Named-Entity Recognition (NER) dataset for botanical french literature. The dataset covers plant species names and plant morphological terms for both plant organs/characteristics and their descriptors. the descriptors are annotated both in a coarse-grained manner as the named entity type "DESCRIPTOR" and in a fine-grained manner, categorized into the following named-entity types: Form, Measure, Surface, Color, Position, Disposition, Structure, and Development. FloraNER is distantly annotated using a specialized botanical corpus. Consequently, it's important to note that not all named entities within the text are captured and annotated for the coarse-grained and fine-grained datasets.</p>
Named-Entity Recognition for Modern Tibetan Newspapers: Tagset, Guidelines and Training Data
<p>This dataset, tagset and guidelines were the output of a six-month incubator project on the feasibility of developing Named-Entity Recognition (NER) for modern Tibetan, primarily for use with contemporary Tibetan-language newspapers and media published inside the PRC. The project was carried out by the Mongolian and Inner Asian Studies Unit at Cambridge University’s Department of Social Anthropology. It was funded by an incubator grant from Cambridge Language Sciences. The project title was “Named-Entity Recognition in Tibetan and Mongolian Newspapers.” The Project PI was Dr Hildegard Diemberger (Cambridge), the Coordinator and Lead Author was Dr Robert Barnett (SOAS), and Senior Advisers were Dr Nathan Hill (SOAS), Dr Marieke Meelen (Cambridge), and Dr Thomas White (Cambridge). <br> <br> Although some forms of NER and other NLP procedures have been developed within China for modern Tibetan (see Liu, Nuo <em>et al</em>, 2011), the data underlying those initiatives have not been made publicly available and their findings cannot be tested or reproduced. Significant work on developing NLP for Tibetan has been carried out outside China, but has focused largely on classical Tibetan and religious texts (see Hill & Garrett, Edward, 2017). </p> <p>The Cambridge incubator project therefore produced a tagset, guidelines and training data for developing NER for modern Tibetan, with a focus on historical and political analysis of contemporary newspapers, media and other public documents in Tibetan. We compiled 3.11m syllables of data in Tibetan extracted from articles downloaded from Chinese-language news aggregator sites within China, primarily tibet.cpc.people.com.cn and tibet.people.com.cn. From this data, we selected texts containing 280,000 syllables in Tibetan, grouped in 26,000 utterances/sentences (available on request). Using Lighttag, an online annotation site, we developed a tagset for NER consisting of 17 tags (and one for wrong segmentation if using segmented data). We annotated approximately 186,000 syllables, leading to 9,884 annotations. Of these, after discounting flawed data, we produced training data containing c.6,700 annotations. We carried out the secondary, manual review offline (for our method of converting Lighttag data for offline review, see the attached report “Using Spreadsheets to Review Annotations Offline.pdf”), and found an error rate of 3.6%. The final total of reviewed annotations was 6,624. </p> <p>The dataset, tagset, guidelines and reports were developed and documented by Robert Barnett, with assistance from Tsering Samdrup, Dr Hill and Dr Meelen. Primary annotation was by Tsering Samdrup, assisted by Dr Barnett.<br> <br> The datasets published here include: </p> <ol> <li>The <strong>tagseet guidelines and annotation manual</strong>, including the 17-tag tagset, guidelines, and recommendations ("NER for Modern Tibetan-tagset and guidelines.pdf").</li> <li>The <strong>tagged training data </strong>in .csv format ("Tibetan NER Training Data-tagged, reviewed wth context-v10-UTF-8.csv") and .xls format ("Tibetan NER Training Data-tagged with context-v10-UTF-8.xlsx"). This includes 6,624 reveiwed annotations, arranged according to the Tibetan alphabet together with the tags and context (utterance) for each annotation.</li> <li>The <strong>raw annotation results </strong>downloaded from Lighttag as .json files ("Raw Training Data for NER in Modern Tibetan -Jobs2-11-JSON.zip") and as .xls files ("Training Data for NER in Modern Tibetan -Jobs2-11-XLS.zip"). These include 10 "tasks" or datasets of articles scraped from Tibetan-language websites within Tibet. </li> <li>A <strong>guide to preparing Lighttag annotation results for manual review offline </strong>(“Using Spreadsheets to Review Annotations Offline.pdf”).</li> </ol> <p>The project's findings regarding the status of NER and NLP for vertical Mongolian are available at DOI: 10.5281/zenodo.5103499.</p>
French indirectly named entities
<p>A set of indirectly entities in French. (indirectly named entities are named entities described by a periphrasis)</p>
Line-level Named Entity Recognition annotation for the George Washington and IAM datasets
<p>Line-level Named Entity annotation for the George Washington and IAM datasets. The word-level annotations from Oliver Tüselmann [3] were extended to line-level to enable experimentation with line-level coupled HTR+NER models. We also publish the line-level partition files that result from the partition proposed by [3].</p>
Benchmark for the evaluation of named entity recognition over ancient documents
<p>The dataset consists of a multilingual noisy corpora for named entity recognition (NER).<br> The noisy versions are simulated from the CoNLL-02 (Spanish and Dutch) and CoNLL-03 (English) NER corpora.<br> The original collections are re-OCRed and four types of noises at two different levels are added in order to simulate various OCR output.</p> <p>More precisely, we first extracted raw texts and converted them into images. These images have been contaminated by adding some common noises when using a scanner. We further extract OCRed data using tesseract open source<br> OCR engine v-3.04.01. Consequently to the image noise insertions, OCRed data contains degradations. Original and noisy texts are finally aligned.</p> <p>This archive contains three folders (one per language). The folders contain the degraded images, the noisy texts extracted by the OCR and their aligned version with clean data.</p> <p>These are the supplementary materials for the TPDL 2020 paper <a href="https://zenodo.org/record/4734376#.YJKAcKE6-Uk">Assessing and minimizing the impact of OCR quality on named entity recognition</a>. If you end up using whole or parts of this resource,<br> please cite this paper:</p> <pre><code>@InProceedings{10.1007/978-3-030-54956-5_7, author="Hamdi, Ahmed and Jean-Caurant, Axel and Sid{\`e}re, Nicolas and Coustaty, Micka{\"e}l and Doucet, Antoine", editor="Hall, Mark and Mer{\v{c}}un, Tanja and Risse, Thomas and Duchateau, Fabien", title="Assessing and Minimizing the Impact of OCR Quality on Named Entity Recognition", booktitle="Digital Libraries for Open Knowledge", year="2020", publisher="Springer International Publishing", address="Cham", pages="87--101", isbn="978-3-030-54956-5" }</code></pre> <p><strong>Acknowledgments</strong><br> This work has been supported by the European Union's Horizon 2020 research and innovation programme under grant 770299 [NewsEye](https://www.newseye.eu/).</p>
GeoEDdA: A Gold Standard Dataset for Named Entity Recognition and Span Categorization Annotations of Diderot & d'Alembert's Encyclopédie
<p>This repository contains a gold standard dataset for named entity recognition and span categorization annotations from Diderot & d’Alembert’s Encyclopédie entries.</p> <p>The dataset is available in the following formats:</p> <ul> <li>JSONL format provided by <a href="https://prodi.gy/" rel="nofollow">Prodigy</a></li> <li>binary spaCy format (ready to use with the spaCy train pipeline)</li> </ul> <p>The Gold Standard dataset is composed of 2,200 paragraphs out of 2,001 Encyclopédie's entries randomly selected. All paragraphs were written in 19th-century French.</p> <p>The spans/entities were labeled by the project team along with using pre-labelling with early machine learning models to speed up the labelling process. A train/val/test split was used. Validation and test sets are composed of 200 paragraphs each: 100 classified under 'Géographie' and 100 from another knowledge domain. The datasets have the following breakdown of tokens and spans/entities.</p> <h2>Tagset</h2> <ul> <li><strong>NC-Spatial</strong>: a common noun that identifies a spatial entity (nominal spatial entity) including natural features, e.g. <code>ville</code>, <code>la rivière</code>, <code>royaume</code>.</li> <li><strong>NP-Spatial</strong>: a proper noun identifying the name of a place (spatial named entities), e.g. <code>France</code>, <code>Paris</code>, <code>la Chine</code>.</li> <li><strong>ENE-Spatial</strong>: nested spatial entity , e.g. <code>ville de France</code> , <code>royaume de Naples</code>, <code>la mer Baltique</code>.</li> <li><strong>Relation</strong>: spatial relation, e.g. <code>dans</code>, <code>sur</code>, <code>à 10 lieues de</code>.</li> <li><strong>Latlong</strong>: geographic coordinates, e.g. <code>Long. 19. 49. lat. 43. 55. 44.</code></li> <li><strong>NC-Person</strong>: a common noun that identifies a person (nominal spatial entity), e.g. <code>roi</code>, <code>l'empereur</code>, <code>les auteurs</code>.</li> <li><strong>NP-Person</strong>: a proper noun identifying the name of a person (person named entities), e.g. <code>Louis XIV</code>, <code>Pline</code>.</li> <li><strong>ENE-Person</strong>: nested people entity, e.g. <code>le czar Pierre</code>, <code>roi de Macédoine</code>.</li> <li><strong>NP-Misc</strong>: a proper noun identifying entities not classified as spatial or person, e.g. <code>l'Eglise</code>, <code>1702</code>, <code>Pélasgique</code></li> <li><strong>ENE-Misc</strong>: nested named entity not classified as spatial or person, e.g. <code>l'ordre de S. Jacques</code>, <code>la déclaration du 21 Mars 1671</code>.</li> <li><strong>Head</strong>: entry name</li> <li><strong>Domain-Mark</strong>: words indicating the knowledge domain (usually after the head and between parenthesis), e.g. <code>Géographie</code>, <code>Geog.</code>, <code>en Anatomie</code>.</li> </ul> <h2>HuggingFace</h2> <p>The GeoEDdA dataset is available on the HuggingFace Hub: <a href="https://huggingface.co/datasets/GEODE/GeoEDdA">https://huggingface.co/datasets/GEODE/GeoEDdA</a></p> <h2>spaCy Custom Spancat trained on Diderot & d’Alembert’s Encyclopédie entries</h2> <p>This dataset was used to train and evaluate a custom spancat model for French using <a href="https://spacy.io/" rel="nofollow">spaCy</a>. The model is available on HuggingFace's model hub: <a href="https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda" rel="nofollow">https://huggingface.co/GEODE/fr_spacy_custom_spancat_edda</a>.</p> <h2>Acknowledgement</h2> <p>The authors are grateful to the <a href="https://aslan.universite-lyon.fr/" rel="nofollow">ASLAN project</a> (ANR-10-LABX-0081) of the Université de Lyon, for its financial support within the French program "Investments for the Future" operated by the National Research Agency (ANR). Data courtesy the <a href="https://artfl-project.uchicago.edu/" rel="nofollow">ARTFL Encyclopédie Project</a>, University of Chicago.</p>
Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences in Biomedical Literature and Its Integration with PubChemRDF
<p>This Zenodo record contains the co-occurrence RDF data generated in the work described in the paper “<strong>A resource description framework (RDF) model of named entity co-occurrences in biomedical literature and its integration with PubChemRDF</strong>” by Li et al., published in the Journal of Cheminformatics (<a href="https://doi.org/10.1186/s13321-025-01017-0" target="_blank" rel="noopener">https://doi.org/10.1186/s13321-025-01017-0</a>). It also contains the SPARQL query examples, the RDF schema in SHACL and ShEx, and the validation scripts.</p> <p>All content in this Zenodo record is for archival purposes. The latest version of the co-occurrence RDF data and other PubChemRDF data can be accessed via the PubChem FTP site (<a href="https://ftp.ncbi.nlm.nih.gov/pubchem/RDF/" target="_blank" rel="noopener">https://ftp.ncbi.nlm.nih.gov/pubchem/RDF/</a>). The up-to-date RDF schema in various formats is available on the PubChemRDF Schema page (<a href="https://pubchem.ncbi.nlm.nih.gov/docs/rdf-schema" target="_blank" rel="noopener">https://pubchem.ncbi.nlm.nih.gov/docs/rdf-schema</a>). A set of SPARQL query examples can be found on the PubChemRDF use case pages (<a href="https://pubchem.ncbi.nlm.nih.gov/docs/rdf-use-cases" target="_blank" rel="noopener">https://pubchem.ncbi.nlm.nih.gov/docs/rdf-use-cases</a>).</p>
CLEF-HIPE-2020 Shared Task Named Entity Datasets
<p>CLEF-HIPE-2020 (Identifying Historical People, Places and other Entities) is a <strong>evaluation campaign on named entity processing</strong> on <strong>historical newspapers</strong> in <strong>French</strong>, <strong>German</strong> and <strong>English</strong>, which was organized in the context of the <a href="http://impresso-project.ch"><em>impresso</em></a> project and run as a <a href="https://clef2020.clef-initiative.eu/">CLEF 2020</a> Evaluation Lab. Data consists of manually annotated historical newspapers in French, German and English.<br> <br> For more information, please refer to:</p> <ul> <li>the CLEF-HIPE-2020 <a href="https://impresso.github.io/CLEF-HIPE-2020/">website</a>;</li> <li>the <a href="https://github.com/impresso/CLEF-HIPE-2020-eval">CLEF-HIPE-2020-eval repository</a>, for the necessary material to replicate the results of the shared task;</li> <li>the <a href="https://zenodo.org/record/3539085">CLEF-HIPE-2020 poster</a> presented at CLEF 2019 in Lugano, Switzerland;</li> <li>the CLEF-HIPE-2020 <a href="https://zenodo.org/record/3677171">participation guidelines</a> (v1.1);</li> <li>the <em>impresso</em> <a href="https://doi.org/10.5281/zenodo.3585749">Named Entity Annotation Guidelines</a> (v2.2.0);</li> <li>the <a href="https://infoscience.epfl.ch/record/281054">CLEF-HIPE-2020 Extended Overview</a> paper (bibtex below);</li> <li>the participant team <a href="http://ceur-ws.org/Vol-2696/">CEUR working note papers</a>;</li> <li>the workshop presentation <a href="https://www.youtube.com/playlist?list=PLB45F159nVx-3bee7G_1jdTfUAtsLD0FU">video records</a>;</li> </ul> <p>A second edition of HIPE is organised in 2022: <a href="https://hipe-eval.github.io/HIPE-2022/">https://hipe-eval.github.io/HIPE-2022/ </a></p> <p>Please cite this paper if you are using the datasets or find the shared task results relevant to your research:</p> <pre><code>@inproceedings{ehrmann_extended_2020, title = {Extended {Overview} of {CLEF HIPE} 2020: {Named Entity Processing} on {Historical Newspapers}}, booktitle = {{CLEF 2020 Working Notes}. {Working Notes} of {CLEF} 2020 - {Conference} and {Labs} of the {Evaluation Forum}}, author = {Ehrmann, Maud and Romanello, Matteo and Fl{\"u}ckiger, Alex and Clematide, Simon}, editor = {Cappellato, Linda and Eickhoff, Carsten and Ferro, Nicola and N{\'e}v{\'e}ol, Aur{\'e}lie}, year = {2020}, volume = {2696}, pages = {38}, publisher = {{CEUR-WS}}, address = {{Thessaloniki, Greece}}, doi = {10.5281/zenodo.4117566}, url = {https://infoscience.epfl.ch/record/281054}, }</code></pre> <p> </p>
Dataset for Named Entity Recognition and Entity Linking from Greek Wikipedia Events
<p>An automated benchmark dataset for (Named Entity Recognition) NER and (Named Entity Linking) NEL tools, based on Greek Wikipedia events pages.</p> <p>Note: This data includes data from the following sources:<br> - Wikipedia el.wikipedia.org</p> <p><strong>Description</strong></p> <p>The dataset is provided in the form of three JSON-formatted subsets i.e., train, validation and test in an analogy of 70-20-10. The current version of the dataset contains 18,617 events annotated with 40,798 entity mentions and 36,189 links to elWikipedia (and wikidata ids). The dataset contains annotations belonging to 8 entity types: person, organization, location, gpe, event, facility, product and work of art.</p> <table> <caption>Overall dataset statistics</caption> <thead> <tr> <th scope="col"> </th> <th scope="col">Docs</th> <th scope="col">Tokens</th> <th scope="col">Sentences</th> <th scope="col">Surface Mentions</th> <th scope="col">Valid Links</th> <th scope="col">Red Links</th> </tr> </thead> <tbody> <tr> <td><strong>Train</strong></td> <td>13,031</td> <td>332,077</td> <td>16,927</td> <td>28,593</td> <td>25,365</td> <td>3,228</td> </tr> <tr> <td><strong>Validation</strong></td> <td>3,722</td> <td>94,746</td> <td>4,844</td> <td>8,168</td> <td>7,240</td> <td>928</td> </tr> <tr> <td><strong>Test</strong></td> <td>1,862</td> <td>47,450</td> <td>2,427</td> <td>4,037</td> <td>3,584</td> <td>453</td> </tr> <tr> <td><strong>Total</strong></td> <td>18,617</td> <td>474,361</td> <td>24,200</td> <td>40,798</td> <td>36,189</td> <td>4,609</td> </tr> </tbody> </table> <p><strong>Example</strong></p> <p>A record example is given below.</p> <p>{</p> <p>"json_file": "February 2012_39_0 events",<br> "text": "Sudan and South Sudan sign non-aggression pact.",<br> "ground_truth_mentions": [<br> {"start": 0, "end": 4, "surface_mention": "Sudan", "mention_type": "GPE"},<br> {"start": 10, "end": 20, "surface_mention": "South Sudan", "mention_type": "GPE"}<br> ],<br> "ground_truth_links": [<br> {"enwiki": "Sudan","wikidata": "Q1049"},<br> {"enwiki": "South_Sudan", "wikidata": "Q958"}<br> ]<br> }</p> <p><strong>Code</strong></p> <p><a href="https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark">https://gitlab.isl.ics.forth.gr/debatelab/elwiki_events_benchmark</a></p> <p><strong>Acknowledgments</strong></p> <p>This work has received funding from the Hellenic Foundation for Research and Innovation (HFRI) and the General Secretariat for Research and Technology (GSRT), under grant agreement No 4195.</p>
LivingNER corpus: Named entity recognition, normalization & classification of species, pathogens and food
<p><strong>LivingNER Gold Standard corpus (includes training, validation, test and background sets + MULTILINGUAL RESOURCES</strong>)</p><p> </p><p><strong>Please cite if you use this dataset:</strong></p><p>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization & classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</p><p>@article{amiranda2022nlp, title={Mention detection, normalization \& classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources}, author={Miranda-Escalada, Antonio and Farr{\'e}-Maduell, Eul{`a}lia and Lima-L{\'o}pez, Salvador and Estrada, Darryl and Gasc{\'o}, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, year={2022} }</p><p> </p><p><i><strong>1. Introduction</strong></i></p><p>The LivingNER Gold Standard corpus is a collection of<strong> 2000 clinical case reports</strong> covering a <strong>broad range of medical specialities</strong>, i.e. infectious diseases (including Covid-19 cases), cardiology, neurology, oncology, dentistry, pediatrics, endocrinology, primary care, allergology, radiology, psychiatry, ophthalmology, urology, internal medicine, emergency and intensive care medicine, tropical medicine, and dermatology <strong>annotated with species</strong> [SPECIES] (including <strong>living organisms</strong> and <strong>microorganisms</strong>) and <strong>infectious diseases</strong> [ENFERMEDAD] mentions. Species mentions include many <strong>pathogens</strong> and infectious agents, but also <strong>food</strong>, allergens, <strong>pets</strong> or other species, taxonomic groups and organisms of clinical relevance. </p><p>The LivingNER corpus has also annotations of mentions of <strong>humans</strong> (tag HUMAN), including the patients itself, <strong>family members</strong>, healhcare professionals or other persons mentioned in the case reports. Thus it can be useful to extract family history information of patients or information about the social and healthcare personal environment and interactions.</p><p>All mentions have been exhaustively manually mapped by experts to their corresponding <a href="https://www.ncbi.nlm.nih.gov/taxonomy"><strong>NCBI Taxonomy</strong></a> identifiers. </p><p>It was used for the <a href="https://temu.bsc.es/livingner/">LivingNER</a> Shared Task on pathogens and living beings detection and normalization in Spanish medical documents, which was celebrated as part of IberLEF 2022.</p><p> </p><p><i><strong>2. Training, validation, test and background sets</strong></i></p><p>The training set is composed of 1000 clinical case reports. The validation set includes 500 clinical case reports with the same characteristics and the test set includes 485. The background set is a collection of around 13k unannotated case reports that were originally added to prevent manual annotations in the test set during the competition and to create a Silver Standard.</p><p><i><strong>2.1 Annotations format</strong></i></p><p>Annotations and text files are distributed separately. The texts are in plain text (.txt in UTF-8) format, while the annotations are are distributed in a tab-separated file (.tsv) file with one row per annotation:</p><p>- For <strong>subtask 1 (LivingNER-Species NER track)</strong>, the .tsv file has the following columns:</p><ul><li>filename: document name</li><li>mark: identifier mention mark</li><li>label: mention type (SPECIES or HUMAN)</li><li>off0: starting position of the mention in the document</li><li>off1: ending position of the mention in the document</li><li>span: textual span</li></ul><p> - For <strong>subtask 2 (LivingNER-Species Norm track)</strong>, the .tsv file has the same columns as the previous one, plus:</p><ul><li>isH: whether the span is narrower than the NCBITax assigned code</li><li>isN: whether the mention corresponds to a nosocomial infection</li><li>iscomplex: whether the span has assigned a combination of NCBITax codes</li><li>NCBITax: mention code in the NCBI Taxonomy</li></ul><p>- For <strong>subtask 3 (LivingNER-Clinical IMPACT track)</strong>, the .tsv file has the following columns:</p><ul><li>filename</li><li>isPet (Yes/No)</li><li>PetIDs (NCBITaxonomy codes of pet & farm animals present in document)</li><li>isAnimalInjury (Yes/No)</li><li>AnimalInjuryIDs (NCBITaxonomy codes of animals causing injuries present in document)</li><li>IsFood (Yes/No)</li><li>FoodIDs (NCBITaxonomy codes of food mentions present in document)</li><li>isNosocomial (Yes/No)</li><li>NosocomialIDs (NCBITaxonomy codes of nosocomial species mentions present in document)</li></ul><p><i><strong>2.2 Important notes about subtask 3 (LivingNER-Clinical IMPACT track):</strong></i></p><ul><li><strong>Less clinical case reports</strong>. Subtask 3 (LivingNER-Clinical IMPACT track) contains half of the clinical case reports (500 in the training partition, 250 in the validation partition). The list of valid clinical case reports for task 3 is included in the data (train_files_task3.txt and validation_files_task3.txt)</li><li><strong>Enriched dataset.</strong> The GS format is the one described above (a TSV with one line per clinical case report). However, we believe participants may find useful and <strong>enriched dataset. </strong>Then, we provide an additional dataset, with the mentions of the NER track classified in the 4 Clinical impact categories (food, pet&farm animals, animals causing injuries and nosocomial). It is a TSV file with one row per annotation, and with the following columns: filename, mark, label, off0, off1, span, isPet, isAnimalInjury, isFood, isNosocomial, isH, iscomplex, code</li></ul><p> </p><p><i><strong>3. Multilingual resources</strong></i></p><p>We have generated the annotated training and validation sets in <strong>7 languages</strong>:</p><ul><li><i><strong>English</strong></i></li><li><i><strong>Portuguese</strong></i></li><li><i><strong>Catalan</strong></i></li><li><i><strong>Galician</strong></i></li><li><i><strong>Italian</strong></i></li><li><i><strong>French</strong></i></li><li><i><strong>Romanian</strong></i></li></ul><p> </p><p>The process was:</p><ol><li>The text files were translated with a neural machine translation system.</li><li>The annotations were translated with the same neural machine translation system.</li><li>The translated annotations were transferred to the translated text files using an annotation transfer technology.</li></ol><p>The text files are stored in the multilingual_resources/<strong>training-text-files</strong> and multilingual_resources/<strong>validation-text-files </strong>subfolders.</p><p>The annotated TSV files are stored in the multilingual_resources/<strong>annotation_transfer </strong>subfolder.</p><p>For the sake of comparison, we incorporate as well the annotations that resulted from the <a href="https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-11-85">LINNAEUS tool</a> in the multilingual_resources/<strong>linneaus</strong> subfolder.</p><p>If you want to visualize the multilingual resources, check out this Brat server: <a href="https://temu.bsc.es/mLivingNER/#/translations/">https://temu.bsc.es/mLivingNER/#/translations/</a></p><p>For instance, you can see the parallel annotations in <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/en/annotation_transfer/train/casos_clinicos_cardiologia34?diff=/translations/fr/annotation_transfer/train/">English vs in French</a>, or <a href="https://temu.bsc.es/mLivingNER/diff.xhtml#/translations/cat/annotation_transfer/train/casos_clinicos_cardiologia35?diff=/gold-standard/train/">in Spanish (the gold standard) vs in Catalan.</a></p><p> </p><p><strong>Resources</strong></p><ul><li><a href="https://temu.bsc.es/livingner/"><strong>Task Web</strong></a></li><li><strong>Citation: </strong>A. Miranda-Escalada, E. Farré-Maduell, S. Lima-López, D. Estrada, L. Gascó, M. Krallinger, Mention detection, normalization & classification of species, pathogens, humans and food in clinical documents: Overview of LivingNER shared task and resources, <i>Procesamiento del Lenguaje Natural</i> (2022)</li><li><a href="https://doi.org/10.5281/zenodo.6385162"><strong>Annotation guidelines</strong></a></li><li><a href="https://github.com/tonifuc3m/livingner-evaluation-library"><strong>Evaluation library</strong></a></li><li><a href="https://doi.org/10.5281/zenodo.6390506">LivingNER terminology</a></li><li><a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6444"><strong>Overview paper</strong></a></li><li><a href="https://ceur-ws.org/Vol-3202/"><strong>Proceedings participant papers</strong></a></li><li><a href="https://www.youtube.com/watch?v=8VcZw8ywyJY&list=PL5uSCzf1azhA_gMLC3DBZe6NvmMJiggTg"><strong>Youtube videos</strong></a></li><li><a href="https://www.slideshare.net/MartinKrallinger/mention-detection-normalization-classification-of-species-pathogens-humans-and-food-in-clinical-documents-overview-of-the-livingner-shared-task-and-resources-talk-at-iberlef-sepln-2022"><strong>LivingNER overview talk sides at IberLEF/SEPLN</strong></a></li></ul><p> </p><p><strong>License</strong></p><p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p><p><strong>Contact</strong></p><p>If you have any questions or suggestions, please contact us at:</p><p><br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p><p><strong>Additional resources and corpora</strong></p><p>If you are interested in SympTEMIST, you might want to check out these corpora and resources:</p><ul><li><a href="https://zenodo.org/records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8413866">SympTEMIST</a> (Corpus of symptoms, sign and findings mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization, different document collection)</li><li><a href="https://zenodo.org/records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization, different document collection)</li><li><a href="https://zenodo.org/records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI), differentdocument collection)</li><li><a href="https://zenodo.org/records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization, different document collection)</li><li><a href="https://zenodo.org/records/3837305">CodiESp</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries), different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags, different document collection, some overlapping documents)</li><li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, different document collection, some overlapping documents)</li></ul>
Berlin State Library (2023). Named Entity Disambiguation German Database for the Named Entity Linking System of the Berlin State Library (SBB)
<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>
Berlin State Library (2023). Named Entity Disambiguation French Database for the Named Entity Linking System of the Berlin State Library (SBB)
<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>
Berlin State Library (2023). Named Entity Disambiguation English Database for the Named Entity Linking System of the Berlin State Library (SBB)
<p>This database is a part of a BERT-based entity recognition and three-stage entity linking (EL) system. Its components consist of three models as well as three related databases, one of which is published here.</p>
CORD-19 Named Entities Knowledge Graph (CORD19-NEKG)
<p>CORD-19 Named Entities Knowledge Graph (CORD19-NEKG) is an RDF dataset describing named entities identified in the scholarly articles of the <a href="https://www.semanticscholar.org/cord19">COVID-19 Open Research Dataset</a> (CORD-19), a resource of over 47,000 articles about COVID-19 and the coronavirus family of viruses.</p> <p>Homepage: https://github.com/Wimmics/cord19-nekg</p> <p>License: see LICENCE file in the dataset.</p>
Historical German Children's Playbooks - 6 Digitized Books with Images, OCR-Fulltext, and Named Entity Recognition
<p>The dataset consists of 6 digitized books with 1750 images and OCR-fulltext.</p> <p>Additionally, named entity recognition has been carried out on basis of flair's de-ner model, see https://github.com/flairNLP for details.</p>
CrowdTruth/Crowdsourcing-NamedEntities-GoldStandard: Data release for crowdsourcing named entity gold standards
<p>This repository contains the experimental results of identifying and typing named entities in English Wikipedia sentences by using a hybrid Multi-NER - crowd-enhanced approach.</p>
AI4PROFHEALTH - Automatic Silver Gazetteer for Named Entity Recognition and Normalization
<p>This dataset comprises a professions gazetteer generated with automatically extracted terminology from the Mesinesp2 corpus, a manually annotated corpus in which domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts, as well as clinical case reports. </p> <p>A silver gazetteer for mention classification and normalization is created combining the predictions of automatic Named Entity Recognition models and normalization using Entity Linking to three controlled vocabularies SNOMED CT, NCBI and ESCO. The sources are 265,025 different documents, where 249,538 correspond to <a href="https://zenodo.org/records/4707104">MESINESP2 Corpora</a> and 15,487 to clinical cases from open clinical journals. From them, 5,682,000 mentions are extracted and 4,909,966 (86.42%) are normalized to any of the ontologies: SNOMED CT (4,909,966) for diseases, symptoms, drugs, locations, occupations, procedures and species; ESCO (215,140) for occupations; and NCBI (1,469,256) for species.</p> <p>The repository contains a .tsv file with the following columns:</p> <ul> <li><strong><code>filenameid</code></strong>: A unique identifier combining the file name and mention span within the text. This ensures each extracted mention is uniquely traceable. Example: biblio-1000005#239#256 refers to a mention spanning characters 239–256 in the file with the name biblio-1000005.</li> <li> <p><strong><code>span</code></strong>: The specific text span (mention) extracted from the document, representing a term or phrase identified in the dataset. Example: centro oncológico.</p> </li> <li> <p><strong><code>source</code></strong>: The origin of the document, indicating the corpus from which the mention was extracted. Possible values: mesinesp2, clinical_cases.</p> </li> <li> <p><strong><code>filename</code></strong>: The name of the file from which the mention was extracted. Example: biblio-1000005.</p> </li> <li> <p><strong><code>mention_class</code></strong>: Categories or semantic tags assigned to the mention, describing its type or context in the text. Example: ['ENFERMEDAD', 'SINTOMA'].</p> </li> <li> <p><strong><code>codes_esco</code></strong>: The normalized ontology codes from the European Skills, Competences, Qualifications, and Occupations (ESCO) vocabulary for the identified mention (if applicable). This field may be empty if no ESCO mapping exists. Example: 30629002.</p> </li> <li> <p><strong><code>terms_esco</code></strong>: The human-readable terms from the ESCO ontology corresponding to the <code>codes_esco</code>. Example: ['responsable de recursos', 'director de recursos', 'directora de recursos'].</p> </li> <li> <p><strong><code>codes_ncbi</code></strong>: The normalized ontology codes from the NCBI Taxonomy vocabulary for species (if applicable). This field may be empty if no NCBI mapping exists.</p> </li> <li> <p><strong><code>terms_ncbi</code></strong>: The human-readable terms from the NCBI Taxonomy vocabulary corresponding to the <code>codes_ncbi</code>. Example: ['Lacandoniaceae', 'Pandanaceae R.Br., 1810', 'Pandanaceae', 'Familia'].</p> </li> <li> <p><strong><code>codes_sct</code></strong>: The normalized ontology codes from SNOMED CT (Systematized Nomenclature of Medicine - Clinical Terms) vocabulary for diseases, symptoms, drugs, locations, occupations, procedures, and species (if applicable). Example: 22232009.</p> </li> <li> <p><strong><code>terms_sct</code></strong>: The human-readable terms from the SNOMED CT ontology corresponding to the <code>codes_sct</code>. Example: ['adjudicador de regulaciones del seguro nacional'].</p> </li> <li> <p><strong><code>sct_sem_tag</code></strong>: The semantic category tag assigned by SNOMED CT to describe the general classification of the mention. Example: environment.</p> </li> </ul> <p> </p> <p><strong>Suggestion</strong>: If you load the dataset using python, it is recommended to read the columns containing lists as follows</p> <div> <blockquote> <div>import ast</div> <div>df["mention_class"] = df["mention_class"].apply(lambda x: ast.literal_eval(x) if isinstance(x, str) else x)</div> </blockquote> </div> <p> </p> <p><strong>License</strong></p> <p>This dataset is licensed under <strong>Creative Commons Attribution 4.0 International (CC BY 4.0)</strong>. This means you are free to:</p> <ul> <li>Share: Copy and redistribute the material in any medium or format.</li> <li>Adapt: Remix, transform, and build upon the material for any purpose, even commercially.</li> </ul> <p><strong>Attribution Requirement</strong>: Please credit the dataset creators appropriately, provide a link to the license, and indicate if changes were made.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p> <p><strong>Additional resources and corpora</strong></p> <p>If you are interested, you might want to check out these corpora and resources:</p> <ul> <li><a href="https://zenodo.org/records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts, different document collection)</li> <li><a href="10.5281/zenodo.5070540" target="_blank" rel="noopener">MEDDOPROF corpus </a></li> <li> <p><a href="https://zenodo.org/record/4722741">Codes Reference List</a> (for MEDDOPROF-NORM)</p> </li> <li> <p><a href="https://zenodo.org/record/4720833">Annotation Guidelines</a></p> </li> <li> <p><a href="https://doi.org/10.5281/zenodo.4524658">Occupations Gazetteer</a></p> </li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.