Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

42

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

42 results for “Named Entities”

Learn how ShareScore rates datasets ↗
zenodo36/100

Risultati dell'analisi di frequenza sul dataset di named entities italiane

<p>Risultati dell&#39;analisi di frequenza sulle named entities estratte dall&#39;elenco delle vie&nbsp;italiane, divise per territorio e per fasce di popolazione, elencate in ordine di frequenza discendente. I file che fanno riferimento alla divisione per territorio sono &quot;nord-est&quot;, &quot;nord-ovest&quot;, &quot;centro&quot;, &quot;sud&quot;, &quot;isole&quot;. Quelli che fanno riferimento alla divisione per fasce di popolazione iniziano con una lettera maiuscola seguita da un punto (vanno dalla &quot;A.&quot; alla &quot;L.&quot;). E&#39; presente anche il file &quot;italia.txt&quot; che riporta tutti i risultati italiani (tutti gli altri file riportano solo le prime 100 named entities).</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Multilingual named entity recognition for medieval charters. Datasets and models

<p>Annotated dataset for training named entities recognition models for medieval charters in Latin, French and Spanish.</p> <p>&nbsp;</p> <p>The original raw texts for all charters were collected from four charters collections</p> <p>- HOME-ALCAR corpus : <a href="https://zenodo.org/record/5600884">https://zenodo.org/record/5600884</a></p> <p>- CBMA : <a href="https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=&amp;ved=2ahUKEwisvNOa3qP3AhULyIUKHenpBoAQFnoECA0QAQ&amp;url=http%3A%2F%2Fwww.cbma-project.eu%2F&amp;usg=AOvVaw0blsASXzOKSNz_EkixwfJT">http://www.cbma-project.eu</a></p> <p>- Diplomata Belgica : <a href="https://www.diplomata-belgica.be">https://www.diplomata-belgica.be</a></p> <p>- CODEA corpus :<a href="http://https://corpuscodea.es/"> https://corpuscodea.es/</a></p> <p>&nbsp;</p> <p>We include (i) the annotated training datasets, (ii) the contextual and static embeddings trained on medieval multilingual texts and (iii) the named entity recognition models trained using two architectures: Bi-LSTM-CRF + stacked embeddings and fine-tuning on Bert-based models (mBert and RoBERTa)</p> <p>Codes, datasets and notebooks&nbsp;used to train models can&nbsp;be consulted in our&nbsp;gitlab repository:&nbsp;<a href="https://gitlab.com/magistermilitum/ner_medieval_multilingual">https://gitlab.com/magistermilitum/ner_medieval_multilingual</a></p> <p>Our best RoBERTa model is also available in the HuggingFace library:&nbsp;<a href="https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner">https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner</a></p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

BiodivNERE: Gold Standard Corpora for Named Entity Recognition and Relation Extraction in Biodiversity Domain

<p>BiodivNER+RE are two gold-standard manually annotated corpora that are meant to be for Named Entity Recognition (NER) and Relation Extraction (RE) tasks based on abstracts and metadata files from the biodiversity domain</p> <p>Such corpora are designed for machine learning techniques,&nbsp;for example, NER as&nbsp;TokenClassification technique, and RE as&nbsp;SequenceClassification technique.</p>

opencc-by-4.0Apr 2022View details →
zenodo36/100

HIPE-2022 Shared Task Named Entity Datasets

<p>HIPE-2022 datasets used for the <a href="https://hipe-eval.github.io/HIPE-2022/">HIPE 2022 shared task</a> on <strong>named entity recognition and classification (NERC) and entity linking (EL) in multilingual historical documents</strong>.&nbsp;</p> <p>HIPE-2022 datasets are based on six primary datasets assembled and prepared for the shared task. Primary datasets are composed of historical newspapers and classic commentaries covering ca. 200 years, feature several languages and different entity tag sets and annotation schemes. They originate from several European cultural heritage projects, from HIPE organizers&rsquo; previous research project, and from the previous HIPE-2020 campaign. Some are already published, others are released for the first time for HIPE-2022.</p> <p>The HIPE-2022 shared task assembles and prepares these primary datasets in HIPE-2022 release(s), which correspond to a single package composed of neatly structured and homogeneously formatted files.</p> <p>Primary datasets undergo the following preparation steps:</p> <ul> <li>conversion to the HIPE format (with correction of data inconsistencies and metadata consolidation);</li> <li>rearrangement or composition of train and dev splits.</li> </ul> <p>Please also refer to:</p> <ul> <li>HIPE-2022 shared task website: <a href="https://hipe-eval.github.io/HIPE-2022/">https://hipe-eval.github.io/HIPE-2022/</a></li> <li>HIPE-2022 data repository: <a href="https://github.com/hipe-eval/HIPE-2022-data">https://github.com/hipe-eval/HIPE-2022-data</a></li> </ul> <p>Here is an overview of the primary datasets:</p> <table> <tbody> <tr> <td> <p><strong>Dataset alias</strong></p> </td> <td> <p><strong>Readme</strong></p> </td> <td> <p><strong>Document type</strong></p> </td> <td> <p><strong>Languages</strong></p> </td> <td> <p><strong>Suitable for</strong></p> </td> <td> <p><strong>Project</strong></p> </td> </tr> <tr> <td> <p>hipe2020</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-hipe2020.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de, fr, en</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://impresso.github.io/CLEF-HIPE-2020">CLEF-HIPE-2020</a></p> </td> </tr> <tr> <td> <p>newseye</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-newseye.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de, fi, fr, sv</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://www.newseye.eu/">NewsEye</a></p> </td> </tr> <tr> <td> <p>sonar</p> </td> <td> <p>link</p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de</p> </td> <td> <p>NERC-Coarse, EL</p> </td> <td> <p><a href="https://sonar.fh-potsdam.de/">SoNAR</a></p> </td> </tr> <tr> <td> <p>letemps</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-letemps.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>fr</p> </td> <td> <p>NERC-Coarse, NERC-Fine</p> </td> <td> <p>LeTemps</p> </td> </tr> <tr> <td> <p>topres19th</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-topres19th.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>en</p> </td> <td> <p>NERC-Coarse, EL</p> </td> <td> <p><a href="https://livingwithmachines.ac.uk/">Living with Machines</a></p> </td> </tr> <tr> <td> <p>ajmc</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-ajmc.md">link</a></p> </td> <td> <p>classical commentaries</p> </td> <td> <p>de, fr, en</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://mromanello.github.io/ajax-multi-commentary/">AjMC</a></p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>The HIPE-2022 team expresses her greatest appreciation to the partnering projects, namely <a href="https://mromanello.github.io/ajax-multi-commentary/">AJMC</a>, <em><a href="https://impresso-project.ch/">impresso</a>, </em><a href="https://impresso.github.io/CLEF-HIPE-2020/">HIPE-2020</a>, <a href="https://livingwithmachines.ac.uk/"><em>Living with Machines</em></a>, <a href="https://www.newseye.eu/"><em>NewsEye</em></a>, and <a href="https://sonar.fh-potsdam.de/"><em>SoNAR</em></a>, for contributing their NE-annotated datasets (and hiding a part thereof for the time of the evaluation campaign).</p>

opencc-by-nc-4.0Feb 2022View details →
zenodo36/100

Benchmark for the evaluation of Named Entity Linking over ancient documents

<p><strong>Benchmark for the evaluation of Named Entity Linking over ancient documents</strong><br> Elvys Linhares Pontes, Ahmed Hamdi, Nicolas Sidere, and Antoine Doucet<br> University of Avignon: elvys.linhares-pontes@univ-avignon.fr; University of La Rochelle: {elvys.linhares_pontes,ahmed.hamdi,nicolas.sidere,antoine.doucet}@univ-lr.fr</p> <p>These are the supplementary materials for the ICADL 2019 paper <strong><em>Impact of OCR Quality on Named Entity Linking</em></strong>. If you end up using whole or parts of this resource, please use the following citation:</p> <ul> <li>Linhares Pontes, E., Hamdi, A., Sidere, N., and Doucet, A. (2019). Impact of OCR Quality on Named Entity Linking. In Proceedings of 21st International Conference on Asia-Pacific Digital Libraries ICADL 2019, Kuala Lumpur, Malaysia.</li> </ul> <p>or alternatively use the following `bib`:</p> <pre><code>@inproceedings{linhares2019icadl, title="Impact of OCR Quality on Named Entity Linking.", author={Linhares Pontes, Elvys, and Hamdi, Ahmed, and Sidere, Nicolas, and Doucet, Antoine}, year={2019}, booktitle={Proceedings of 21st International Conference on Asia-Pacific Digital Libraries ICADL 2019} }</code></pre> <p><strong>Files</strong><br> This archive contains six folders -- one per dataset -- as well as this README. The folders contain the degraded images, the noisy texts extracted by the OCR and their aligned version with clean data. This work is licensed under a [Creative Commons Attribution-ShareAlike 4.0 International License](http://creativecommons.org/licenses/by-sa/4.0/).</p> <p><strong>Acknowledgments</strong><br> This work has been supported by the European Union&#39;s Horizon 2020 research and innovation programme under grant 770299 [NewsEye](https://www.newseye.eu/).</p>

opencc-by-4.0Oct 2019View details →
zenodo36/100

Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction

<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models.&nbsp;</p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p>&nbsp; &nbsp; English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br>&nbsp; &nbsp; German: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br>&nbsp; &nbsp; Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br>&nbsp; &nbsp; Greek: https://huggingface.co/HUMADEX/german_medical_ner<br>&nbsp; &nbsp; Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br>&nbsp; &nbsp; Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br>&nbsp; &nbsp; Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p>&nbsp;</p> <p><strong>Dataset Building&nbsp;</strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model&nbsp;</li> <li>Translation into the targeted language</li> <li>Word Alignment&nbsp;</li> <li>Data Augmentation&nbsp;</li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website:&nbsp;</span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Annotated Dataset for Named Entity Recognition and Relation Extraction in French Building Technical Specifications (BTS)

<p>This dataset contains 233 raw requirements extracted from French Building Technical Specifications (BTS), referred to as "<a href="https://www.aglo.ai/cctp/#:~:text=Le%20CCTP%20(Cahier%20des%20Clauses%20Techniques%20Particuli%C3%A8res)%20est%20un%20document,code%20du%20march%C3%A9%20public%20donc."><em>Cahier des Clauses Techniques Particuli&egrave;res (CCTP)</em></a>", specifically focused on carpentry ("<em>lot menuiserie</em>") in public French construction projects. The requirements have been collected from 72 CCTP documents, resulting in a total of 19,725 sentences and 651,948 words.</p> <p>The dataset has been annotated using <a title="Open-source text annotation tool" href="https://github.com/doccano/doccano">Doccano </a>for Named Entity Recognition (NER) and Relation Extraction (RE). The annotations involve identifying entities and the relationships between them within the domain of building requirements. This dataset is intended for research on Natural Language Processing (NLP) models for Requirements Engineering (RE) in the Architecture, Engineering, and Construction (AEC) sector. Potential applications include requirements extraction, compliance analysis, and knowledge management in construction.</p> <p>The dataset includes the following components:</p> <ol> <li><strong>CCTP Documents</strong>: The original CCTP files from which the raw requirements were extracted.</li> <li><strong>Annotated Dataset</strong>: A JSONLines file containing the annotated dataset, including labels for Named Entity Recognition (NER) and Relation Extraction (RE).</li> </ol> <p>Key features of the dataset:</p> <ul> <li>Language: French</li> <li>Number of requirements: 233</li> <li>Number of sentences: 19,725</li> <li>Number of words: 651,948</li> <li>Annotation tasks: Named Entity Recognition (NER) and Relation Extraction (RE)</li> </ul> <p>This dataset is relevant for NLP research focused on structured information extraction from domain-specific texts in the construction industry.</p>

opencc-by-4.0Oct 2024View details →
zenodo36/100

Datasets of "An Automatically Generated Annotated Corpus for Albanian Named Entity Recognition"

<p>This is an Albanian named entities annotation corpus generated automatically (silver-standard)&nbsp;from Wikipedia and WikiData. It is offered in Apache OpenNLP annotation format.</p> <p>Details of the generation approach may be found in the respective published paper:&nbsp;https://doi.org/10.2478/cait-2018-0009</p> <p>Attached are also the files that were used for generating the Albanian named entities gazetteer and the gazetteer itself in JSON format.</p>

opencc-by-4.0Feb 2018View details →
zenodo36/100

Schema.org mark-up data for named entities

<p>This dataset contains two files: original_data.zip, and website_5folds.zip</p> <p><strong>original_data.zip&nbsp;</strong>will unpack into three .csv files, Place.csv, CreativeWork.csv, and LocalBusiness.csv. Each file contains one entity on each row, and this entity belongs to a subclass of the class indicated by the file name. There are 8 columns:</p> <ul> <li>the first 2 columns are simply the index of the row</li> <li>description_t: the long textual description of the entity</li> <li>schemaorg_class: the schema.org class assigned to the entity</li> <li>name_tpage_domain: always empty</li> <li>name_t: the name of the entity</li> <li>page_domain: the website where the entity mark-up data is found</li> <li>label: an index for the schemaorg_class</li> <li>description: this is the name of the entity (name_t) plus the first sentence of its description (from description_t)</li> </ul> <p><strong>website_5folds.zip</strong> is a transformation of the original_data.zip. It unzips into three folders, Place, LocalBusiness, and CreativeWork. Inside each folder, there are five folders: 0, 1, 2, 3 and 4 indicating five folds. Inside each of the numbered sub-folder there is a train.csv and test.csv file. Then each csv file contains one entity on each row, with the following columns:</p> <ul> <li>the first column is simply the index of the row</li> <li>schemaorg_class: the schema.org class assigned to the entity</li> <li>name_t:&nbsp;the name of the entity</li> <li>description:&nbsp;this is the name of the entity (name_t) plus the first sentence of its description (from description_t)</li> <li>page_domain:&nbsp;the name of the entity plus the processed domain name. The process includes parsing the domain URL, extract the host name, applying word segmentation (tescobank -&gt; tesco bank), and removing stopwords and TLDs (co, uk, com, fr)</li> </ul> <p>As mentioned,&nbsp;website_5folds.zip is a transformation of the original_data.zip and in fact contains multiple replications of original_data.zip. It is created for 5 fold validation experiment while ensuring that there are no overlap in the page_domain of entities in training and test sets.&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo32/100

An Arabic Dataset for Disease Named Entity Recognition with Multi-Annotation Schemes

<p>This is a novel data descriptor that provides the Arabic natural language processing community with a dataset dedicated to named entity recognition tasks for diseases. The dataset comprises more than 60 thousand words, which were annotated manually by two independent annotators using the inside-outside (IO) annotation scheme. To ensure the reliability of the annotation process, the inter-annotator agreements rate was calculated, and it scored 95.14\%. Due to the lack of research efforts in the literature dedicated for studying Arabic multi-annotation schemes, a distinguishing and a novel aspect of this dataset is the inclusion of six more annotation schemes that will bridge the gap by allowing the researchers to explore and compare the effects of these schemes on the performance of the Arabic named entity recognizers. These annotation schemes are IOE, IOB, BIES, IOBES, IE, and BI. Additionally, five linguistic features, including part-of-speech tags, stopwords, gazetteers, lexical markers, and the presence of the definite article, are provided for each record in the dataset.</p>

opencc-byJun 2020View details →
zenodo32/100

Named Entity Recognition Datasets

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2024View details →
zenodo32/100

Romanian micro-blogging named entity recognition (MicroBloggingNERo)

<p>MicroBloggingNERo is a manually annotated corpus for named entity recognition in Romanian micro-blogging texts.&nbsp;<br> It provides gold annotations for organizations, locations, persons, time expressions, legal references, medical devices,<br> chemicals, anatomical parts and disorders found in micro-blogging texts. The text was anonymized, by replacing all<br> URLs with &lt;url&gt;, user references with &lt;user&gt;, person names, specific locations and organizations with new randomized names.&nbsp;<br> Anonymization was realized in the same way, regardless of the micro-blogging platform specific format.</p> <p>Since names were replaced with new random ones, any resemblance to real individuals is by pure chance of the random names<br> generator. No real person is depicted in the included messages.</p> <p><br> DATA</p> <p>The MicroBloggingNERo corpus is available in different formats: text, span-based, and token-based.&nbsp;</p> <p>Text files are in the folder &quot;text&quot; with .txt extension, in UTF-8 encoding.</p> <p>Span-based annotations are given in BRAT (https://brat.nlplab.org/) ann format. These annotations can be found in folders starting with &quot;ann_&quot;.</p> <p>Token-based annotations are given in CONLLUP files, following the CoNLL-U Plus format https://universaldependencies.org/ext-format.html .<br> Part-of-speech tagging was realized using UDPIPE.&nbsp;<br> Named entity annotations are placed in the column &quot;RELATE:NE&quot; (the 11th column) as defined in the &quot;global.columns&quot; metadata field.<br> Automatic processing was performed through the RELATE platform (https://relate.racai.ro).</p> <p>The archive contains:&nbsp;</p> <p>- ann_EVERYTHING&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations of organizations and time entities inside legal references were allowed.&nbsp;</p> <p>- ann_EVERYTHING_LARGEST_SPAN&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated. This affects primarily the legal references class.</p> <p>- ann_LEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; There are no overlapping annotations.&nbsp;</p> <p>- ann_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; There are no overlapping annotations.&nbsp;</p> <p>- ann_BIOMEDICAL<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: medical devices, chemicals, anatomical parts and disorders.&nbsp;<br> &nbsp; &nbsp; There are no overlapping annotations.&nbsp;</p> <p>- conllup_EVERYTHING_LARGEST_SPAN<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations, time, chemicals, medical devices, anatomical parts and disorders.&nbsp;<br> &nbsp; &nbsp; There are no overlapping annotations.&nbsp;</p> <p>- conllup_LEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .conllup format and contains annotations of: legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;</p> <p>- conllup_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .conllup format and contains annotations of: persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;</p> <p>- conllup_BIOMEDICAL<br> &nbsp; &nbsp; Folder in which all the files are in .conllup format and contains annotations of: medical devices, chemicals, anatomical parts and disorders.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;</p> <p>- text&nbsp;<br> &nbsp; &nbsp; Folder containing the raw texts.</p> <p>- splits.tsv<br> &nbsp; &nbsp; Proposed splits into train,test,valid following a distribution of 70-15-15% for each entity class, based on the ann_EVERYTHING_LARGEST_SPAN folder</p> <p>LICENSING</p> <p>This work is provided under the license CC BY-NC-ND 4.0 (Attribution-NonCommercial-NoDerivatives 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-nd/4.0/&nbsp;<br> and the full text here: https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode .&nbsp;</p> <p><br> CONTACT</p> <p>Research Institute for Artificial Intelligence &quot;Mihai Draganescu&quot;, Romanian Academy<br> Web: http://www.racai.ro&nbsp;<br> Contact emails: vasile@racai.ro , maria@racai.ro , vergi@racai.ro , elena@racai.ro</p>

opencc-by-nc-nd-4.0Jul 2022View details →
zenodo32/100

Romanian Named Entity Recognition in the Legal domain (LegalNERo)

<p>LegalNERo is a manually annotated corpus for named entity recognition in the Romanian legal domain.&nbsp;<br> It provides gold annotations for organizations, locations, persons, time and legal resources mentioned in legal documents (legal references).Starting with version 4, the legal references were annotated using fine-grained legal document types: Law. Ordinance, Publication, Decree, Decision, Treaty, Report, Order, Regulation, Directive, EmergencyOrdinance, Norm, Convention, Code, and Other.<br> Additionally it offers GEONAMES codes for the named entities annotated as location (where a link could be established).&nbsp;</p> <p>The LegalNERo corpus is available in different formats: span-based, token-based and RDF.&nbsp;<br> The Linguistic Linked Open Data (LLOD) version is provided in RDF-Turtle format.</p> <p>CONLLUP files conform to the CoNLL-U Plus format <a href="https://universaldependencies.org/ext-format.html">https://universaldependencies.org/ext-format.html</a> .<br> Part-of-speech tagging was realized using UDPIPE.&nbsp;<br> Named entity annotations are placed in the column &quot;RELATE:NE&quot; (the 11th column) as defined in the &quot;global.columns&quot; metadata field.<br> Similarly GEONAMES references are in the column &quot;RELATE:GEONAMES&quot; (the 12th column, last).<br> Automatic processing was performed through the RELATE platform (<a href="https://relate.racai.ro">https://relate.racai.ro</a>).</p> <p>ANN files conform to BRAT format (<a href="https://brat.nlplab.org/">https://brat.nlplab.org/</a>).<br> &nbsp;<br> The archive contains:&nbsp;</p> <p>- ann_LEGAL_PER_LOC_ORG_TIME_overlap&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations of organizations and time entities inside legal references were allowed.&nbsp;</p> <p>- ann_FGLEGAL_PER_LOC_ORG_TIME_overlap&nbsp;<br> &nbsp; &nbsp; Folder (corresponding to the above entry) in which all the files are in .ann format and contains annotations of: fine-grained legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations of organizations and time entities inside legal references were allowed.&nbsp;</p> <p>- ann_LEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;</p> <p>- ann_FGLEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder (corresponding to the above entry) in which all the files are in .ann format and contains annotations of: fine-grained legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;</p> <p>- ann_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .ann format and contains annotations of: persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; There are no overlapping annotations.&nbsp;</p> <p>- conllup_LEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .conllup format and contains annotations of: legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;<br> &nbsp; &nbsp; The annotation of these files was enhanced with GEONAMES codes (where linking was possible). &nbsp;</p> <p>- conllup_FGLEGAL_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder (corresponding to the above entry) in which all the files are in .conllup format and contains annotations of: fine-grained legal references, persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;<br> &nbsp; &nbsp; The annotation of these files was enhanced with GEONAMES codes (where linking was possible). &nbsp;</p> <p>- conllup_PER_LOC_ORG_TIME&nbsp;<br> &nbsp; &nbsp; Folder in which all the files are in .conllup format and contains annotations of: persons, locations, organizations and time.&nbsp;<br> &nbsp; &nbsp; Overlapping annotations were not allowed and only the longest named entities were annotated.&nbsp;<br> &nbsp; &nbsp; The annotation of these files was enhanced with GEONAMES codes (where linking was possible).</p> <p>- rdf&nbsp;<br> &nbsp; &nbsp; Folder containing the corpus in RDF-Turtle format.<br> &nbsp; &nbsp; All the annotations are available here in both span and token format.</p> <p>- text&nbsp;<br> &nbsp; &nbsp; Folder containing the raw texts.</p> <p>- splits_FGLEGAL_PER_LOC_ORG_TIME.tsv<br> &nbsp; &nbsp; This is a proposed split of the documents for training a NER system using the fine-grained entity classes.&nbsp;<br> &nbsp; &nbsp; The split was created randomly, while trying to ensure 15% of each entity type for validation, 15% for testing and 70% for training.<br> &nbsp;</p> <p><strong>NER System</strong></p> <p>A NER model generated using the LegalNERo corpus can be used online in the RELATE platform:&nbsp;https://relate.racai.ro/index.php?path=ner/demo</p> <p>This system was described in:&nbsp;Păiș, Vasile and Mitrofan, Maria and Gasan, Carol Luca and Coneschi, Vlad and Ianov, Alexandru. Named Entity Recognition in the Romanian Legal Domain. In Proceedings of the Natural Legal Language Processing Workshop 2021. Association for Computational Linguistics, Punta Cana, Dominican Republic, pp. 9--18, nov 2021</p> <p><br> <strong>LICENSING</strong></p> <p>This work is provided under the license CC BY-NC-ND 4.0 (Attribution-NonCommercial-NoDerivatives 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-nd/4.0/&nbsp;<br> and the full text here: https://creativecommons.org/licenses/by-nc-nd/4.0/legalcode .&nbsp;</p> <p><br> <strong>CONTACT</strong></p> <p>Research Institute for Artificial Intelligence &quot;Mihai Draganescu&quot;, Romanian Academy<br> Web: <a href="http://www.racai.ro">http://www.racai.ro</a>&nbsp;<br> Contact emails: vasile@racai.ro , maria@racai.ro</p>

opencc-by-nc-nd-4.0May 2021View details →
zenodo32/100

Enhancing Name Entity Recognition Through Hybrid Deep Learning in Natural Language Processing

Open the record for dataset details and reuse information.

opencc-by-4.0Jun 2024View details →
zenodo32/100

Named Entity Recognition in the Regesta of Emperor Frederick III. of the Holy Roman Empire

<p>This dataset contains Named Entity Recognition and Entity Linking across 8883 Regesta of Emperor Frederick III. of the Holy Roman Empire. The Regesta are from the Regesta Imperii Edition project: http://www.regesta-imperii.de/en/the-project.html.</p> <p>There are 4554 different names normalized to 2849 distinct entities, identified by the index entry identifier of the yet to be fully published index of the Regesta of Frederick III. All in all, 17258 instances of names are identified.</p> <p><strong>Not all named entities are identified!</strong></p> <p>This dataset is intended as a step stone for recording already identified entities and training NER classifiers for the Regesta Imperii. Some named entities were deliberately omitted.</p> <p>Only those that could be identified with reasonable certainty via the index were included, for example: &quot;Mgf. Albrecht von Brandenburg&quot; is included, while &quot;Albrecht&quot;, even when refering to the same entity as the former name, is not. Even of this subset, some entities may have been missed. A newer version may rectify that in the future.</p> <p>The JSONL format is modeled after the <a href="https://github.com/doccano/doccano">Doccano</a> format for entity linking.</p>

opencc-by-4.0Sep 2023View details →
zenodo28/100

Temporally-Informed Analysis of Named Entity Recognition

<p>This repository contains the data set developed for the paper:</p> <p>&ldquo;Shruti Rijhwani and Daniel Preoțiuc-Pietro. <em>Temporally-Informed Analysis of Named Entity Recognition.</em> In Proceedings of the Association for Computational Linguistics (ACL). 2020.&rdquo;</p> <p>It includes 12,000 tweets annotated for the named entity recognition task. The tweets are uniformly distributed over the years 2014-2019, with 2,000 tweets from each year. The goal is to have a temporally diverse corpus to account for data drift over time when building NER models.</p> <p>The entity types annotated are locations (LOC), persons (PER) and organizations (ORG). The tweets are preprocessed to replace usernames and URLs with a unique token. Hashtags are left intact and can be annotated as named entities.</p> <p><strong>Format</strong></p> <p>The repository contains the annotations in JSON format.</p> <p>Each year-wise file has the tweet IDs along with token-level annotations. The Public Twitter Search API (<a href="https://developer.twitter.com/en/docs/tweets/search">https://developer.twitter.com/en/docs/tweets/search</a>) can be used extract the text for the tweet corresponding to the tweet IDs.</p> <p><strong>Data Splits</strong></p> <p>Typically, NER models are trained and evaluated on annotations available at the model building time, but are used to make predictions on data from a future time period. This setup makes the model susceptible to temporal data drift, leading to lower performance on future data as compared to the test set.</p> <p>To examine this effect, we use tweets from the years 2014-2018 as the training set and random splits of the 2019 tweets as the development and test sets. These splits simulate the scenario of making predictions on data from a future time period.</p> <p>The development and test splits are provided in the JSON format.</p> <p><strong>Use</strong></p> <p>Please cite the data set and the accompanying paper if you found the resources in this repository useful.</p>

opencc-by-4.0Jun 2020View details →
zenodo28/100

NoSta-D Named Entity Annotation for German

<p>Dataset of the GermEval2014 Named Entity Recognition task.</p>

opencc-by-4.0Jul 2014View details →
zenodo28/100

Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences Derived from Biomedical Literature in the PubChemRDF

<p>This is the dataset used for the publication of &quot;Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences Derived from Biomedical Literature in the PubChemRDF&quot;.&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo24/100

Tutorial: Stanford Named Entity Recognizer installieren und deutsche Kategorien laden

<p>In der Videoreihe &bdquo;Named Entity Recognition und digitale Literaturanalyse&ldquo; zeigen wir, wie sich der <em>Stanford Named Entity Recognizer</em> installieren und zur Literaturanalyse nutzen l&auml;sst. Au&szlig;erdem stellen wir drei M&ouml;glichkeiten zur Verbesserung der Ergebnisse vor und zeigen, unter anderem, wie sich eine generierte Output-Datei in eine XML-Datei umwandeln l&auml;sst oder wie du dein eigenes NER-Modell trainieren kannst.&nbsp; <br>In diesem Video zeigen wir Schritt f&uuml;r Schritt, wie das Named Entity Recognition (NER)-Tool aus der Stanford Natural Language Processing Group installiert wird und die deutschen Kategorien &bdquo;Orte&ldquo;, &bdquo;Personen&ldquo;, &bdquo;Organisationen&ldquo; und &bdquo;Vermischtes&ldquo; geladen werden.</p> <p>&nbsp;</p> <p>Mehr Infos:&nbsp;</p> <ul> <li>Webseite der Stanford NER Group: <a href="https://nlp.stanford.edu">https://nlp.stanford.edu </a></li> <li>Stanford NER Herunterladen: <a href="https://nlp.stanford.edu/software/CRF-NER.shtml">https://nlp.stanford.edu/software/CRF-NER.shtml</a></li> <li>Schriftliche Einf&uuml;hrung in die Methodik der NER: <a href="https://fortext.net/routinen/methoden/named-entity-recognition-ner">https://fortext.net/routinen/methoden/named-entity-recognition-ner</a></li> </ul> <p>&Uuml;bersicht der Videoreihe auf Zenodo:</p> <ol> <li><a href="../records/10372231">Tutorial: Stanford Named Entity Recognizer installieren und deutsche Kategorien laden&nbsp;</a></li> <li><a href="../records/10372239">Tutorial: Stanford Named Entity Recognizer zur digitalen Literaturanalyse nutzen</a>&nbsp;</li> <li><a href="../records/10371086">Tutorial: Ein eigenes NER-Modell f&uuml;r die digitale Literaturanalyse trainieren</a></li> <li><a href="../records/10250582">Fallbeispiel: Figurenkonstellationen in Goethes Werther und Plenzdorfs neuem Werther</a></li> </ol> <p><a href="https://www.youtube.com/watch?v=hYed-ZqEzs8&amp;list=PLu-M0KuYw64pVltr9EazXvscw5DpRv01Q">Hier</a> zur Videoreihe auf Youtube</p>

opencc-by-4.0Dec 2023View details →
zenodo24/100

Tutorial: Stanford Named Entity Recognizer zur digitalen Literaturanalyse nutzen

<p>In der Videoreihe &bdquo;Named Entity Recognition und digitale Literaturanalyse&ldquo; zeigen wir, wie sich der <em>Stanford Named Entity Recognizer</em> installieren und zur Literaturanalyse nutzen l&auml;sst. Au&szlig;erdem stellen wir drei M&ouml;glichkeiten zur Verbesserung der Ergebnisse vor und zeigen, unter anderem, wie sich eine generierte Output-Datei in eine XML-Datei umwandeln l&auml;sst oder wie du dein eigenes NER-Modell trainieren kannst.&nbsp;<br>In diesem Schritt-f&uuml;r-Schritt-Tutorial zeigen wir schnell und einfach, wie der Stanford Named Entity Recognizer f&uuml;r literarische Texte genutzt werden kann.</p> <p>Mehr Infos:&nbsp;</p> <ul> <li>Webseite der Stanford NER Group: <a href="https://nlp.stanford.edu">https://nlp.stanford.edu </a></li> <li>Stanford NER Herunterladen: <a href="https://nlp.stanford.edu/software/CRF-NER.shtml">https://nlp.stanford.edu/software/CRF-NER.shtml</a></li> <li>Schriftliche Einf&uuml;hrung in die Methodik der NER: <a href="https://fortext.net/routinen/methoden/named-entity-recognition-ner">https://fortext.net/routinen/methoden/named-entity-recognition-ner</a></li> </ul> <p>&Uuml;bersicht der Videoreihe auf Zenodo:</p> <ol> <li><a href="../records/10372231">Tutorial: Stanford Named Entity Recognizer installieren und deutsche Kategorien laden&nbsp;</a></li> <li><a href="../records/10372239">Tutorial: Stanford Named Entity Recognizer zur digitalen Literaturanalyse nutzen</a>&nbsp;</li> <li><a href="../records/10371086">Tutorial: Ein eigenes NER-Modell f&uuml;r die digitale Literaturanalyse trainieren</a></li> <li><a href="../records/10250582">Fallbeispiel: Figurenkonstellationen in Goethes Werther und Plenzdorfs neuem Werther</a></li> </ol> <p><a href="https://www.youtube.com/watch?v=hYed-ZqEzs8&amp;list=PLu-M0KuYw64pVltr9EazXvscw5DpRv01Q">Hier</a> zur Videoreihe auf Youtube</p>

opencc-by-4.0Aug 2018View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record