Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
51
datasets available to search
ShareScore release 0.7.1
Dataset results
51 results for “relation extraction”
Relation Extraction Dataset for Dutch Biographical Texts
<p>A manually annotated dataset with relations relevant for biographical texts. The texts are in Dutch and are originally available in the Biographical portal of the Netherlands (http://www.biografischportaal.nl/)</p>
CrowdTruth Corpus for Open Domain Relation Extraction from Sentences
<p>This repository contains a ground truth corpus for open domain relation extraction from sentences, acquired with crowdsourcing and processed with <strong><a href="http://crowdtruth.org/">CrowdTruth</a></strong> metrics that capture ambiguity in annotations by measuring inter-annotator disagreement.</p> <p>The dataset contains annotations for 4,100 sentences sampled from Angeli et al. (1) and Riedel et al. (2), over 16 relations, with each sentence annotated by 15 workers. The sentences have been pre-processed with Distant Supervision (3) using the Freebase knowledge base, in order to identify the term pairs in each sentence that are likely to express a relation. The crowdsourced data was collected from <a href="http://figure-eight.com/">Figure Eight</a> and <a href="https://www.mturk.com/">Amazon Mechanical Turk</a>.</p> <p>This corpus has been discussed in the following papers:</p> <ul> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1809.00537">Crowdsourcing Semantic Label Propagation in Relation Classification</a></strong>. <a href="http://fever.ai/">FEVER</a> Workshop at <a href="http://emnlp2018.org/">EMNLP 2018</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="https://arxiv.org/abs/1711.05186">False Positive and Cross-relation Signals in Distant Supervision Data</a></strong>. <a href="http://www.akbc.ws/">AKBC</a> Workshop at <a href="http://nips.cc/">NIPS 2017</a>.</li> <li>Anca Dumitrache, Lora Aroyo and Chris Welty: <strong><a href="http://crowdtruth.org/wp-content/uploads/2017/03/collint17-open-domain.pdf">Disagreement in Crowdsourcing and Active Learning for Better Distant Supervision Quality</a></strong>. <a href="http://collectiveintelligenceconference.org/">Collective Intelligence 2017</a>.</li> </ul> <p>Sentence-level data is available in file: <code>|--data/output/aggregated_sentences.csv</code></p> <p>Worker-level data is available in file: <code>|--data/output/aggregated_workers.csv</code></p> <p>Raw crowdsourcig data is available in folder: <code>|--data/input/</code></p> <p>Results of the relation classification model are available in folder: <code>|--data/model_results/</code></p> <p> </p> <p>References</p> <p>(1) Angeli, Gabor, et al. "Combining distant and partial supervision for relation extraction." Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). 2014.</p> <p>(2) Riedel, Sebastian, et al. "Relation extraction with matrix factorization and universal schemas." Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). 2013.</p> <p>(3) Mintz, Mike, et al. "Distant supervision for relation extraction without labeled data." Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2-Volume 2. Association for Computational Linguistics, 2009.</p>
Definitions of terms extracted from data-related European Union laws, version 3
<p>Collection of definitions of terms in English, French, German, Italian and Spanish extracted from the following data-related European laws:</p> <ol> <li> <p><a href="http://data.europa.eu/eli/dir/2007/2/oj?locale=en">Directive 2007/2/EC of the European Parliament and of the Council of 14 March 2007 establishing an Infrastructure for Spatial Information in the European Community (<strong>INSPIRE</strong>)</a></p> </li> <li> <p><a href="http://data.europa.eu/eli/reg/2016/679/2016-05-04?locale=en">Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (<strong>General Data Protection Regulation</strong>) (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reco/2018/790/oj?locale=en">Commission Recommendation (EU) 2018/790 of 25 April 2018 on <strong>access to and preservation of scientific information</strong></a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2018/1807/oj?locale=en">Regulation (EU) 2018/1807 of the European Parliament and of the Council of 14 November 2018 on a framework for the <strong>free flow of non-personal data</strong> in the European Union (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://data.europa.eu/eli/dir/2019/790/oj?locale=en">Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on <strong>copyright and related rights in the Digital Single Market</strong> and amending Directives 96/9/EC and 2001/29/EC (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/dir/2019/1024/oj?locale=en">Directive (EU) 2019/1024 of the European Parliament and of the Council of 20 June 2019 on open data and the re-use of public sector information (recast) (<strong>Open Data Directive</strong>)</a></p> </li> <li><a href="https://eur-lex.europa.eu/eli/reg/2021/695/oj?locale=en">Regulation (EU) 2021/695 of the European Parliament and of the Council of 28 April 2021 establishing <strong>Horizon Europe</strong> – the Framework Programme for Research and Innovation, laying down its rules for participation and dissemination, and repealing Regulations (EU) No 1290/2013 and (EU) No 1291/2013 (Text with EEA relevance)</a></li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2022/868/oj?locale=en">Regulation (EU) 2022/868 of the European Parliament and of the Council of 30 May 2022 on European data governance and amending Regulation (EU) 2018/1724 (<strong>Data Governance Act</strong>) (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2022/1925/oj?locale=en">Regulation (EU) 2022/1925 of the European Parliament and of the Council of 14 September 2022 on contestable and fair markets in the digital sector and amending Directives (EU) 2019/1937 and (EU) 2020/1828 (<strong>Digital Markets Act</strong>) (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2022/2065/oj?locale=en">Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services and amending Directive 2000/31/EC (<strong>Digital Services Act</strong>) (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg_impl/2023/138/oj?locale=en">Commission Implementing Regulation (EU) 2023/138 of 21 December 2022 laying down a list of specific <strong>high-value datasets</strong> and the arrangements for their publication and re-use (Text with EEA relevance)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2023/2854/oj?locale=en">Regulation (EU) 2023/2854 of the European Parliament and of the Council of 13 December 2023 on harmonised rules on fair access to and use of data and amending Regulation (EU) 2017/2394 and Directive (EU) 2020/1828 (<strong>Data Act</strong>)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2024/903/oj?locale=en">Regulation (EU) 2024/903 of the European Parliament and of the Council of 13 March 2024 laying down measures for a high level of public sector interoperability across the Union (<strong>Interoperable Europe Act</strong>)</a></p> </li> <li> <p><a href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj?locale=en">Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828 (<strong>Artificial Intelligence Act</strong>) Text with EEA relevance.</a></p> </li> <li><a href="https://eur-lex.europa.eu/eli/reg/2024/2847/oj?locale=en">Regulation (EU) 2024/2847 of the European Parliament and of the Council of 23 October 2024 on horizontal cybersecurity requirements for products with digital elements and amending Regulations (EU) No 168/2013 and (EU) 2019/1020 and Directive (EU) 2020/1828 (<strong>Cyber Resilience Act</strong>) (Text with EEA relevance)</a></li> </ol>
Relative abundance of the CHC extracts of each population replicate's of I. uriae ticks from Iceland after log centered ratio transformation
<p>Relative abundance of the CHC extracts of each population replicate’s of I. uriae ticks from Iceland after log centered ratio transformation.</p>
TBGA: A Large-Scale Gene-Disease Association Dataset for Biomedical Relation Extraction
<p>This repository contains the TBGA dataset. TBGA is a large-scale, semi-automatically annotated dataset for Gene-Disease Association (GDA) extraction. The dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong> sentence from which the GDA was extracted.</li> <li><strong>relation:</strong> relation name associated with the given GDA.</li> <li><strong>h: </strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id: </strong>NCBI Entrez ID associated with the gene entity.</li> <li><strong>name:</strong> NCBI official gene symbol associated with the gene entity.</li> <li><strong>pos: </strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong> JSON object representing the disease entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated with the disease entity.</li> <li><strong>name:</strong> UMLS preferred term associated with the disease entity.</li> <li><strong>pos:</strong> list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>TBGA contains over 200,000 instances and 100,000 bags.<br> The zip file consists of one folder, named TBGA, containing the files corresponding to the dataset.</p> <p>If you use or extend our work, please cite the following: https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-022-04646-6#citeas<br> TBGA paper can be found at: <a href="https://rdcu.be/cKkY2">https://rdcu.be/cKkY2</a><br> TBGA code is available at: https://github.com/GDAMining/gda-extraction</p>
Building Large-Scale Gene-Disease Association Datasets for Biomedical Relation Extraction
<p>This repository contains the GDAb and GDAt datasets. GDAb and GDAt are large-scale, distantly supervised, and manually enhanced datasets for Gene-Disease Association (GDA) extraction. Each dataset consists of three text files, corresponding to train, validation, and test sets, plus an additional JSON file containing the mapping between relation names and IDs. Each record in train, validation, or test files corresponds to a single GDA extracted from a sentence. Records are represented as JSON objects with the following structure:</p> <ul> <li><strong>text:</strong> sentence from which the GDA was extracted.</li> <li><strong>relation:</strong> relation name associated to the given GDA.</li> <li><strong>h: </strong>JSON object representing the gene entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the gene entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the gene entity.</li> <li><strong>pos: </strong>list consisting of starting position and length of the gene mention within text.</li> </ul> </li> <li><strong>t:</strong> JSON object representing the disease entity, composed of: <ul> <li><strong>id: </strong>UMLS CUI associated to the disease entity.</li> <li><strong>name:</strong> UMLS preferred name associated to the disease entity.</li> <li><strong>pos:</strong> list consisting of starting position and length of the disease mention within text.</li> </ul> </li> </ul> <p>Both datasets contain over 2,500,000 sentences and 500,000 bags.<br> The zip file consists of two folders, GDAb and GDAt, containing the files corresponding to the two datasets, respectively.</p> <p> </p>
Evaluation dataset for Relation Extraction of relationships between Organisms and Natural-Products
<p>A curated evaluation dataset for end-to-end Relation Extraction of relationships between organisms and natural-products.</p> <p>Details about the manual annotation:</p> <ul> <li>For Chemicals: <ul> <li>The chemical labels are annotated as they appear in the abstract.</li> <li>In abstracts, singular chemicals and classes of chemicals produced by a specific organism were distinguished.</li> <li>The "type" attribute {“chemical”, “class”} is used to indicate the nature of the mentioned name.</li> <li>A "class" attribute for chemical entities has also been included if class information is present in the abstract.</li> <li>A Wikidata and PubChem identifiers were assigned to chemicals and classes when available.</li> </ul> </li> <li>For Organisms: <ul> <li>The organism labels are annotated as they appear in the abstract.</li> <li>If in an abstract, the genus name was mention first, e.g. "Plakinastrella sp." and then the specie name e.g "Plakinastrella clathrata" is precise, then only the specie name is used.</li> <li>A Wikidata identifier was assigned to all organisms.</li> <li>In some abstracts, only the genus name is mentioned.</li> </ul> </li> <li>For Relations: <ul> <li>Only the relations explicitly mentioned in the abstract are reported in the output labels.</li> <li>Relations are reported in their order of appearance in the abstract.</li> </ul> </li> </ul>
silverCID — a silver standard corpus for Chemical-induced Diseases relation extraction — Edit
<p>silverCID — a silver standard corpus for Chemical-induced Diseases relation extraction — Edit</p>
Medical Relation Extraction Gold Standard with CrowdTruth
<p>The lack of annotated datasets for training and benchmarking is one of the main challenges of Clinical Natural Language Processing. In addition, current methods for collecting annotations attempt to minimize disagreement between annotators, and therefore fail to model the ambiguity inherent in language. We propose the <strong>CrowdTruth</strong> method for collecting medical ground truth through crowdsourcing, based on the observation that disagreement between annotators can signal ambiguity in the text, target semantics, or the worker's interpretation.</p> <p>This repository contains a dataset of 3,984 English sentences for medical relation extraction, centering on the cause and treat medical relations, that have been processed with CrowdTruth disagreement analytics to capture ambiguity. In addition, we provide the raw crowdsourcing data used to compile this ground truth, as well as the task templates used to collect the data on CrowdFlower.</p>
Synthetic dataset for end-to-end Relation Extraction of relationships between Organisms and Natural-Products with Vicuna-13b-v1.5
<p>A new synthetic dataset (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products.</p><p>The new dataset was generated using <a href="https://huggingface.co/lmsys/vicuna-13b-v1.5">Vicuna-13b-v1.5</a>, derived from LLaMA 2. Like the model, the produced synthetic data are also submitted to the License of the model used for generation, see the original <a href="https://github.com/facebookresearch/llama/blob/main/LICENSE">LLaMA 2 license</a>.</p><p>The new dataset was created based on the top-1000 (per biological kingdom) LOTUS literature references extracted with the <a href="https://github.com/idiap/gme-sampler">GME-sampler</a>.</p><p>The dataset contains 10,405 items in the training set and 547 items in the validation set.</p><p>The dataset was generated using the same protocol as described in the <a href="https://github.com/idiap/gme-sampler">article</a>.</p>
BioPropaPhenKG on Multi-Relation Extraction Methods on Online Newspapers
<p>The coronavirus disease (COVID-19) spread rampantly around the world at the beginning of 2020 before the governments of each country could prevent it by making decisions based on medical data analysis. With proper formalization, the terabytes of new textual data available online every day could have been used for the early description and detection of cases of this virus. Since then, the number of Event-Based Surveillance (EBS) applications has increased exponentially. These applications aim to mine channels of unstructured data to detect signs of possible public health events. However, one problem with such systems is the need for expert intervention to define which event will be captured, which relevant terms should be used in the search, and to analyze the events to modify the search procedure constantly. Another problem is that many of these applications do not consider both spatial and temporal characteristics. Addressing such limitations, this article presents a novel approach. We propose the biomedical domain specialization of the Core Propagation Phenomenon Ontology (PropaPhen) to capture spatiotemporal characteristics of the propagation of health-related phenomena. We also propose the Description-Detection-Framework (DDF), which leverages PropaPhen, UMLS, and OpenStreetMaps to detect new medical events automatically. Finally, we demonstrate a use case with experiments on extracts from online newspapers about COVID-19. The results show that DDF can be useful for detecting clusters of suspicious cases of possible emerging health-related phenomena.</p> <p>BioPropaPhenKG, its ontology and other useful information can be found in <a href="../records/10911980">https://zenodo.org/records/10911980</a>. The code used for this use case can be found in <a href="https://github.com/Gabriel382/DDPF-Health-Risks">https://github.com/Gabriel382/DDPF-Health-Risks</a> . Finally, the datasets used where UMLS MetamorphoSys, OpenStreetMaps, Wikidata, <a href="https://aylien.com/blog/free-coronavirus-news-dataset">Aylien</a> (data only from November of 2019).</p> <p> </p> <p>To read, you just need to load it with Neo4j:4.4.3. Alternatively, you can open it with docker using the following command: </p> <p>docker run --interactive --tty --rm \<br> --publish=7474:7474 --publish=7687:7687 \<br> --volume=/path-to-data-folder:/data --user="$(id -u):$(id -g)"\<br> neo4j:4.4.3 \<br>neo4j-admin load --from=/data/BioPropaPhenKG-Journal-MultiRE.dump --database "neo4j" --force</p>
An annotated corpus of clinical trial publications supporting schema-based relational information extraction
<p>Repository of an annotated corpus of clinical trial abstracts supporting schema-based relational information extraction and the code for the inter-annotation agreement calculation and the baseline information extraction method.</p>
Synthetic dataset for end-to-end Relation Extraction of relationships between Organisms and Natural-Products with Mixtral-8x7B-Instruct-v0.1
<p>A new synthetic dataset (training/validation) for end-to-end Relation Extraction of relationships between Organisms and Natural-Products.</p> <p>The new dataset was generated using <a href="https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1">Mixtral-8x7B-Instruct-v0.1</a>. Like the model, the produced synthetic data are also submitted to the License of the model used for generation (apache-2.0).</p> <p>The new dataset was created based on the top-1000 (per biological kingdom) LOTUS literature references extracted with the <a href="https://github.com/idiap/gme-sampler">GME-sampler</a>.</p> <p>The dataset contains 8,913 items in the training set and 344 items in the validation set.</p> <p>The dataset was generated using the same protocol as described in the article.</p>
EnzChemRED, a rich enzyme chemistry relation extraction dataset
<h1>Abstract</h1> <p>Expert curation is essential to capture knowledge of enzyme functions from the scientific literature in FAIR open knowledgebases but cannot keep pace with the rate of new discoveries and new publications. In this work we present EnzChemRED, for Enzyme Chemistry Relation Extraction Dataset, a new training and benchmarking dataset to support the development of Natural Language Processing (NLP) methods such as (large) language models that can assist enzyme curation. EnzChemRED consists of 1,210 expert curated PubMed abstracts in which enzymes and the chemical reactions they catalyze are annotated using UniProtKB and ChEBI identifiers. We show that fine-tuning language models with EnzChemRED significantly boosts their ability to identify proteins and chemicals in text (86.30% F1 score) and to extract the chemical conversions in which they participate (86.66% F1 score), and the enzymes that catalyze those conversions (83.79% F1 score). We apply our methods to abstracts at PubMed scale to create a draft map of enzyme functions in literature to guide curation efforts in UniProtKB and the reaction knowledgebase Rhea.</p> <p><strong>Corresponding authors:</strong> Alan Bridge (alan.bridge@sib.swiss) and Zhiyong Lu (zhiyong.lu@nih.gov)</p> <h1>Content</h1> <p>This repository contains data to support the development of natural language processing (NLP) methods to mine biochemical reactions from text for Rhea and UniProt.</p>
Pharmacokinetic Relation Extraction Database (PRED)
<p>Annotated data to perform Relation Extraction and extract pharmacokinetic (PK) parameter estimates from scientific text. </p> <p>Training, development and test files are released in <a href="https://jsonlines.org/">JSONL format</a> and store the annoated data for training and evaluating end-to-end relation extraction models. </p> <p>Each line in the JSONL files corresponds to an annotated sentence with the following information: </p> <ul> <li><strong>text</strong>: Raw sentence</li> <li><strong>relations: </strong>List or relations each containing: <ul> <li><strong>head_span </strong>: Head entity of the relation as a dictionary containing start character, end character and entity type label</li> <li><strong>child_span: </strong>Child entity of the relation </li> <li><strong>label</strong>: relation type label (i.e. either C_VAL, D_VAL or RELATED)</li> </ul> </li> <li><strong>spans</strong>: List of entities mentioned in the sentence, specifying: (1) the character-level boundaries and (2) the entity type label of each annotation (i.e. either PK, VALUE, UNITS, RANGE or COMPARE). This field is not strictly required to train the model since all spans are defined within the relations.</li> <li><strong>sentence_hash</strong>: Unique sentence ID</li> <li><strong>metadata</strong>: Metadata with unique article, paragraph and sentence identifiers and the article section from which the sentence was extracted</li> </ul> <p> </p>
Average daily air temperature, precipitation and relative sunshine duration for Vallon de Nant catchment, extracted from gridded MeteoSwiss data (1961-2020)
<p>This excel file contains time series of daily temperature, precipitation and relative sunshine duration obtained as the spatial average of gridded data sets. The underlying original gridded data sets produced by MeteoSwiss are known as RhiresD, TabsD and SrelD. All meta data are included in the excel file.</p> <p>The data can e.g. be used for hydrological modelling. Comparison to local station data is not included.</p> <p><strong>This data set as well as the original data set should be cited</strong>.</p>
Relations from Italian Wikipedia using Unsupervised Information Extraction
<p>This dataset contains relations extracted from the Italian Wikipedia by the WikiOIE framework.<br> WikiOIE is based on UDPipe and the Universal Dependencies project for text processing.<br> It easily allows customizing the information extraction (IE) approach to automatically extract triples (subject, predicate, object).<br> This dataset contains relations extracted by two unsupervised IE methods. The former (<strong>simple</strong>) is based only on PoS-tag patterns; the latter (<strong>simpledep</strong>) also uses syntactic dependencies. <br> The extraction process is provided in JSON format.</p> <p>More information and the Java code are available here https://github.com/pippokill/WikiOIE</p> <p>Pierluigi Cassotti, Lucia Siciliani, Pierpaolo Basile,Marco de Gemmis, and Pasquale Lops. 2021. Extracting relations from Italian Wikipedia using unsupervised information extraction. In Proceedings of the 11th Italian Information Retrieval Workshop 2021 (IIR 2021). CEUR-WS.</p>
MEDDOPLACE Corpus: Gold Standard annotations for Medical Documents Place-related Content Extraction
<p><strong>MEDDOPLACE</strong> stands for MEDical DOcument PLAce-related Content Extraction. It is a shared task and set of resources focused on the detection, normalization (entity linking/toponym resolution) and classification of different kinds of places, as well as related types of information such as clinical departments, nationalities or patient movements, in medical documents in Spanish.</p> <p>This repository includes the corpus' <strong>train and test sets</strong> in multiple formats, as well as the <strong>SNOMED gazetteer</strong>, <strong>cross-mapping</strong> between SNOMED and MeSH and the <strong>multilingual silver standard in 8 languages </strong>(Catalan, English, French, Italian, Dutch, Portuguese, Romanian and Swedish). For more information, please check the attached README file.</p> <p>MEDDOPLACE was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of IberLEF 2023. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a>.</p> <p> </p> <p><strong>Please cite if you use this resource:</strong></p> <p>Salvador Lima-López, Eulàlia Farré-Maduell, Antonio Miranda-Escalada, Vicent Brivá-Iglesias and Martin Krallinger. NLP applied to occupational health: MEDDOPROF shared task at IberLEF 2021 on automatic recognition, classification and normalization of professions and occupations from medical texts. In Procesamiento del Lenguaje Natural, 67. 2021.</p> <pre><code>@article{meddoplace, title={MEDDOPLACE Shared Task overview: recognition, normalization and classification of locations and patient movement in clinical texts}, author={Lima-López, Salvador and Farré-Maduell, Eulàlia and Brivá-Iglesias, Vicent and Gasco-Sanchez, Luis and Krallinger, Martin}, journal = {Procesamiento del Lenguaje Natural}, volume = {71}, year={2023}, issn = {1135-5948},<br>DOI = {10.26342/2023-71-23}, url = {http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561/3961}, pages = {301--311} }</code></pre> <p><strong>Related Links:</strong></p> <p>- MEDDOPLACE website: <a href="https://temu.bsc.es/meddoplace">https://temu.bsc.es/meddoplace</a></p> <p>- MEDDOPLACE overview paper: <a href="http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561">http://journal.sepln.org/sepln/ojs/ojs/index.php/pln/article/view/6561</a></p> <p>- Annotation Guidelines (Spanish): <a href="https://doi.org/10.5281/zenodo.7775234">https://doi.org/10.5281/zenodo.7775234</a></p> <p>- Annotation Guidelines (English): <a href="https://doi.org/10.5281/zenodo.7928145">https://doi.org/10.5281/zenodo.7928145</a></p> <p><strong>License</strong></p> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <p><strong>Contact</strong></p> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p>
BiodivNERE: Gold Standard Corpora for Named Entity Recognition and Relation Extraction in Biodiversity Domain
<p>BiodivNER+RE are two gold-standard manually annotated corpora that are meant to be for Named Entity Recognition (NER) and Relation Extraction (RE) tasks based on abstracts and metadata files from the biodiversity domain</p> <p>Such corpora are designed for machine learning techniques, for example, NER as TokenClassification technique, and RE as SequenceClassification technique.</p>
Extracted microbial terms from Wikipedia - Marine Microbiology related pages (unfiltered)
<p>Marine microorganism related terms were extracted from various Wikipedia pages to scan the ODIS graph using the Wikipedia API library. The pages include topics such as "Marine microorganisms", "Marine microbiome", "Marine viruses", "Marine bacteria", "Bacterioplankton", "Bacterial motility", "Marine prokaryotes", "Marine archaea", "Marine protists", "Marine fungi", "Mycoplankton", "Marine microanimals", "Ichthyoplankton", "Marine primary production", "Algae", "Marine microplankton", "Marine microbenthos", "Sea ice microbial communities", "Hydrothermal vent microbial communities", "deep biosphere", and "microbial dark matter". </p> <p> </p> <p>We used the english version of the spaCy library, a language processing (NLP) tool, to process text content extracted from Wikipedia and extract meaningful terms. We passed the combined text content to the spaCy model to generate a Doc object, which contains words and their associated linguistic information. After extracting these words from the Doc object, spaCy keeps track of their frequencies using a dictionary and filters out words that are not purely alphabetic, common stop words, and those shorter than four characters. Then, these words are each converted to lowercase and sorted based on their frequency.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.