Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36
datasets available to search
ShareScore release 0.9.0
Dataset results
36 results for “text corpus”
Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text
<p>Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular sentiments of videos on social media based on viewer comments. However, comments from social media do not follow strict rules of grammar, and they contain mixing of more than one language, often written in non-native scripts. Non-availability of annotated code-mixed data for a low-resourced language like Tamil also adds difficulty to this problem. To overcome this, we created a gold standard Tamil-English code-switched, sentiment-annotated corpus containing 15,744 comment posts from YouTube. In this paper, we describe the process of creating the corpus and assigning polarities. We present inter-annotator agreement and show the results of sentiment analysis trained on this corpus as a benchmark.</p>
Archi text corpus
<p>Archi belongs to the Lezgic group of the Nakh-Daghestanian (North-East Caucasian) languages, being quite loosely related to the rest of the group. It has been long time surrounded by non-Lezgic languages and therefore has kept and/or acquired a number of peculiar features.</p> <p>Here is presented a sample of texts collected in the village of Archi in 2006 and 2007. In total, over 50 texts of various genres have been recorded, including stories, conversations, tales, legends and songs. Most of them were recorded in both video and audio.</p> <p>Two kinds of texts were recorded. First, some 30 previously published (1977) texts were re-recorded in video and audio, read by one of three speakers. (The original recordings do not exist anymore). Second, new texts have been collected, mostly dialogues and stories. Texts available in this version were all originally published in [Kibrik et al. 1977].</p> <p><em>Кибрик А. Е., Кодзасов С. В., Оловянникова И. П., Самедов Д. С.</em> Арчинский язык. Тексты и словари. — М.: МГУ, 1977.<br> [Kibrik, Aleksandr E.; Kodzasov, S. V.; Olovjannikova, I. P. & Samedov, D. S. (1977). <em>Arčinskij jazyk. Teksiy i slovari</em>. Moscow: Izdatel'stvo moskovskogo universiteta.]</p> <p>The project was generously supported by NSF grant #0553546 «Five languages of Eurasia» (PI under the Documenting Endangered Languages Program, and by RFBR grants № 05-06-80351 «Minority languages and cultures: On the verge of extinction» and № 08-06-00345 «Multimedia corpora for endangered languages».</p>
OpenITI: a Machine-Readable Corpus of Islamicate Texts
<p><strong>Co-PIs</strong>: Matthew Thomas Miller (University of Maryland, College Park), Maxim G. Romanov (University of Hamburg), Sarah Bowen Savant (Aga Khan University—ISMC, London).</p> <p><em>Open Islamicate Texts Initiative</em> (<strong>OpenITI</strong>, see <a href="https://openiti.org/">https://openiti.org/</a>) is a multi-institutional effort to construct the first machine-actionable scholarly corpus of premodern Islamicate texts. Led by researchers at the Aga Khan University, Institute for the Study of Muslim Civilisations (AKU-ISMC), University of Hamburg (UH), and the Roshan Institute for Persian Studies at the University of Maryland (College Park) and an interdisciplinary advisory board of leading digital humanists and Islamic, Persian, and Arabic studies scholars, <strong>OpenITI</strong> aims to provide the essential textual infrastructure in Arabic, Persian and other Islamicate languages for new forms of textual analysis and digital scholarship. In the process, OpenITI will enable new synergies between Digital Humanities and the inter-related Islamicate fields of Islamic, Persian, and Arabic Studies. In addition to support from the researchers’ home institutions, it is supported by funding from the <a href="https://erc.europa.eu/">European Research Council</a> under the European Union’s Horizon 2020 research and innovation programme, awarded to the <a href="http://kitab-project.org/">KITAB</a> project (Grant Agreement No. 772989, PI Sarah Bowen Savant) and the <a href="https://www.qnl.qa/en">Qatar National Library</a>.</p> <p>Currently, <strong>OpenITI</strong> contains almost exclusively Arabic texts, which were first assembled into a corpus within the <strong>OpenArabic</strong> project, developed first at Tufts University (at <em>The Perseus Project</em>, 2013–2015) and then at Leipzig University (at the Alexander von Humboldt Chair for Digital Humanities, 2015–2017)—in both cases with the support and under the patronage of Prof. Gregory Crane. The much more limited number of Persian texts were compiled during 2015–2016 in the Persian Digital Library (PDL) pilot (see <a href="https://persdigumd.github.io/PDL/">Persian Digital Library by PersDigUMD</a>) at Roshan Institute for Persian Studies at the University of Maryland. These texts have not been made fully compatible with OpenITI mARkdown yet and will be made fully available in next releases.</p> <p>This release contains all digital versions of the same text that are available in the OpenITI corpus . <strong>We also release a <a href="https://doi.org/10.5281/zenodo.7764025">'primary' version of the corpus</a></strong> that contains a single digital version for each text in the corpus that is marked as 'PRI' in the corpus metadata and may be more convenient for some use cases.</p> <p><strong>Note on Release Numbering</strong>: Version <strong>2019.1.1</strong>—where <strong>2019</strong> is the year of the release, the first dotted number—<strong>.1</strong>—is the ordinal release number in 2019, and the second dotted number—<strong>.1</strong>—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.</p> <p>For more details: <a href="https://github.com/OpenITI/RELEASE">https://github.com/OpenITI/RELEASE</a></p> <p><strong>Note: </strong>In case of any issues with unzipping the files on Windows using built-in utilities, please use free softwares, such as WinRAR and 7zip.</p> <p> </p>
The Machine-Actionable Ancient Text (MAAT) Corpus
<p>The Machine-Actionable Ancient Text (MAAT) Corpus is a new resource providing training and evaluation data for restoring lacunae in ancient Greek, Latin, and Coptic texts. Current text restoration systems require large amounts of data for training and task-relevant means for evaluation. The MAAT Corpus addresses this need by converting texts available in EpiDoc XML format into a machine-actionable format that preserves the most textually salient aspects needed for machine learning: the text itself, unclear letters, restorations, and lacunae. Structured test cases are generated from the corpus that align with the actual text restoration task performed by papyrologists and epigraphist, enabling more realistic evaluation than the synthetic tasks used previously. The initial 1.0 beta release contains approximately 134,000 text editions, 178,000 text blocks, and 750,000 individual restorations, with Greek and Latin predominating. This corpus aims to facilitate the development of computational methods to assist scholars in accurately restoring ancient texts.</p>
InTeReC: In-text Reference Corpus - Single References Dataset
<p>This dataset contains a set of sentences extracted from articles published by the Public Library of Science (PLOS) up to September 2013. Information is given on the position of the sentences relative to the article and the section in which they appear, the section type with respect to the four main types of the IMRaD structure, as well as verb phrases that occur in the sentence. Each sentence contains one single in-text reference.</p> <p>The dataset is in the CSV format. Size: 314023 sentences.</p> <p>Column list:</p> <ul> <li><em>journal</em>: journal title</li> <li><em>doi</em>: DOI of the article from which the sentence was extracted</li> <li><em>article-length</em>: size of the article, as number of sentences</li> <li><em>article-pos</em>: position of the sentence in the article, as number of sentences from the beginning of the article</li> <li><em>section-length</em>: size of the section, as number of sentences</li> <li><em>section-pos</em>: position of the sentence in the section, as number of sentences from the beginning of the section</li> <li><em>section-type</em>: section type (see below)</li> <li><em>sentence-text</em>: full text of the sentence</li> <li><em>verb-phrases</em>: a list of verb phrases that occur in the sentence, comma separated</li> </ul> <p>Possible section types are:</p> <ul> <li>I: Introduction</li> <li>M: Methods</li> <li>R: Results</li> <li>D: Discussion</li> <li>MR: Methods and Results</li> <li>RD: Results and Discussion</li> </ul> <p> </p> <p>Full description of the construction of the dataset is published in:</p> <p>Marc Bertin and Iana Atanassova (2018) InTeReC : an In-text Reference corpus for applying Natural Language Processing to Bibliometrics. Bibliometric-enhanced Information Retrieval: 7th International BIR workshop (7th BIR workshop) at the 40th European Conference on Information Retrieval (ECIR).</p>
Webis Wikipedia Text Reuse Corpus 2018 (Webis-Wikipedia-Text-Reuse-18)
<p>The Wikipedia Text Reuse Corpus 2018 (Webis-Wikipedia-Text-Reuse-18) containing text reuse cases extracted from within Wikipedia and in between Wikipedia and a sample of the Common Crawl.</p> <p>The corpus has following structure:</p> <ul> <li>wikipedia.jsonl.bz2: Each line, representing a Wikipedia article, contains a json array of article_id, article_title, and article_body</li> <li>within-wikipedia-tr-01.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (source article id), t_id (target article id), s_text (source text), t_text (target text)</li> <li>within-wikipedia-tr-02.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (source article id), t_id (target article id), s_text (source text), t_text (target text)</li> <li>preprocessed-web-sample.jsonl.xz: Each line, representing a web page, contains a json object of d_id, d_url, and content</li> <li>without-wikipedia-tr.jsonl.bz2: Each line, representing a text reuse case, contains a json array of s_id (Wikipedia article id), d_id (web page id), s_text (article text), d_content (web page content)</li> </ul> <p>The datasets were extracted in the work by Alshomary et al. 2018 that aimed to study the text reuse phenomena related to Wikipedia at scale. A pipeline for large scale text reuse extraction was developed and used on Wikipedia and the CommonCrawl.</p>
A Corpus of Persian Literary Text
<p>A Corpus of Persian Literary Text</p> <p>A corpus of Persian literary text mainly focusing on poetry, covering the 9th to 21st century. Annotated for century and style, with additional partial annotation of rhetorical figures (e.g. metaphor, metonymy, simile, etc.)</p>
OpenITI: a Machine-Readable Corpus of Islamicate Texts; Primary Version
<p>This is a partial release of the data in the <a href="https://doi.org/10.5281/zenodo.3082463">OpenITI corpus</a>. Whereas the main release of the corpus often contains multiple digital versions of the same text, this release contains a single digital version for each text in the corpus (the version marked as "primary" (PRI) in the corpus metadata).</p> <p>The version numbers corresponds to the versions of the main releases.</p> <p> </p> <p> </p>
Malicious Forensic Texts Corpus
<p>The Malicious Forensic Texts (MFT) corpus is a corpus of authentic malicious forensic texts that has been compiled in order to study their register variation.</p>
Text corpus of Kepler's Astronomia nova
<p>The JSON file contains preprocessed paragraphs of Kepler’s Astronomia Nova for machine learning. The database is derived from Donahue’s translation: Kepler, Johannes, New Astronomy, rev. edition, tr. by William H. Donahue, Green Lion Press, 2015. The text was digitized using OCR and automated text processing aiming at “pure” text containing machine-readable sentences in UTF8. Special characters, reference marks, and other markings were removed. OCR artefacts and errors may remain. For the authoritative text see Donahue’s edition. Digital Latin version cf. Kepler’s Gesammelte Werke.</p>
Classification of hierarchical text using geometric deep learning: the case of clinical trials corpus
<p>We consider the hierarchical representation of documents as graphs and use geometric deep learning to classify them into different categories. While graph neural networks can efficiently handle the variable structure of hierarchical documents using the permutation invariant message passing operations, we show that we can gain extra performance improvements using our proposed selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents. We applied our model to classify clinical trial (CT) protocols into completed and terminated categories. We use bag-of-words based as well as pre-trained transformer-based embeddings to featurize the graph nodes, achieving f1-scores $\simeq 0.85$ on a publicly available large scale CT registry of around 360K protocols. We further demonstrate how the selective pooling can add insights into the CT termination status prediction.</p>
Text-fig. 4. Briveichthys chantepieorum gen. et sp. nov. a, b: photograph and drawing of the parasphenoid in dorsal view, GMC 15, whitened, scale bars 5 mm; c: maxillary plate in medial view with horizontal lamina along the ventral edge of the bone, segment of the lower jaw with assembly of large slender teeth of the inner row and coronoids with small teeth, GMC 15, whitened, scale bar 5 mm; d: detail of the sculpture on the maxilla and dentalosplenial, GMC 126, whitened, scale bar 5 mm; e: detail of the teeth of the inner and outer row and coronoids on the lower jaw, the frame delineates the area illustrated in (f) at higher magnification, GMC 15, scale bar 2 mm; f: microsculpture formed by elliptical proximo-distally elongated protuberances on the large conical teeth, GMC 15, scale bar 100 µm; g: small fringing fulcra tightly attached to the anterior edge of a lepidotrichium, individual fulcral scales are indicated by arrows, GMC 18, scale bar 2 mm. Abbreviation: bhf – bucco-hypophysial foramen, Cor – coronoids, cp – corpus parasphenoidis, De – dentalosplenial, hl – horizontal lamina, mc – pores of the mandibular sensory canal, Mx – maxilla, paa – processus ascendens anterior, pap – processus ascendens posterior. in New Actinopterygians From The Permian Of The Brive Basin, And The Ichthyofaunas Of The French Massif Central
Text-fig. 4. Briveichthys chantepieorum gen. et sp. nov. a, b: photograph and drawing of the parasphenoid in dorsal view, GMC 15, whitened, scale bars 5 mm; c: maxillary plate in medial view with horizontal lamina along the ventral edge of the bone, segment of the lower jaw with assembly of large slender teeth of the inner row and coronoids with small teeth, GMC 15, whitened, scale bar 5 mm; d: detail of the sculpture on the maxilla and dentalosplenial, GMC 126, whitened, scale bar 5 mm; e: detail of the teeth of the inner and outer row and coronoids on the lower jaw, the frame delineates the area illustrated in (f) at higher magnification, GMC 15, scale bar 2 mm; f: microsculpture formed by elliptical proximo-distally elongated protuberances on the large conical teeth, GMC 15, scale bar 100 µm; g: small fringing fulcra tightly attached to the anterior edge of a lepidotrichium, individual fulcral scales are indicated by arrows, GMC 18, scale bar 2 mm. Abbreviation: bhf – bucco-hypophysial foramen, Cor – coronoids, cp – corpus parasphenoidis, De – dentalosplenial, hl – horizontal lamina, mc – pores of the mandibular sensory canal, Mx – maxilla, paa – processus ascendens anterior, pap – processus ascendens posterior.
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - model weights
<p>In the related <a href="https://github.com/HybridNLP2018/tutorial">notebook </a>we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains the model weights trained on such large corpora.</p>
Cross-modal (text and figures) Analysis of a Scientific Corpus from Semantic Scholar - images
<p>In this notebook we show the application of cross-modal techniques to improve the categorization of scientific papers through content related to figures, using both the textual part (captions) and the visual part (figures, diagrams, images) jointly. To this purpose, we use several CNN models and execute some experiments, illustrating our approach. This deposit contains high quality versions of the images used in the analysis.</p>
BVS Corpus: A Multilingual Parallel Corpus and Translation Experiments of Biomedical Scientific Texts
<p>The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME in agreement with the Pan American Health Organization (OPAS). Abstracts are available in English, Spanish, and Portuguese, with a subset in more than one language, thus being a possible source of parallel corpora. In this article, we present the development of parallel corpora from BVS in three languages: English, Portuguese, and Spanish. Sentences were automatically aligned using the Hunalign algorithm for EN/ES and EN/PT language pairs, and for a subset of trilingual articles also. We demonstrate the capabilities of our corpus by training a Neural Machine Translation (OpenNMT) system for each language pair, which outperformed related works on scientific biomedical articles. Sentence alignment was also manually evaluated, presenting an average 96\% of correctly aligned sentences across all languages. Our parallel corpus is freely available, with complementary information regarding article metadata.</p> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
Complete Medline abstracts corpus between 2015-2019 annotated Whatizit text annotation tool
<p><strong>Background: </strong></p> <p>Whatizit is a text processing system that allows you to do text-mining tasks on text. It is great at identifying molecular biology terms and linking them to publicly available databases. Identified terms are wrapped with XML tags that carry additional information, such as the primary keys to the databases where all the relevant information is kept. The wrapping XML is translated into HTML hypertext links. This service is highly appreciated by people who are reading literature and need to quickly find more information about a particular term, e.g. its Gene Ontology term.</p> <p>Whatizit is used in identifying formalized language patterns, specialized, syntactically formalized, technical notation. The annotation speed of a given pipeline is almost independent of the size of the vocabulary behind it and is currently based on pattern matching. In addition, several vocabularies can be integrated in a single pipeline.</p> <p><strong>Methodology:</strong></p> <p>The pipeline used is comprised of 175k Gene Ontology terms (preferred labels + synonyms).<br> The annotation on Medline 2015-2019 corpus is done with <em>Gene Ontology (GO)</em> integrated dictionary.</p> <p>The .zip file contains 10 XML files - each file is for half an year of MEDLINE annotated abstracts.<br> In addition to the abstract, the title is also annotated for further information enrichment.</p> <p>Respective DOIs, PMIDs are also included in the XML, when applicable.</p> <p><strong>Further development:</strong></p> <p>The XML files can be converted into JSON, JSON-LD format.</p> <p> </p> <p> </p>
DrugProt corpus: Biocreative VII Track 1 - Text mining drug and chemical-protein interactions
<p>Gold Standard annotations of the DrugProt corpus (training and development sets). Also, test and background sets.</p><p> </p><p><strong>Please cite if you use any DrugProt resource:</strong></p><p>Antonio Miranda-Escalada, Farrokh Mehryary, Jouni Luoma, Darryl Estrada-Zavala, Luis Gasco, Sampo Pyysalo, Alfonso Valencia, Martin Krallinger, Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical–protein relations, <i>Database</i>, Volume 2023, 2023, baad080</p><blockquote><p><i>@article{miranda2023overview, title={Overview of DrugProt task at BioCreative VII: data and methods for large-scale text mining and knowledge graph generation of heterogenous chemical--protein relations}, author={Miranda-Escalada, Antonio and Mehryary, Farrokh and Luoma, Jouni and Estrada-Zavala, Darryl and Gasco, Luis and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, journal={Database}, volume={2023}, pages={baad080}, year={2023}, publisher={Oxford University Press UK} }</i></p></blockquote><p>Miranda, Antonio, et al. "Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations." <i>Proceedings of the seventh BioCreative challenge evaluation workshop</i>. 2021.</p><blockquote><p><i>@inproceedings{miranda2021overview, title={Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene/protein relations}, author={Miranda, Antonio and Mehryary, Farrokh and Luoma, Jouni and Pyysalo, Sampo and Valencia, Alfonso and Krallinger, Martin}, booktitle={Proceedings of the seventh BioCreative challenge evaluation workshop}, year={2021} }</i></p></blockquote><p> </p><p><strong>Introduction</strong></p><p>The aim of the DrugProt track (similar to the previous CHEMPROT task of BioCreative VI) is to promote the development and evaluation of systems that are able to automatically detect in relations between chemical compounds/drug and genes/proteins. We have therefore generated a manually annotated corpus, the <i>DrugProt corpus</i>, where domain experts have exhaustively labeled:(a) all chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types (<i>DrugProt relation classes</i>). There is also an increasing interested in the integration of chemical and biomedical data understood as curation of relationships between biological and chemical entities from text and storing such information in form of structured annotation databases. Such databases are of key relevance not only for biological but also for pharmacological and clinical research. A range of different types chemical-protein/gene interactions are of key relevance for biology, including metabolic relations (e.g. substrates, products) inhibition, binding or induction associations.</p><p>The DrugProt track aims to address these needs and to promote the development of systems able to extract chemical-protein interactions that might be of relevance for precision medicine as well as for drug discovery and basic biomedical research.</p><p>The DrugProt track in BioCreative VII (BC VII) will explore recognition of chemical-protein entity relations from abstracts.</p><p>Teams participating in this track are provided with:</p><ul><li>PubMed abstracts</li><li>Manually annotated chemical compound mentions</li><li>Manually annotated gene/protein mentions</li><li>Manually annotated chemical compound-protein relations</li></ul><p> </p><p><strong>Zip structure:</strong></p><ul><li>Training set folder with<ul><li>drugprot_training_abstracts.tsv: PubMed records</li><li>drugprot_training_entities.tsv: manually labeled mention annotations of chemical compounds and genes/proteins</li><li>drugprot_training_relations.tsv: chemical-protein relation annotations</li></ul></li><li>Development set folder with<ul><li>drugprot_development_abstracts.tsv</li><li>drugprot_development_entities.tsv</li><li>drugprot_development_relations.tsv</li></ul></li><li>Test+background set folder with<ul><li>test_background_abstracts.tsv</li><li>test_background_entities.tsv</li></ul></li></ul><p> </p><p><strong>Data format description</strong></p><p>The <strong>input text files</strong> for the DrugProt track are plain-text, UTF8-encoded PubMed records in a tab-separated format with the following three columns:</p><ol><li>Article identifier (PMID, PubMed identifier)</li><li>Title of the article</li><li>Abstract of the article</li></ol><p> </p><p>DrugProt <strong>entity mention annotation files</strong> contain manually labeled mention annotations of chemical compounds and genes/proteins. Such files consist of tab-separated fields containing the following six columns:</p><ol><li>Article identifier (PMID)</li><li>Term number (for this record)</li><li>Type of entity mention (CHEMICAL, GENE-Y, GENE-N)</li><li>Start character offset of the entity mention</li><li>End character offset of the entity mention</li><li>Text string of the entity mention</li></ol><p>Each line contains one entity, and <i>each entity is uniquely identified by its PMID and the Term Number</i>. Besides, each annotation contains an annotation type, the start-offset -the index of the first character of the annotated span in the text-, the end-offset -the index of the first character after the annotated span- and the text spanned by the annotation.</p><p>Example DrugProt <i>training</i> entity mention annotations:</p><p>11808879 T1 GENE-Y 1860 1866 KIR6.2 11808879 T2 GENE-N 1993 2016 glutamate dehydrogenase 11808879 T3 GENE-Y 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p> </p><p>Example DrugProt <i>development</i> entity mention annotations (no distinction between GENE-Y and GENE-N):</p><p>11808879 T1 GENE 1860 1866 KIR6.2 11808879 T2 GENE 1993 2016 glutamate dehydrogenase 11808879 T3 GENE 2242 2253 glucokinase 23017395 T1 CHEMICAL 216 223 HMG-CoA 23017395 T2 CHEMICAL 258 261 EPA</p><p><br>DrugProt <strong>relation annotations</strong> are distributed as a file that contains the detailed chemical-protein relation annotations prepared for the DrugProt track. There are no relation annotations for the test+background set (the goal of the task is to predict them). It consists of tab-separated columns containing:</p><ol><li>Article identifier (PMID)</li><li>DrugProt relation</li><li>Interactor argument 1 (<i>of type CHEMICAL</i>)</li><li>Interactor argument 2 (<i>of type GENE</i>)</li></ol><p>Each line contains one relation, and <i>each relation is identified by the PMID, the relation type and the two related entities</i>. In the below example, to find the entities involved in the first relation, you must find the entities with Term Identifier T1 and T52 <i>within the PMID 12488248.</i></p><p>Example DrugProt relation annotations:</p><p>12488248 INHIBITOR Arg1:T1 Arg2:T52 12488248 INHIBITOR Arg1:T2 Arg2:T52 23220562 ACTIVATOR Arg1:T12 Arg2:T42 23220562 ACTIVATOR Arg1:T12 Arg2:T43 23220562 INDIRECT-DOWNREGULATOR Arg1:T1 Arg2:T14</p><p> </p><p>Please, cite:</p><p>@inproceedings{krallinger2017overview, title={Overview of the BioCreative VI chemical-protein interaction Track}, author={Krallinger, Martin and Rabal, Obdulia and Akhondi, Saber A and P{\'e}rez, Mart{\i}n P{\'e}rez and Santamar{\'\i}a, Jes{\'u}s and Rodr{\'\i}guez, Gael P{\'e}rez and others}, booktitle={Proceedings of the sixth BioCreative challenge evaluation workshop}, volume={1}, pages={141--146}, year={2017}}</p><p> </p><p><strong>Summary statistics:</strong></p><p>Training set Development set Documents 3500 750 Tokens 1001168 199620 Annotated Entities 89529 18858 Annotated Relations 17288 3765</p><p> </p><p>Annotated Entities:</p><p>Training Entities Development Entities CHEMICAL 46274 9853 GENE-Y [Normalizable] 28421 - GENE-N [Non-Normalizable] 14834 - Gene Total (N+Y) 43255 9005 Total 89529 18858</p><p> </p><p>Annotated Relations:</p><p>Training Relations Development Relations INDIRECT-DOWNREGULATOR 1330 332 INDIRECT-UPREGULATOR 1379 302 DIRECT-REGULATOR 2250 458 ACTIVATOR 1429 246 INHIBITOR 5392 1152 AGONIST 659 131 AGONIST-ACTIVATOR 29 10 AGONIST-INHIBITOR 13 2 ANTAGONIST 972 218 PRODUCT-OF 921 158 SUBSTRATE 2003 495 SUBSTRATE_PRODUCT-OF 25 3 PART-OF 886 258 Total 17288 3765</p><p> </p><p>For further information, please visit <a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/</a> or email us at krallinger.martin@gmail.com and antoniomiresc@gmail.com</p><p> </p><p><strong>Related resources:</strong></p><ul><li><a href="https://biocreative.bioinformatics.udel.edu/tasks/biocreative-vii/track-1/">Web</a></li><li><a href="https://github.com/tonifuc3m/drugprot-evaluation-library">Evaluation library</a></li><li><a href="https://codalab.lisn.upsaclay.fr/competitions/8293">Online evaluation (CodaLab)</a></li><li><a href="https://doi.org/10.5281/zenodo.4957137">Relation annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957576">Gene and protein annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.4957518">Chemicals and drugs annotation guidelines</a></li><li><a href="https://doi.org/10.5281/zenodo.7252201">DrugProt Silver Standard Knowledge Graph</a></li><li><a href="https://doi.org/10.5281/zenodo.5042178">FAQ</a></li><li><a href="https://doi.org/10.5281/zenodo.5119878">DrugProt Large Scale Additional SubTrack</a></li><li><a href="https://doi.org/10.5281/zenodo.5656991">DrugProt Large Scale document collection protocol</a></li><li><a href="https://doi.org/10.5281/zenodo.8246229">DrugProt Complete PubMed Knowledge Graph</a><br> </li></ul>
Corpus de textes sur la musique
<p>Le corpus est monolingue (français), couvre une période allant de 1600 à 2000 et il est spécialisé dans la musique.</p>
Spanish text corpus for NLP/linguistics research
<p>Spanish text-corpus extracted from Wikipedia, using the platform described on <em>Cadavid Rengifo, Héctor Fabio, and Jonatan Gómez Perdomo. "Web text corpus extraction system for linguistic tasks." Ingeniería e Investigación 29.3 (2009): 54-60, and the related master thesis <a href="https://www.researchgate.net/publication/313819658_Sistema_de_aprendizaje_no_supervisado_de_lenguajes_naturales_con_soporte_morfologico">available on ResearchGate</a>.</em></p> <p>rawdata.dat: raw outcome of the extraction process from Wikipedia.</p> <p>sentences.txt: sentences extracted from the raw data after cleaning/filtering.</p>
The Annotated Corpus of Classical Tibetan (ACTib), Part I - Segmented version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL.
<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.