Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
36
datasets available to search
ShareScore release 0.9.0
Dataset results
36 results for “text corpus”
The Annotated Corpus of Classical Tibetan (ACTib), Part II - POS-tagged version, based on the BDRC digitised text collection, tagged with the Memory-Based Tagger from TiMBL
<p>This corpus is a part-of-speech tagged version of</p> <p>Wallman, Jeff, Rowinski, Zach, Ngawang Trinley, Tomlinson, Chris, & Keutzer, Kurt. (2017). Collection of Tibetan etexts compiled by the Buddhist Digital Resource Center [Data set]. Zenodo. http://doi.org/10.5281/zenodo.821218</p> <p>using the training data of</p> <p>Hill, Nathan W., & Garrett, Edward. (2017). A part-of-speech (POS) tagged corpus of Classical Tibetan [Data set]. Zenodo. http://doi.org/10.5281/zenodo.574878</p> <p>Please note that the files are not post-processed or manually corrected and that a small number of files in the KarmaDelek directory were still annotated, although the original xml-input was corrupted already.</p> <p> </p> <p>using the memory based tagger of</p> <p>https://languagemachines.github.io/mbt/</p>
The morphologically glossed Rigveda - The Zurich annotation corpus revised and extended. Hosted by VedaWeb - Online Research Platform for Old Indic Texts.
<p>This file contains morphological and lexicographic annotations for the Rigveda. It was created in the DFG-funded research project Vedaweb and used as source data for the linguistic research platform <a href="https://vedaweb.uni-koeln.de">vedaweb.uni-koeln.de</a>.</p> <p>Prof. Dr. Paul Widmer and Dr. Salvatore Scarlata from the "Institut für Vergleichende Sprachwissenschaft" (Universität Zürich) provided the VedaWeb project a Filemaker file that was later transformed in Cologne into an Excel file. This data contained a version of the Rigveda by Prof. Dr. A. Lubotsky ("Indo-European Linguistics", Leiden University) that had been morphosytactically annotated over the course of more than 10 years at the University of Zurich. It also contained for each token, if available, a reference to an entry in Grassmann's dictionary for the Rigveda.</p> <p> </p> <p><strong>Modifications made by Jakob Halfmann and Natalie Korobzow to the data in 2020:</strong></p> <p>Disambiguation of the relevant categories, if unspecified in Zurich data, according to the Grassmann dictionary (updates from 6th edition partially included up to page 274):</p> <ul> <li>case, gender and number for nouns, pronouns (columns G–I)</li> <li>number, person, mood, tense and voice for verbs (columns I–M) up to line 109216</li> <li>case, gender, number, tense and voice for participles (columns G–I, L–M) up to line 109216</li> <li>absolutives are marked as Abs. in columns N and V</li> <li>Inconsistencies between the original file from Zurich and the Grassmann dictionary as well as internal inconsistencies in Grassmann are noted in column AE, whenever they were noticed.</li> <li>Zurich data was overwritten by conflicting Grassmann data in columns G–M but retained elsewhere.</li> <li>Verb classes according to Whitney (1885) and Jamison (1983) for class 10 in column Y, differences in root spelling between Whitney and Grassmann are noted in column Z. All potential verb classes provided by Whitney are given for every occurrence of the root.</li> <li>Local particles and verbal forms containing them are marked as LP in column AF.</li> <li>Comparatives and superlatives are marked as such in column X and desideratives as Des. in column Y.</li> </ul> <p> </p> <p><strong>Modifications made by Anna Fischer (data transformation, technical realisation) to the data:</strong></p> <p>New structure of data table for linguistic annotations with new column titles:</p> <ul> <li>A - "VERS_NR": renamed column (from "belege::stelleMMSSSRR")</li> <li>B - "PADA_NR": renamed column (from "belege::pada")</li> <li>C - "PADA_TEXT_LUBOTSKY": renamed column (from "belege::lubotskypada")</li> <li>D - "TOKEN_NR_VERS": renamed column (from "belege::wortnummer rc")</li> <li>E - "TOKEN_NR_PADA": renamed column (from "belege::wortnummer pada")</li> <li>F - "FORM": renamed column (from "belege::form")</li> <li>G - "KASUS": renamed column (from "belege::kasus")</li> <li>H - "GENUS": renamed column (from "belege::genus")</li> <li>I - "NUMERUS": renamed column (from "belege::numerus")</li> <li>J - "PERSON": renamed column (from "belege::person")</li> <li>K - "TEMPUS": moved and renamed column (from L "belege::tempus")</li> <li>L - "PRAESENSKLASSE": created new column for present stem class for each form</li> <li>M - "LEMMA_PRAESENSKLASSEN": created column for present stem classes of respective lemma: Moved and renamed column (from Y "formen::zusätzliche merkmale verb"), moved values "Abs." and "Inf." to column P "INFINIT", moved values "Prek." and "si-Ipv." to column N "MOOD", moved value "Des." to column Q "ABGELEITETE_KONJUGATION": moved value "se-Form" to column W "WEITERE_WERTE"</li> <li>N - "MODUS": moved and renamed column (from K "belege::modus")</li> <li>O - "DIATHESE": moved and renamed column (from M "belege::diathese")</li> <li>P - "INFINIT": created new column for infinite forms "Abs.", "Inf.", "Ptz.", "ta-Ptz.", "na-Ptz."</li> <li>Q - "ABGELEITETE_KONJUGATION": created new column for secondary conjugation "Des.", "Int.", "Kaus."</li> <li>R - "GRADUS": created new column for degree: "Comp.", "Sup."</li> <li>S - "LOKALPARTIKEL": moved and renamed column (from AF "LP")</li> <li>T - "LEMMA_ZÜRICH": moved and renamed column (from AA "lemmata klassisch::lemma")</li> <li>U - "LEMMA_ZÜRICH_LEMMATYP": moved and renamed column (from AB "lemmata klassisch::lemmatyp")</li> <li>V - "LEMMA_ZÜRICH_BEDEUTUNG": moved and renamed column (from AC "lemmata klassisch::bedeutung")</li> <li>W - "WEITERE_WERTE": created new column for all miscellaneous values: e.g. "Hyperchar.", "n-haltig", "se-Form"</li> <li>X - "KOMMENTAR": created new column merging former columns Z "formen::HELPformbestimmung", AD "lemmata klassisch::HELPbedeutung" and AE "anmerkungen abweichungen"</li> </ul> <p>Columns that were removed due to redundant information:</p> <ul> <li>"formen::zusätzliche merkmale nomen": values "superlative" And "comparative" were renamed "sup." and "comp." and moved to new column R "GRADUS", all other values were moved to new column for miscellaneous W "WEITERE_WERTE"</li> <li>"belege::belegbestimmung summe simpel": values "Ptz.", "ta-Ptz." and "na-Ptz." were moved to new new column P "INFINIT"</li> <li>"belege::kasus bestof"</li> <li>"belege::genus bestof"</li> <li>"belege::numerus bestof"</li> <li>"belege::person bestof"</li> <li>"belege::modus bestof"</li> <li>"belege::tempus bestof"</li> <li>"belege::diathese bestof"</li> <li>"belege::belegbestimmung bestof summe sophistiziert"</li> </ul> <p> </p> <p><strong>Revisions and additions made by Antje Casaretto to the data in 2023:</strong></p> <ul> <li>F-T: - revision and correction (wherever necessary) of all annotations (books 1-7)</li> <li>G,H,I - disambiguation of case forms, reg. pronouns and nominal forms, if unspecified in Zurich data (books 1-7)</li> <li>L - disambiguation of present stem classes (book 7 and book 1 up to line 21050 vers 01.125.01)</li> <li>M - disambiguation of denominal verbs from primary verbs of the 10th class (books 1-10)</li> <li>N - disambiguation of precative and optative forms wherever possible (books 1-7)</li> <li>Q - new annotations for "Int." (intensives) and "Kaus." (causatives) (books 1-7)</li> </ul> <p> </p> <p dir="ltr"><strong>Revisions and additions made by Antje Casaretto to the data in 2024 with support in data modeling and automation by Anna Fischer:</strong></p> <ul> <li>F-T: revision and correction (wherever necessary) of all annotations (books 8-10)</li> <li>G,H,I: disambiguation of case and gender forms in nominal and pronominal forms, if unspecified in Zurich data (books 8-10)</li> <li>L, M: disambiguation of present stem classes (books 1-10)</li> <li>N: disambiguation of precative and optative forms wherever possible (books 8-10)</li> <li>P: new annotations for "Gdv." (gerundives)</li> <li>Q: new annotations for "Den." (denominatives) (books 1-10) and further annotations of “Kaus.” (causatives) and “Int.” (intensives) (books 8-10)</li> <li>T: revision of lemmatization (books 1-10)</li> <li>V: update of meanings according to revised lemmatization; minimal revision</li> <li>W: revised annotation of ending -se (“se-Form”) (books 1-10); no systematic revision</li> <li>X: no systematic revision</li> <li>A-U: general revision of formal inconsistencies and typing errors (book 1-10)</li> </ul> <p> </p> <p dir="ltr"><strong>Revisions made by Natalie Korobzov and Pascal Coenen to the data in 2024 with computational support by Anna Fischer:</strong></p> <ul> <li>Y - "LEMMA_GRASSMANN_ID": new column for references to Grassmann dictionary (books 1-10) and revision of Grassmann references</li> </ul>
PT2vec - A Portuguese text corpus created from online newspapers
<p>A Portuguese text corpus created from online newspapers with a total of 394,825,480 tokens and 33,089,734 sentences.</p> <p>If you use this repository, please cite this paper:</p> <p>Pinto JP, Viana P, Teixeira I, Andrade M. 2022. Improving word embeddings in Portuguese: increasing accuracy while reducing the size of the corpus. PeerJ Computer Science 8:e964 <a href="https://doi.org/10.7717/peerj-cs.964">https://doi.org/10.7717/peerj-cs.964</a></p>
Webis Text Reuse Corpus 2012
<p>The Webis Text Reuse Corpus 2012 (Webis-TRC-12) compiles manually written documents obtained from a completely controlled, yet representative environment that emulates the web. Each document in the corpus is about one of the 150 topics used at the TREC Web Tracks 2009–2011, thus forming a strong connection with existing evaluation efforts. Writers, hired at the crowdsourcing platform oDesk, had to retrieve sources for a given topic and to reuse text from what they found. Part of the corpus are detailed interaction logs that consistently cover the search for sources as well as the creation of documents. This will allow for in-depth analyses of how text is composed if a writer is at liberty to reuse texts from a third party.</p> <p> </p>
Hypernyms extracted from a large text corpus using Hearst lexical-syntactic patterns
<p><br> The list of hyponym-hypernym pairs was obtained by applying lexical-syntactic patterns described in Hearst (1992) on the corpus prepared by Panchenko et al. (2016). This corpus is a concatenation of the English Wikipedia (2016 dump), Gigaword, ukWaC and English news corpora from the Leipzig Corpora Collection. The lexical-syntactic patterns proposed by Marti Hearst (1992) and further extended and implemented in the form of FSTs by Panchenko et al. (2012) for extracting (noisy) hyponym-hypernym pairs are as follows -- (i) such NP as NP, NP[,] and/or NP; (ii) NP such as NP, NP[,] and/or NP; (iii) NP, NP [,] or other NP; (iv) NP, NP [,] and other NP; (v) NP, including NP, NP [,] and/or NP; (vi) NP, especially NP, NP [,] and/or NP. Pattern extraction on the corpus yields a list of 27.6 million hyponym-hypernym pairs along with the frequency of their occurrence in the corpus. </p>
Arabic Raw Text Corpus (unfiltered, extracted from the Clueweb WARC Files)
<p>The corpus is created from 29,192,662 ClueWeb html. Initially archived in WARC files. Each text file contains in average 30,000 different webpages. The size of the corpus is 18,482,719 terms. </p>
Bilateral Cultural Agreements, 1935-1972: a Curated International Text Corpus
<p>This is a curated text corpus composed of the texts of bilateral general cultural agreements signed between 1935 and 1972 and deposited with either the League of Nations Treaty Service (LTS) or United Nations Treaty Service (UNTS) and published by those organizations. All texts are in English, either as an original language or in the translation provided by these treaty services.</p> <p>The complete list of bilateral treaties included in this corpus is <a href="https://zenodo.org/record/7361913">available here</a>. The source documents can be found online in PDF form through the <a href="https://treaties.un.org/Pages/Content.aspx?path=DB/UNTS/pageIntro_en.xml">UNTS website</a>. We were able to identify this selection of agreements thanks to the electronic World Treaty Index (<a href="http://worldtreatyindex.com">worldtreatyindex.com</a>), or eWTI [1]. An edited version of the portions of the eWTI related to cultural agreements is available at: <a href="https://zenodo.org/record/5159745">zenodo.org/record/5159745</a>. Discussion of the principles of selection behind this corpus, including a definition of "general cultural agreements," can be found in this article (<a href="https://doi.org/10.1080/07075332.2022.2048051">B. G. Martin, "The Rise of the Cultural Treaty," <em>International History Review</em>, 2022</a>); as well as in <a href="https://github.com/benjamingmartin/the_culture_of_international_relations/wiki/eWTI-Notes:-Selection-Decisions-Log-2020,-2021">this log</a>, which documents the decision-making process. The 464 agreements collected in this corpus represent almost exactly half of the total number of general cultural agreements signed between January 1935 and December 1972 (49.88% of 931). The remainder were not deposited with the League of Nations or the UN; some were published in national treaty collections.</p> <p>I assembled this corpus, together with Roger Mähler and Andreas Marklund of <a href="https://www.umu.se/en/humlab/">Humlab, Umeå University</a>, between 2018 and 2020, as part of the research project "The Culture of International Society," which was funded by a generous grant from Riksbankens Jubileumsfond (the Swedish Foundation for Humanities and Social Sciences, P16-0900:1). Sincere thanks to Roger and Andreas, as well as to Milada Jamroskovic, who conducted many hours of painstaking editing.</p> <p>Read about the project at: <a href="http://www.benjamingmartin.com/cultintsoc">benjamingmartin.com/cultintsoc</a>.</p> <p>Code related to the project, including text analysis tools designed for this corpus, is available at <a href="https://github.com/benjamingmartin/the_culture_of_international_relations">the project's GitHub page</a>.</p> <p>Please contact me with questions or suggestions: benjamin [dot] martin [at] idehist.uu.se</p> <p> </p> <p>[1] Paul Poast, Michael James Bommarito and Daniel Martin Katz, ‘The Electronic World Treaty Index: Collecting the Population of International Agreements in the 20th Century’, SSRN Scholarly Paper (Rochester, NY: Social Science Research Network, February 15, 2010), online at: <a href="http://papers.ssrn.com/abstract=2652760">http://papers.ssrn.com/abstract=2652760</a>.</p>
Swedish Diachronic Corpus - User-generated Data - Blog text
<p>The Swedish Diachronic Corpus (https://www2.lingfil.uu.se/person/pettersson/svediakorp/) is a project funded by Swe-Clarin (<a href="https://sweclarin.se/eng">https://sweclarin.se/eng</a>). The purpose of the project is to provide a corpus of texts covering the time period from Old Swedish to present day, with a wide variety of text types and freely available for download and search. The texts are provided in a plain text format and in a uniform CoNLL format, with placeholders for linguistic annotation. </p> <p>The dataset provided here is the blog section of the User-generated Data in the Swedish Diachronic Corpus. For other datasets within the corpus, see further the corpus website: https://www2.lingfil.uu.se/person/pettersson/svediakorp/. </p> <p>The project members are Eva Pettersson (Uppsala University) and Lars Borin (University of Gothenburg). For questions or comments, or if you are aware of any corpus resource that could be included in the Swedish Diachronic Corpus, don't hesitate to contact us!<br> <br> Eva Pettersson Department of Linguistics and Philology, Uppsala University eva.pettersson@lingfil.uu.se<br> Lars Borin Department of Swedish, University of Gothenburg lars.borin@svenska.gu.se</p>
Corpus of historic Basque texts
<p>A corpus of historic Basque texts analyzed for this study, classified according to textual genre and dialect, with source references.</p>
Komnzo text corpus
<p>This dataset contains a zip-file of the most up to date version of the Komnzo text corpus. The original footage files can be found here at Zenodo or under: </p> <ul> <li>Döhler, C. 2010-2015. <em>DoBeS Documentation: Nen and Komnzo, two languages of Southern New Guinea</em>. Nijmegen: The Language Archive. URL: <a href="https://hdl.handle.net/1839/00-0000-0000-0017-B0AC-C">https://hdl.handle.net/1839/00-0000-0000-0017-B0AC-C</a></li> </ul> <p>The material was recorded by Christian Döhler as part of a language documentation project for his PhD. The project was located at the <a href="http://chl.anu.edu.au/">School of Culture, History and Language</a> at the <a href="http://anu.edu.au/">Australian National University, Canberra</a>. For the most part it was funded by the <a href="http://dobes.mpi.nl/">DOBES project</a> of the <a href="https://www.volkswagenstiftung.de/en/foundation">Volkswagen Foundation</a>.</p>
BHAAV (भाव) - A Text Corpus for Emotion Analysis from Hindi Stories
<p>The first and largest Hindi text corpus, named BHAAV (भाव), which means emotions in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus consists of 20,304 sentences collected from 230 different short stories spanning across 18 genres such as प्रेरणादायक (Inspirational) and रहस्यमयी (Mystery). Each sentence has been annotated into one of the five emotion categories anger, joy, suspense, sad, and neutral) by three native Hindi speakers with at least ten years of formal education in Hindi.</p>
Appendix: corpus of historic (18th-21st century) PCM texts
<p>Corpus of historic (18th-21st) PCM texts analyzed for the study.</p>
GraSCCo_PII_V2 - Graz Synthetic Clinical text Corpus with PII Annotations
<p><strong>GraSCCo_PII_V2 - Graz Synthetic Clinical text Corpus with Personally Identifiable Informations</strong></p> <p>Additional source of:</p> <ul> <li>Lohr C, Faller J, Riedel A, Nguyen HM, Wolfien M, Hofenbitzer J, Modersohn L, Romberg J, Prasser F, Omeirat J, Wen Y, Galusch O, Hahn U, Seiferling M, Dieterich C, Klügl P, Matthies F, Kind J, Boeker M, Löffler M, Meineke F. <strong>GeMTeX's De-Identification in Action: Lessons Learned & Devil's Details</strong>. Stud Health Technol Inform. 2025 Sep 3;331:274-282. doi: 10.3233/SHTI251406. PMID: 40899551. (<a href="https://pubmed.ncbi.nlm.nih.gov/40899551/">https://pubmed.ncbi.nlm.nih.gov/40899551/</a>)</li> </ul> <p>GraSCCo is a collection of artificially generated semi-structured and unstructured German-language clinical summaries. These summaries are formulated as letters from the hospital to the patient's GP after in-patient or out-patient care. Details:</p> <ul> <li>Stefan Schulz. (2022). GraSCCo (Version v1) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.6539131" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.6539131</a></li> <li>Modersohn L, Schulz S, Lohr C, Hahn U. GRASCCO - The First Publicly Shareable, Multiply-Alienated German Clinical Text Corpus. <em>Stud Health Technol Inform</em>. 2022;296:66-72. doi:10.3233/SHTI220805</li> </ul> <p><strong>First Version - GraSSCo</strong> <strong>with</strong> PHI (now named as <strong>PII</strong>) annotations as external source of</p> <ul> <li><strong>Lohr C, Matthies F, Faller J, et al. De-Identifying GRASCCO - A Pilot Study for the De-Identification of the German Medical Text Project (GeMTeX) Corpus. <em>Stud Health Technol Inform</em>. 2024;317:171-179. doi:10.3233/SHTI240853 (<a href="https://pubmed.ncbi.nlm.nih.gov/39234720/">https://pubmed.ncbi.nlm.nih.gov/39234720/</a>)<br></strong></li> </ul> <p>This repository contains the annotations in XMI and JSON exports created with the INCEpTION annotation platform (<a href="https://inception-project.github.io/">https://inception-project.github.io/</a>), also the annotation guideline document, TypeSystem.xml and layer.json (needed for import in INCEpTION), see also <a href="https://github.com/dkpro/dkpro-cassis">https://github.com/dkpro/dkpro-cassis</a>.</p>
James Joyce's Complete Correspondence Text Corpus
<p>Text corpus of James Joyce's complete correspondence, including all the extant published and unpublished letters. The corpus was created within my PhD project titled "James Joyce’s Correspondence: A Scholarly Digital Edition and a Distant Reading Analysis", and is accompanied by an annotated digital edition of Joyce's unpublished letters that can be found here <a href="https://jamesjoycecorrespondence.org/" target="_blank" rel="noopener">https://jamesjoycecorrespondence.org/</a>. The corpus counts a total of 3618 files, containing a variety of letters, postcards, notes, calling cards, and telegrams Joyce sent to his correspondents over the course of his life. This upload is accompanied by the <a href="../records/10987518">James Joyce's Complete Correspondence Metadata</a>.</p>
A collection of text embeddings of the arXiv corpus by title and abstract
<p>A popular online repository of <a href="http://arxiv.org">arXiv</a> is home to numerous preprints in many scientific domains. Other than playing a role of disseminating up-to-date knowledge in pertaining domains, arXiv is an interesting complex system by itself from text analytics point of view. In this repository, we provide a collection of text embedding outputs for (almost) all papers' from the arXiv corpus by their titles and abstracts in order to provide multi-faceted characteristics of scientific knowledge.</p>
TeCoPhy: A Text Corpus of German Physics Texts
<p>TeCoPhy is a Text Corpus of German Physics Texts. Most of the texts are taken from textbooks at school and university level, but other sources were included as well. The corpus is a collection of sentences. From each book, at most 14% of the text is included in the corpus. The corpus consists of 236,278 sentences with 5,32 Million tokens collected from 223 different sources.</p> <p>The distribution consists of two files: an XML file with metadata on the sources and a file with the sentences taken from those sources.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.