Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

56 results for “corpora”

Learn how ShareScore rates datasets ↗
zenodo36/100

EuroparlExtract - Directional Parallel Corpora Extracted from the European Parliament Proceedings Parallel Corpus

<p>This dataset contains directional parallel corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. EuroparlExtract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>

opencc-by-4.0Nov 2017View details →
zenodo36/100

EuroparlExtract - Comparable Corpora Extracted from the European Parliament Proceedings Parallel Corpus

<p>This dataset contains comparable translational corpora extracted from the European Parliament Proceedings Corpus (Europarl) v7 created by Philipp Koehn (see http://www.statmt.org/europarl/). For the extraction, the EuroparlExtract corpus processing toolkit by Michael Ustszewski (2017) was used. Europarl Extract is freely available under the MIT License (see https://github.com/mustaszewski/europarl-extract).</p>

opencc-by-4.0Nov 2017View details →
zenodo36/100

Metadata for Historical Corpora. Realization of the Metamodel for Corpus Metadata with the help of TEI Customization

<p>TEI ODD Customization for the documentation of historical corpora:</p> <p>The TEI ODD customizations map the Metamodel for Corpus Metadata (MCM) to a TEI p5 header structure for each of the objects of the classes &#39;Corpus&#39;, &#39;Document&#39; and &#39;Preparation&#39;. The MCM is realized with a subset of the TEI p5 guidelines.</p> <p>Each ODD contains further information and explanations regarding the MCM and the customization of the TEI. Additionally, for each ODD, an HTML documentation is provided.</p>

opencc-by-4.0Feb 2017View details →
zenodo36/100

UNIC Example implementations of the metadata schema for interpreting corpora

<p>Four example implementations of the metadata schema for interpreting corpora, namely the European Parliament Interpreting Corpus v2.0 (Russo et al. 2012), the interpreted subcorpus of the <em>Dolmetschen im Krankenhaus</em> (&lsquo;Interpreting in Hospitals&rsquo;) corpus v1.1&nbsp;(B&uuml;hrig et al. 2012), the Speech Corpus of Interpreted Premier Press Conferences v1.0 (Liu 2023), and the Belgian Covid Sign language corpus v1.1 (Vandeghinste et al. 2022).&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

UNIC Example implementations of the metadata schema for interpreting corpora

<div> <p>Four example implementations of the metadata schema for interpreting corpora, namely the European Parliament Interpreting Corpus v2.0 (Russo et al. 2012), the interpreted subcorpus of the&nbsp;<em>Dolmetschen im Krankenhaus</em> (&lsquo;Interpreting in Hospitals&rsquo;) corpus v1.1&nbsp;(B&uuml;hrig et al. 2012), the Speech Corpus of Interpreted Premier Press Conferences v1.0 (Liu 2023), and the Belgian Covid Sign language corpus v1.1 (Vandeghinste et al. 2022).&nbsp;</p> </div>

opencc-by-4.0Aug 2024View details →
zenodo36/100

Appendix B: Language Corpora Available for Text Mining

<p>Appendix B is associated with <em><strong>Chapter 3: Text Pre-Processing</strong></em> of the book -- Manika Lamba and Margam Madhusudhan (2021) Text Mining for Information Professionals: An Uncharted Territory, SpringerNature.</p>

openother-openJul 2021View details →
zenodo36/100

PatchFuzz evaluation corpora

<p>This package contains several compressed corpora used to evaluate PatchFuzz. These corpora are originally derived from the ones archived by the `magma` project, and minimized using `afl-cmin`.</p>

opencc-by-4.0Mar 2023View details →
zenodo36/100

Corpora Scraped from /r/youtubers and /r/uberdrivers

<p>To investigate the status of platform workers&rsquo; deliberations, two corpuses were produced from &ldquo;conversation data from a large online community dedicated to&rdquo; the platforms of interest, YouTube and Uber <a href="#_ftn1">[1]</a>. Following Bucher et al. (2021), both corpuses were built to supplement existing literature. Instead of relying solely on others&rsquo; surveys of platform workers&rsquo; experiences, the research question dealing with the articulation of platform workers&rsquo; voices was probed directly. These data, gathered from &ldquo;/r/youtubers&rdquo; and &ldquo;/r/uberdrivers&rdquo; on Reddit, thus sought to reveal the content of platform workers&rsquo; deliberations and to what extent they evidence inquiry and continuous learning in the pragmatist sense. The top 1000 &ldquo;discussion threads from the online community&rdquo; were scraped. These included 1000 initial posts and 5 comments for each post, producing a total of 6000 total posts from each subreddit.</p> <p>&nbsp;</p> <p><a href="#_ftnref1">[1]</a> Bucher, Schou, and Waldkirch, &ldquo;Pacifying the Algorithm &ndash; Anticipatory Compliance in the Face of Algorithmic Management in the Gig Economy.&rdquo;</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

Corpora, Corpus Analysis, Resources and Tools

<p>During the fourth Project Presentation Session on <strong>Tuesday 24.07.2018</strong> the following 3 projects were presented:</p> <ul> <li><strong>Yu-Hua Chen</strong> (University of Nottingham Ningbo China, China): &quot;Beyond Borders, Beyond Words: Issues &amp; Challenges in Developing An Open-Access Multimodal Corpus of L2 Academic English from An EMI University in China&quot;</li> <li><strong>Theodora Konach</strong> (Jagiellonian University,Krakow, Poland): &quot;Towards an Ethical Framework for the Digitalisation of Intangible Cultural Heritage in Museums - Digital Humanities in the Comparative Intellectual Property Law&quot;</li> <li><strong>Marco Passarotti</strong> (Universit&agrave; Cattolica del Sacro Cuore, Milan, Italy): &quot;What&#39;s Going On in Milan? A Practical Introduction to Resources and Tools for Latin at the CIRCSE Research Centre&quot;</li> </ul>

opencc-by-4.0Jul 2018View details →
zenodo32/100

General Terminology Induction: Two Corpora of Experimental Ontologies

<p>Two corpora of ontologies that contain some data and hence can be used to evaluate ontology learning algorithms</p>

opencc-by-4.0Feb 2017View details →
zenodo32/100

Authentic human translations corpora EN-MNE, EN-THA

<p>Data for the research has been gathered from two authentic human translations corpora <a href="https://www.clarin.si/repository/xmlui/handle/11356/1176?show=full">EN-MNE</a>, and <a href="https://opus.nlpl.eu/">EN-THA</a>. Pairs of sentences were selected according to the length (100-150 characters) and processed according to the methodology explained within the research.</p>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Interview Corpora for the Study of Multilingual Repertoires in South-South Migration Dynamics I: Haitians in Chapecó (SC, Brazil)

<p>Corpus of 19 multilingual interviews with Haitian migrants in Chapec&oacute; (Santa Catarina, Brazil), conducted in various languages: in order from the most to the least documented in the interviews, Portuguese, French, Spanish, and Haitian Creole. This corpus is part of broader research on the evolution of multilingual repertoires in South-South migration dynamics. The interviews were conducted in March 2023 by the author in collaboration with Leonie Ette (University of Augsburg) and the two coordinators of the research group <em>Atlas das L&iacute;nguas em Contato na Fronteira</em>, Professors Cristiane Horst and Marcelo Krug (UFFS, Campus Chapec&oacute;).</p>

opencc-by-4.0Mar 2024View details →
zenodo32/100

Predictability of linguistic complexity: samples of the corpora

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2024View details →
zenodo32/100

iRead4Skills Dataset 2: annotated corpora by level of complexity for FR, PT and SP

<div> <table> <tbody> <tr> <td> <p>&nbsp;</p> <p><span>The <em>Dataset 2: annotated corpora by level of complexity for FR, PT and SP </em>is a collection of texts categorized by complexity level and annotated for complexity features, presented in Excel format (.xlsx). These corpora were compiled and annotated under the scope of the project iRead4Skills &ndash; Intelligent Reading Improvement System for Fundamental and Transversal Skills Development, funded by the European Commission (grant number: 1010094837). The project aims to enhance reading skills within the adult population by creating an intelligent system that assesses text complexity and recommends suitable reading materials to adults with low literacy skills, contributing to reducing skills gaps and facilitating access to information and culture (</span><span><a href="https://iread4skills.com/">https://iread4skills.com</a></span><span>).</span></p> <p><span>This dataset is the result of specifically devised classification and annotation tasks, in which selected texts were organized and distributed to trainers in Adult Learning (AL) and Vocational Education Training (VET) Centres, as well as to adult students in AL and VET centres. This task was conducted via the Qualtrics platform.</span></p> <p><span>The <em>Dataset 2: annotated corpora by level of complexity for FR, PT and SP</em> is derived from <em>the iRead4Skills Dataset 1: corpora by level of complexity for FR, PT and SP</em> ( </span><span><a href="https://doi.org/10.5281/zenodo.10055909">https://doi.org/10.5281/zenodo.10055909</a></span><span>), which comprises written texts of various genres and complexity levels. From this collection, a sample of texts was selected for classification and annotation. This classification and annotation task aimed to provide additional data and test sets for the complexity analysis systems for the three languages of the project: French, Portuguese, and Spanish. The sample texts in each of the language corpora were selected taking into account the diversity of topics/domains, genres, and the reading preferences of the target audience of the iRead4Skills project. This percentage amounted to the total of 462 texts per language, which were divided by level of complexity, resulting in the following distribution:</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 140 Very Easy texts</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 140 Easy texts</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 140 Plain texts</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 42 More Complex texts.</span></p> <p><span>Trainers and students were asked to classify the texts according to the complexity levels of the project, here informally defined as:</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span><em><span>Very Easy</span></em><span> (everyone can understand the text or most of the text).</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span><em><span>Easy</span></em><span> (a person with less than the 9th year of schooling can understand the text or most of the text)</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span><em><span>Plain</span></em><span> (a person with the 9th year of schooling can understand the text the first time he/she reads it)</span></p> <p><span>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; </span><em><span>More complex</span></em><span> (a person with the 9th year of schooling cannot understand the text the first time he/she reads it).</span></p> <p><span>Annotators were also asked to mark the parts of the texts considered complex according to various type of features, at word-level and at sentence-level (e.g., word order, sentence composition, etc.), The full details regarding the students and the trainers&rsquo; tasks, data qualitative and quantitative description and inter-annotator agreement are described here: </span><span><a href="https://zenodo.org/records/14653180">https://zenodo.org/records/14653180</a></span></p> <p><span>The results are here presented in Excel format. For each language, and for each group (trainers and students), two pairs of files exist &ndash; the annotation and the classification files &ndash; resulting in four files per language and twelve files, in total.</span></p> <p><span>In all files,</span> <span>the data is organized as a matrix, with each row representing an &lsquo;answer&rsquo; from a particular participant, and the columns</span> <span>containing various details about that specific input, as shown below:</span></p> <table> <tbody> <tr> <td> <p><strong><span>Column name</span></strong></p> </td> <td> <p><strong><span>Data</span></strong></p> </td> </tr> <tr> <td> <p><strong><em><span>Annotator's ID</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>The randomly generated ID code for each annotator, together with information on the dataset assigned to them.&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Progress</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Information on the completion of the task (for each text).&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Duration (seconds)</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Time used in the completion of the task (for each text).&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>File Name</span></em></strong><span>&nbsp;</span></p> <p><strong><em><span>N1 = Very Easy</span></em></strong><span>&nbsp;</span></p> <p><strong><em><span>N2 = Easy</span></em></strong><span>&nbsp;</span></p> <p><strong><em><span>N3 = Plain</span></em></strong><span>&nbsp;</span></p> <p><strong><em><span>N4=More Complex</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>File internal identification, providing its iRead4Skills classification.&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Text</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>The content of the file, i.e. the text itself.&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Annotated Level</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Level assigned by the annotator (trainer).&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Proficiency SubLevel</span></em></strong><span>&nbsp;</span></p> <p><strong><em><span>(Likert Scale - 1 to 5)</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>SubLevel assigned by the annotator (trainer) for FR data.&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Corresponding CEFR Level</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>CEFR level closest to the iRead4Skills&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Additional Info</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Observations made by the trainers/students&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Annotated Term</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Word or set of words selected for annotation&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Term Label</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Annotation assigned to the Annotated Term (difficult word, word order, etc.)&nbsp;&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Term Index</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Position of the annotated term in the text&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Annotator's Proficiency Level</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Level of AL/VET of the student&nbsp;</span></p> </td> </tr> <tr> <td> <p><strong><em><span>Text adequate for user</span></em></strong><span>&nbsp;</span></p> </td> <td> <p><span>Validation of the text by the students&nbsp;</span></p> </td> </tr> </tbody> </table> <p><span>&nbsp;</span></p> <p><span>The content of the column &ldquo;File Name&rdquo; is color-coded, where a green shade alludes to a text with a lower level of complexity and a red one alludes to one with a higher level of complexity. </span></p> <p><span>The complete datasets are available under creative CC BY-NC-ND 4.0.</span></p> <p>&nbsp;</p> </td> </tr> </tbody> </table> </div>

restrictedcc-by-4.0Jul 2024View details →
zenodo32/100

VladislavKaryukin/kk_en_corpora: The Kazakh - English parallel corpora

<p>The full corpora of 380 thousand parallel sentences</p>

openother-openSep 2022View details →
zenodo32/100

Multilingual Medical Corpora

<p>The amount of digital data derived from healthcare processes have increased tremendously in the last years. This applies especially to unstructured data, which are often hard to analyze due to the lack of available tools to process and extract information. Natural language processing is often used in medicine, but the majority of tools used by researchers are developed primarily for the English language. For developing and testing natural language processing methods, it is important to have a suitable corpus, specific to the medical domain that covers the intended target language. To improve the potential of natural language processing research, we developed tools to derive language specific medical corpora from publicly available text sources. n order to extract medicine-specific unstructured text data, openly available pub-lications from biomedical journals were used in a four-step process:(1) medical journal databases were scraped to download the articles,(2) the articles were parsed and consolidated into a single repository,(3) the content of the repository was de-scribed, and (4) the text data and the codes were released. In total, 93 969 articles were retrieved, with a word count of 83 868 501 in three different languages (German, English, and Spanish) from two medical journal databases Our results show that unstructured text data extraction from openly available medical journal databases for the construction of unified corpora of medical text data can be achieved through web scraping techniques.</p>

opencc-by-4.0Sep 2019View details →
zenodo32/100

PoeTree. Poetry Corpora in Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Russian, Slovenian, and Spanish

<p>PoeTree is a dataset comprising nearly 335,000 poems / 90,000,000 tokens in 11 languages (Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Spanish, Slovenian, and Russian). Each corpus has been deduplicated, enriched with Universal Dependencies, provided with additional metadata and converted into a unified JSON structure (schema available at&nbsp;<a href="https://versologie.cz/poetree/json-schema">https://versologie.cz/poetree/json-schema</a>).</p> <ul> <li>cs (~80k poems) <ul> <li>derived from <a href="https://github.com/versotym/corpusCzechVerse" target="_blank" rel="noopener">Corpus of Czech Verse</a></li> </ul> </li> <li>de (~74k poems) <ul> <li>derived from <a href="https://metricalizer.de/" target="_blank" rel="noopener">Metricalizer</a> and <a href="https://github.com/tnhaider/DLK" target="_blank" rel="noopener">Deutsches Lyrik Korpus</a></li> </ul> </li> <li>en (~40k poems) <ul> <li>based on texts from <a href="https://gutenberg.org/" target="_blank" rel="noopener">Project Gutenberg</a></li> </ul> </li> <li>es (~9k poems) <ul> <li>derived from <a href="https://github.com/bncolorado/CorpusSonetosSigloDeOro" target="_blank" rel="noopener">Corpus of Spanish Golden-Age Sonnets</a> and <a href="https://github.com/pruizf/disco" target="_blank" rel="noopener">Diachronic Spanish Sonnet Corpus</a></li> </ul> </li> <li>fr (~18k poems) <ul> <li>derived from <a href="https://crisco4.unicaen.fr/verlaine" target="_blank" rel="noopener">Malherbə</a></li> </ul> </li> <li>hu (~13k poems) <ul> <li>derived from <a href="https://github.com/ELTE-DH/poetry-corpus/" target="_blank" rel="noopener">ELTE Poetry Corpus</a></li> </ul> </li> <li>it (~40k poems) <ul> <li>derived from <a href="http://www.bibliotecaitaliana.it/" target="_blank" rel="noopener">Biblioteca Italiana</a></li> </ul> </li> <li>no (~3k poems) <ul> <li>derived from <a href="https://github.com/norn-uio/norn-poems" target="_blank" rel="noopener">NORN Poems</a></li> </ul> </li> <li>pt (~5k poems) <ul> <li>derived from <a href="https://github.com/adiel-mittmann/poemas" target="_blank" rel="noopener">Poemas</a></li> </ul> </li> <li>ru (~45k poems) <ul> <li>derived from <a href="https://ruscorpora.ru/en/" target="_blank" rel="noopener">Corpus of Russian Poetry</a></li> </ul> </li> <li>sl (~5k poem)<br> <ul> <li>based on texts from <a href="https://en.wikisource.org/" target="_blank" rel="noopener">wikisource</a></li> </ul> </li> </ul> <p><em>new in v. 1.0.0:</em></p> <ul> <li><em>PoeTree.no added</em></li> <li><em>PoeTree.(cs,de,en,fr,hu,it,ru,sl) enriched with geolocation mentions</em></li> <li><em>Updated and corrected metadata in PoeTree.(de,en,es,ru)</em></li> <li><em>Multiple text corrections in PoeTree.ru</em></li> </ul>

openApr 2024View details →
zenodo28/100

ORAL CORPORA FOR BILINGUAL AND PLURILINGUAL CONTEXTS

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →
zenodo28/100

Compilation of Large Spanish Unannotated Corpora

<p>Compilation of open Large Spanish Unannotated&nbsp;Corpora with some post processing.</p> <p>More details in:&nbsp;https://github.com/josecannete/spanish-corpora</p>

openother-pdMay 2019View details →
zenodo28/100

Sense-Tagged corpora for Word Sense Disambiguation in several languages

<p>Copies of Sense-Tagged corpora for Word Sense Disambiguation in several languages</p> <p>English, French, Spanish, Russian</p> <p>&nbsp;</p>

opencc-byOct 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record