Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

56 results for “corpora”

Learn how ShareScore rates datasets ↗
zenodo48/100

Cross-language corpora of privacy policies

<p>The dataset consists of three different privacy policy corpora (in English and Italian) composed of 81 unique privacy policy texts spanning the period 2018-2021. This dataset makes available an example of three corpora of privacy policies. The first corpus is the English-language corpus, the original used in the study by Tang et al. [2]. The other two are cross-language corpora built (one, the source corpus, in English, and the other, the replication corpus, in Italian, which is the language of a potential replication study) from the first corpus.</p> <p>The policies were collected from:</p> <ol> <li>the Alexa top 10 Italy and U.S. websites rank;</li> <li>the Play Store apps rank in the &quot;most profitable games&quot; category of the Play Store for Italy and the U.S.</li> </ol> <p>We manually analyzed the Alexa top 10 Italy websites as of November 2021. Analogously, we analyzed selected apps that, in the same period, had ranked better in the &quot;most profitable games&quot; category of the Play Store for Italy.</p> <p>All the privacy policies are ANSI-encoded text files and have been manually read and verified.<br> The dataset is helpful as a starting point for building comparable cross-language privacy policies corpora. The availability of these comparable cross-language privacy policies corpora helps replicate studies in different languages.&nbsp;<br> Details on the methodology can be found in the accompanying paper.</p> <p>The available files are as follows:</p> <ul> <li><strong>policies-texts.zip</strong> --&gt;&nbsp;contains a directory of text files with the policy texts. File names are the SHA1 hashes of the policy text.</li> <li><strong>policy-metadata.csv</strong> --&gt;&nbsp;Contains a CSV file&nbsp;with the metadata&nbsp;for each privacy policy.</li> </ul> <p>This dataset is the original dataset used in the publication [1]. The original English U.S. corpus is described in the publication [2].</p> <p>[1] F. Ciclosi, S. Vidor and F. Massacci. &quot;Building cross-language corpora for human&nbsp;understanding of privacy policies.&quot; Workshop on Digital Sovereignty in Cyber Security: New Challenges in Future Vision. Communications in Computer and Information Science. Springer International Publishing, 2023, In press.</p> <p>[2] J. Tang, H. Shoemaker, A. Lerner, and E. Birrell. Defining Privacy: How Users&nbsp;Interpret Technical Terms in Privacy Policies. Proceedings on Privacy Enhancing&nbsp;Technologies, 3:70&ndash;94, 2021.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

The Curated Courier: Digital Text Corpora from the UNESCO Courier (1948–2020)

<p>Founded in 1948 as the official magazine of the United Nations Educational, Scientific and Cultural Organization, <i>The UNESCO Courier</i> represents an extraordinary resource for research on global themes in the humanities. The complete <a href="https://en.unesco.org/courier/archives">archive of the magazine</a> is available in PDF form through UNESCO.&nbsp;These files make it possible for users anywhere to read individual issues, but it does not allow for full-text searching, much less any of the computational text analysis methods that have recently made important advances in humanities research.</p><p>The Curated Courier 1.0 is a package of digital text corpora, text analysis tools, and supplementary materials that makes the complete archive of <i>The UNESCO Courier</i> from 1948 to 2020 machine-readable, accessible, and reusable for digital text analysis.&nbsp;</p><p>Here on Zenodo we publish two <i>Courier</i> corpora. The first corpus (curated_courier_article_corpus) consists of the texts of all articles published in the English-language edition of <i>The UNESCO Courier</i> between 1948 and 2020. For this corpus we have extracted and reconstructed&nbsp;the complete&nbsp;text of all articles, for example by pulling together non-contiguous pages where necessary and by removing non-article text&nbsp;(masthead, photo captions, letters to the editor, and so on). We have linked each article&nbsp;to a comprehensive curated metadata index, included in the download (document_index.csv).</p><p>The second corpus&nbsp;(curated_issues) compiles&nbsp;the complete text of all <i>Courier</i> issues (English-language edition), 1948-2020. To prepare this corpus we extracted text from <a href="https://en.unesco.org/courier/archives">the PDFs that UNESCO has made available</a>, used multiple modes of OCR, and rendered each issue as a simple text file. Our test of the OCR quality finds an average&nbsp;error rate of 0.7 %, which should be considered good quality.</p><p>Working data from the process can be found in our <a href="https://github.com/inidun/tagged_courier">GitHub repository "tagged Courier."</a> The products, text analysis tools, and additional documentation are in the <a href="https://github.com/inidun/curated_courier">repository "Curated Courier."</a></p><p>The text of <i>The UNESCO Courier</i> is&nbsp;<a href="https://courier.unesco.org/en/about">available in Open Access</a> under the Attribution-ShareAlike 3.0 IGO (CC-BY-SA 3.0 IGO) license, in the context of <a href="https://en.unesco.org/open-access/">UNESCO's open access publications policy</a>. This dataset is published under the most recent version of the same license: Attribution-ShareAlike 4.0 International (<a href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">CC BY-SA 4.0 Deed</a>).</p><p>These datasets&nbsp;was developed as part of the research project "International Ideas at UNESCO: Digital Approaches to Global Conceptual History" (INIDUN), led by Benjamin G. Martin at Uppsala University and funded by a grant from the Swedish Research Council (Vetenskapsrådet), 2020-2024. For more information, see: <a href="https://inidun.github.io">https://inidun.github.io</a>, as well as the<a href="https://github.com/inidun"> project repository on GitHub</a>, which includes documentation and files related to the curating process.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 5

<p>Dataset containing four .xlsx and .csv files for the exercises in Episode 5 of the&nbsp;<a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a>&nbsp;lesson of the&nbsp;<a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a>&nbsp;project. The original data was collected from&nbsp;<a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

MESINESP2 Corpora: Annotated data for medical semantic indexing in Spanish

<p>Gold Standard annotations of the MESINESP2 corpora (training, development and test sets).&nbsp;</p> <p><strong>Please cite this paper if you use this dataset:</strong></p> <pre><code class="language-bash">@inproceedings{gasco2021overview, title={Overview of BioASQ 2021-MESINESP track. Evaluation of advance hierarchical classification techniques for scientific literature, patents and clinical trials}, author={Gasco, Luis and Nentidis, Anastasios and Krithara, Anastasia and Estrada-Zavala, Darryl and Murasaki, Renato Toshiyuki and Primo-Pe{\~n}a, Elena and Bojo Canales, Cristina and Paliouras, Georgios and Krallinger, Martin and others}, year={2021}, organization={CEUR Workshop Proceedings} }</code></pre> <p>&nbsp;</p> <p><strong>Introduction</strong></p> <p>The main aim of MESINESP2 is to promote the development of practically relevant semantic indexing tools for biomedical content in non-English language. We have generated a manually annotated corpus, where domain experts have labeled a set of scientific literature, clinical trials, and patent abstracts. All the&nbsp;documents were labeled with DeCS descriptors, which is a structured controlled vocabulary created by BIREME to index scientific publications on BvSalud,&nbsp;the largest database of scientific documents in Spanish, which hosts records from the databases LILACS, MEDLINE, IBECS, among others.&nbsp;</p> <p>MESINESP track at BioASQ9 explores the efficiency of systems for assigning DeCS to different types of biomedical documents. To that purpose, we have divided the task into three subtracks depending on the document type. Then,&nbsp;for each one we generated an annotated corpus which was provided to participating teams:</p> <ul> <li><strong>[Subtrack 1 corpus] MESINESP-L &ndash; Scientific Literature:&nbsp;</strong>It contains all Spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.</li> <li><strong>[Subtrack 2 corpus] <strong>MESINESP-T- Clinical Trials&nbsp;</strong></strong>contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;</li> <li><strong>[Subtrack 3 corpus] MESINESP-P &ndash; Patents:&nbsp;</strong>This corpus&nbsp;includes patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.</li> </ul> <p>In addition, we also provide a set of complementary data such as: the DeCS terminology file, a silver standard with the participants&#39; predictions to the task background set and the entities of medications, diseases, symptoms and medical procedures extracted from the BSC NERs documents.</p> <p>&nbsp;</p> <p><strong>Files structure:</strong></p> <p><strong>Silver_Standard_Mesinesp2.zip </strong>contains two separate sections. On the one hand, the union of the labels of the best model of each participating team as long as this model had obtained at least an F-score of 0.2 (folder <em>join</em>). On the other hand, the predictions of the best models of each participant have been included individually and anonymized&nbsp;(folder <em>separated</em>).&nbsp;This silver standard contains a set of <em>8642 scientific articles</em>, <em>1537 text sections from Clinical Practice Guidelines</em>, a set of <em>8458 text segments from Medication Data Sheets</em>, <em>461 clinical trials from REEC and 5170 patents</em>.&nbsp;</p> <p><strong>Subtrack1-Scientific_Literature.zip</strong> contains the corpora generated for subtrack 1. Content:</p> <ul> <li>Subtrack1: <ul> <li>Train:&nbsp; <ul> <li>training_set_track1_all.json: Full training set for subtrack 1.&nbsp;</li> <li>training_set_track1_only_articles.json:&nbsp;Articles training set for subtrack 1.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack1.json:&nbsp;</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack1.json: Test set for subtrack 1.&nbsp;</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack2-Clinical_Trials.zip</strong> contains the corpora generated for subtrack 2. Content:</p> <ul> </ul> <ul> <li>Subtrack2: <ul> <li>Train <ul> <li>training_set_subtrack2.json: Training set for subtrack 2.</li> </ul> </li> <li>Development <ul> <li>development_set_subtrack2.json:&nbsp;Manually annotated&nbsp;development set for subtrack 2.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack2.json: Test set for subtrack 2.</li> </ul> </li> </ul> </li> </ul> <p><strong>Subtrack3-Patents.zip</strong> contains the corpora generated for subtrack 3. Content:</p> <ul> </ul> <ul> <li>Subtrack3: <ul> <li>Development <ul> <li>development_set_subtrack3.json:&nbsp;Manually annotated&nbsp;development set for subtrack 3.</li> </ul> </li> <li>Test <ul> <li>test_set_subtrack3.json: Test set for subtrack 3.</li> </ul> </li> </ul> </li> </ul> <p><strong>Additional data.zip&nbsp;</strong>contains the corpora with additional data for each subtrack of MESINESP2.</p> <p><strong>DeCS2020.tsv</strong> contains a DeCS table with the following structure:</p> <ul> <li>DeCS code</li> <li>Preferred descriptor (the preferred label in the Latin Spanish DeCS 2020&nbsp;set)</li> <li>List of synonyms (the descriptors and synonyms from&nbsp; Latin Spanish DeCS 2020&nbsp;set, separated by pipes.</li> </ul> <p><strong>DeCS2020.obo&nbsp;</strong>contains the *.obo file with the hierarchical relationships between DeCS descriptors.</p> <p>*Note: The <em>obo </em>and <em>tsv </em>files with DeCS2020 descriptors contain some additional COVID19 descriptors that will be included in future versions of DeCS. These items were provided by the Pan American Health Organization (PAHO), which has kindly shared this content to improve the results of the task by taking these descriptors into account.</p> <p>&nbsp;</p> <p><strong>Data format&nbsp;description</strong></p> <p>The&nbsp;<strong>input text files</strong>&nbsp;for the MESINESP track are JSON files with the following structure:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "title": "Metas de control de la presión arterial e impacto sobre desenlaces cardiovasculares en pacientes con diabetes mellitus tipo 2: un análisis crítico de la literatura", "abstractText": "La hipertensión arterial en individuos con diabetes mellitus tipo2 incrementa el riesgo de eventos cardiovasculares. Las guías internacionales de manejo recomiendan iniciar tratamiento farmacológico con valores de presión arterial &gt;140/90mmHg Sin embargo, no existe un punto de corte óptimo a partir del cual se logre reducir los eventos cardiovasculares sin originar eventos adversos; un rango de presión arterial &gt;130/80 y &lt;140/90mmHg parece ser el adecuado. Estos valores pueden alcanzarse mediante intervenciones no farmacológicas (dieta, ejercicio) y farmacológicas (por fármacos que hayan demostrado reducir eventos cardiovasculares). La elección de uno o varios fármacos debe ser individualizada, de acuerdo con factores como etnia, edad, comorbilidades asociadas, entre otros", "journal": "Clín. investig. arterioscler. (Ed. impr.)", "year": 2019, "db": "IBECS", "decsCodes": [ "D006973", "D000959", "D002318", "D003924", "D012307" ] } ] }</code></pre> <p>MESINESP&nbsp;<strong>entity mention files</strong>&nbsp;contain automatically generated mention annotations of medications, diseases, syntoms and medical procedures with the following JSON format:</p> <pre><code class="language-json">{ "articles": [ { "id": "ibc-FGT-907", "diseases": [ {"span": "hipertensión arterial", "start": "3", "end": "24"}, {"span": "diabetes mellitus tipo2", "start": "43", "end": "66"}, {"span": "eventos cardiovasculares", "start": "91", "end": "115"}], "medications": [], "procedures": [], "symptoms": []}] } ] }</code></pre> <p>&nbsp;</p> <p><strong>Dataset description:</strong><br> These corpora contain the data for each of the subtracks of MESINESP2 shared-task:</p> <ul> <li><strong>[Subtrack 1] MESINESP-L &ndash; Scientific Literature&nbsp;</strong>: &nbsp; <ul> <li><em><strong>Training set:&nbsp;</strong></em>It contains all spanish records from LILACS and IBECS databases at the Virtual Health Library (VHL) with non-empty abstract written in Spanish.&nbsp;We have filtered out empty abstracts and non-Spanish abstracts.&nbsp;&nbsp;We have built the training dataset with the data crawled on 01/29/2021. This means that the data is a snapshot of that moment and that may change over time since LILACS and IBECS usually add or modify indexes after the first inclusion in the database.&nbsp;We distribute two different datasets: <ul> <li><strong>Articles training set:&nbsp;</strong>This corpus contains the set of 237574 Spanish scientific papers in VHL that have at least one DeCS code assigned to them.</li> <li><strong>Full training set</strong>: This corpus contains the whole set of 249474 Spanish documents from VHL that have at leas one DeCS code assigned to them.</li> </ul> </li> <li><strong>Development set:&nbsp;</strong>We provided a development set manually indexed by our expert annotators (not VHL ones). This dataset includes 1065 articles annotated with DeCS by three expert indexers in this controlled vocabulary. The articles were initially indexed by 7 annotators, after analyzing the Inter-Annotator Agreement among their annotations we decided to select the 3 best ones, considering their annotations the valid ones to build the test set. From those 1065 records: <ul> <li>213 articles were annotated by more than one annotator. We have selected de union between annotations.</li> <li>852 articles were annotated by only one of the three selected annotators with better performance.</li> </ul> </li> <li><strong>Test set:</strong> We provide a test set containing 491 abstracts&nbsp;from LILACS and IBECS. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 2] <strong>MESINESP-T- Clinical Trials</strong></strong>: &nbsp; <ul> <li><strong>Training set:&nbsp;</strong>The training dataset contains records from&nbsp;<a href="https://reec.aemps.es/reec/public/web.html">Registro Espa&ntilde;ol de Estudios Cl&iacute;nicos (REEC)</a>. REEC doesn&#39;t&nbsp;provide documents with the structure title/abstract needed in BioASQ, for that reason we have built artificial abstracts based on the content available in the data crawled using the REEC&nbsp;<a href="https://github.com/luisgasco/REECapi">API</a>.&nbsp;Clinical trials are not indexed with DeCS terminology, we have used as training data a set of 3560 clinical trials that were automatically annotated in the first edition of MESINESP and that were published as a&nbsp;<a href="https://zenodo.org/record/3946558#.YFHyhZ1KiUk">Silver Standard outcome</a>. Because the performance of the models used by the participants was variable, we have only selected predictions from runs with a MiF higher than 0.41, which corresponds with the submission of the best team.&nbsp;</li> <li><strong>Development set: </strong>We provide a development set manually indexed by expert annotators. This dataset includes 147 clinical trials annotated with DeCS by seven expert indexers in this controlled vocabulary.</li> <li><strong>Test set:&nbsp;</strong>The test dataset contains a collection of 248 items. We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>[Subtrack 3] MESINESP-P &ndash; Patents:&nbsp;</strong> <ul> <li><strong>Development set: </strong>We provide a Development set manually indexed by expert annotators. This dataset includes 115 patents in Spanish extracted from Google Patents which have the IPC code &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;. We have selected these patents based on semantic similarity to the MESINESP-L training set to facilitate model generation and to try to improve model performance.</li> <li><strong>Test set:&nbsp;</strong>We provide a&nbsp;<strong>test set</strong>&nbsp;containing 119 records that correspond to a subset of patents published in Spanish with the IPC codes &ldquo;A61P&rdquo; and &ldquo;A61K31&rdquo;.Similarly to the development set, we selected these records based on semantic similarity to the MESINESP-L training set.&nbsp;We used this subset to evaluate the participating systems.</li> </ul> </li> <li><strong>Additional data:</strong> <ul> <li>&nbsp;We provide this information to the participants as additional data in the &ldquo;Additional Data&rdquo; folder. For each training, development, and test set there is an additional JSON file with the structure shown <a href="https://temu.bsc.es/mesinesp2/resources/">here</a>. Each file contains&nbsp;entities related to medications, diseases, symptoms, and medical procedures extrated with the BSC NERs.</li> </ul> </li> </ul> <p>&nbsp;</p> <p><strong>Summary statistics:</strong></p> <table align="center"> <caption>MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP-L</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">Training</th> <td>237574</td> <td>1988684</td> <td>22434</td> <td>43106663</td> </tr> <tr> <th scope="row">Development</th> <td>1065</td> <td>11283</td> <td>3750</td> <td>211420</td> </tr> <tr> <th scope="row">Test</th> <td>491</td> <td>5398</td> <td>2124</td> <td>93645</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Training</th> <td>3560</td> <td>52257</td> <td>3940</td> <td>4133166</td> </tr> <tr> <th scope="row">Development</th> <td>147</td> <td>2038</td> <td>771</td> <td>146791</td> </tr> <tr> <th scope="row">Test</th> <td>248</td> <td>3271</td> <td>905</td> <td>267031</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> <td>&nbsp;</td> </tr> <tr> <th scope="row">Development</th> <td>109</td> <td>1092</td> <td>520</td> <td>38564</td> </tr> <tr> <th scope="row">Test</th> <td>119</td> <td>1176</td> <td>629</td> <td>9065</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> </tbody> </table> <p>&nbsp; </p><table align="center"> <caption>General MESINESP Corpus statistics</caption> <thead> <tr> <th scope="col">MESINESP</th> <th scope="col">Docs</th> <th scope="col">DeCS</th> <th scope="col">Unique DeCS</th> <th scope="col">Tokens</th> </tr> </thead> <tbody> <tr> <th scope="row">MESINESP-L</th> <td>239130</td> <td>2005365</td> <td>22482</td> <td>43411728</td> </tr> <tr> <th scope="row">MESINESP-T</th> <td>3955</td> <td>57566</td> <td>4410</td> <td>4546988</td> </tr> <tr> <th scope="row">MESINESP-P</th> <td>228</td> <td>2268</td> <td>989</td> <td>47629</td> </tr> <tr> <th scope="row"><em>Total</em></th> <td>243313</td> <td>2065199</td> <td>22641</td> <td>48006345</td> </tr> </tbody> </table> <p></p> <p><strong>Related resources:</strong></p> <ul> <li><a href="http://temu.bsc.es/mesinesp2/">MESINESP2&nbsp;Web</a></li> <li><a href="https://github.com/BioASQ/Evaluation-Measures">Evaluation library</a></li> <li><a href="http://metodologia.lilacs.bvsalud.org/download/E/LILACS-4-ManualIndexacao-es.pdf">Annotation guidelines</a></li> <li><a href="https://www.youtube.com/playlist?list=PL5uSCzf1azhCNKd8zhgD0rLwbhxGqF_wX">Participating teams Youtube Videos</a></li> <li><a href="http://ceur-ws.org/Vol-2936/">Proceedings of BioASQ@CLEF2021</a></li> <li><a href="http://bioasq.org/">BioASQ Web</a></li> </ul> <p>&nbsp;</p> <p>For further information, please&nbsp;email us at luis.gasco@bsc.es</p>

opencc-by-4.0Mar 2021View details →
zenodo44/100

MeSpEn_Parallel-Corpora

<p>MeSpEn consists of a resource of heterogeneous health related documents in Spanish and English useful to build parallel corpora for training and evaluating Spanish &lt;-&gt; English medical machine translation systems, to generate multilingual automatic term extraction tools, and develop other Spanish medical NLP components. MeSpEn provides the combination and harmonization of various bibliographic datasets of biomedical and clinical literature from Spain and Latin America or web-content with trusted information sources about diseases, conditions, and wellness issues for patients.</p> <p>MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish.</p> <p>MeSpEn was used to generate automatically bilingual health related-glossaries through automatic term detection and named entity recognition in English and target candidate term extraction in Spanish through sentence alignment approaches, implying potentially the generation of Silver Standard annotated health texts in Spanish (see Villegas, et al. &quot;The MeSpEN resource for English-Spanish medical machine translation and terminologies: census of parallel corpora, glossaries and term translations.&quot; <em>Proc. LREC 2018 Workshop MultilingualBIO: Multilingual Biomedical Text Processing)</em>.</p> <p>The MeSpEn resource aggregates several datasets, mainly from 4 principal sources: IBECS, SciELO, Pubmed and MedlinePlus:</p> <ul> <li> <p><a href="http://ibecs.isciii.es/cgi-bin/wxislind.exe/iah/online/?IsisScript=iah/iah.xis&amp;base=IBECS&amp;lang=i&amp;form=F">IBECS</a> (Spanish Bibliographical Index in Health Sciences) is a bibliographical database that collects scientific journals covering multiple fields in health sciences. It is maintained by the Spanish National Health Sciences Library (BNCS), at the <a href="http://www.eng.isciii.es/">Carlos III Health Institute</a>.</p> <p>This corpus contains titles and abstracts from 168,198 records in English and Spanish. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>. The original XML file of the record provided by IBECS is provided as well.</p> <p>For more information about IBECS parallel corpora, see IBECS_README file.</p> </li> <li> <p><a href="http://scielo.org/php/index.php?lang=en">SciELO</a> (Scientific Electronic Library Online) gathers electronic publications of complete full text articles from scientific journals of Latin America, South Africa and Spain. Currently is present in 15 countries and supported by the Sao Paulo Research Foundation (<a href="http://www.fapesp.br/en/">FAPESP</a>) and the Brazilian National Council for Scientific and Technological Development (<a href="http://bvsalud.org/en/">BIREME</a>).</p> <p>This corpus contains titles and abstracts from 161,710 records in English and Spanish. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>.</p> <p>For more information about SciELO parallel corpora, see Scielo_README file.</p> </li> <li> <p><a href="https://www.ncbi.nlm.nih.gov/pubmed/">Pubmed</a> is a free search engine used to access the <a href="https://www.medline.com/">MedlineNLM</a>).</p> <p>This corpus contains titles and abstracts from 127,619 records. Users can find the metadata of each record written in <a href="http://dublincore.org/">Dublin Core format</a>. The original XML file of the record provided by PubMed is provided as well.</p> <p>For more information about Pubmed parallel corpora, see Pubmed_README file.</p> <p>Users can access to all Spanish articles in Pubmed by <a href="https://www.ncbi.nlm.nih.gov/pubmed?term=%22spanish%22%5BLanguage%5D">clicking here</a>. Follow these steps to download all articles&#39; metadata in XML format:</p> <ul> <li>Click on <em>Send to</em>.</li> <li>Select <em>File</em> on <em>Choose destination</em>.</li> <li>Select <em>XML</em> on <em>Format</em>.</li> <li>And finally click on <em>Create File</em>.</li> </ul> </li> <li> <p><a href="https://medlineplus.gov/">MedlinePlus</a> is an online information service provided by the U.S. National Library of Medicine (<a href="https://www.nlm.nih.gov/ target=">NLM</a>), and gives free information about health in both English and Spanish. MedlinePlus provides the following information: <a href="https://medlineplus.gov/healthtopics.html">Health topics, </a><a href="https://medlineplus.gov/druginformation.html">Drugs and supplements, </a><a href="https://medlineplus.gov/labtests.html">Laboratory test information, </a><a href="https://medlineplus.gov/encyclopedia.html">Medical encyclopedia.</a></p> <p>There are 2 corpora available for download:</p> <ul> <li>Health topics metadata in <a href="http://dublincore.org/">Dublin Core format</a>: the source code of the site stores metadata information about each topic, we created the DC files based on these metadata. This collection contains a total of 1,063 articles in English and Spanish. For more information about it, see MedlinePlus-health-topics_README.</li> <li>Complete MedlinePlus in <a href="http://www.tei-c.org/index.xml">TEI format</a>: clean raw text and XML files of each article, structured by sections and paragraphs. This collection contains a total of 7,033 articles in English and Spanish. For more information about it, see MedlinePlus-articles_README.</li> </ul> </li> </ul> <p>These corpora are also available at http://temu.bsc.es/mespen/</p> <p>In addition, forty-six bilingual medical glossaries for various language pairs are available at https://zenodo.org/record/2205690#.XefkzdEo9hF</p> <p>Copyright (c) 2019 Secretar&iacute;a de Estado para el Avance Digital</p>

opencc-by-4.0Dec 2019View details →
zenodo44/100

Dataset for the paper "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia"

<p>Dataset for the EMNLP'24 Main conference paper titled "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia".</p>

opencc-by-sa-4.0Oct 2024View details →
zenodo44/100

MIGR-TWIT CORPORA. Migration Tweets of French Left-wing Politics.

<p><strong>Description</strong></p> <p>The&nbsp;<strong>FR-L-MIGR-TWIT Corpus</strong>&nbsp;is part of the&nbsp;<strong><a href="https://www.ortolang.fr/market/corpora/migr-twit-corpus">MIGR-TWIT CORPORA</a></strong>, diachronic bilingual corpus of Tweets about the topic of migration in Europe.<br>Within the framework of the collaborative research project&nbsp;<a href="https://olindinum.huma-num.fr/recherche/">OLiNDiNUM</a>&nbsp;(Observatoire LINguistique du DIscours NUM&eacute;rique, [Linguistic Observatory of Online Debate]), the MIGR-TWIT Corpora are created with the aim to study the evolution of the public discourse on migration in Europe during the past dozen years from 2011 to 2022. First two components of the corpus represent migration discourse of right-wing politics in France and in the UK. The&nbsp;FR-L-MIGR-TWIT Corpus&nbsp;represents French left-wing politics' migration discourse on Twitter.&nbsp;&nbsp;</p> <p>Using the&nbsp;<em>Twitter API v2 Academic Research</em>, the Tweets containing at least one occurrence of lexicon derived from a latin root "<em>migr</em>" of <em>migrare </em>are automatically retrieved from 23 Twitter accounts of French left-wing political figures and parties.<br>&nbsp;</p> <p><strong>Scientific reference : &nbsp;</strong>Jeon, S. (2025). Le discours num&eacute;rique sur l'immigration en France entre 2011 et 2022. Une analyse de corpus (Online Discourse on Immigration in France between 2011 and 2022. A Corpus Analysis), PhD Thesis, Universit&eacute; de Lille, France.</p> <p><strong>Contents</strong><br>The&nbsp;downloadable version of <strong>FR-L-MIGR-TWIT-2011-2022</strong>&nbsp;<strong>Corpus&nbsp;</strong>contains&nbsp;32&nbsp;CSV files (tabular format). The corpus is presented in simplified and complete versions in terms of metadata. The simplified version corresponds to one single file named&nbsp;<em><strong>FR-L-MIGR-TWIT-2011-2022.csv</strong></em>, containing four basic (meta)data, <em>i.e</em>. identifier, text, posting date&nbsp;and username (that is,&nbsp;<em><strong>data__id</strong></em>,&nbsp;<strong>data__text</strong>,&nbsp;<em><strong>data__created_at</strong></em>&nbsp;and&nbsp;<strong><em>author__name</em></strong><em>&nbsp;</em>as the table hearder elements).&nbsp;In addition to these four (meta)data, the elaborate version is provided with all Tweet fields information included as a header element, such as the numbers of Replies, Retweets, Likes and Quotes, etc. This version is also available in one single CSV file named&nbsp;<em><strong>FR-L-MIGR-TWIT-2011-2022_meta.csv</strong></em>.</p> <p>Besides, the elaborate version is provided with three CSV Zip&nbsp;files: 7&nbsp;CSV files in the zip file named <em>FR-L-MIGR-TWIT-</em><strong><em>YEAR</em></strong><em>_meta</em> correspond&nbsp;to grouped years (<em>i.e.&nbsp;FR-L-MIGR-TWIT-<strong>2011-2016</strong>_meta.csv</em>) or each and every year (<em>e.g. FR-L-MIGR-TWIT-<strong>2017</strong>_meta.csv, </em>and so on) for the last dozen years. 23 files in the zip file named <em>FR-L-<strong>NAME</strong>-MIGR-TWIT_meta</em> for each and every component of selected French left-wing political figures and parties (<em>e.g. FR-L-<strong>Arthaud</strong>-TWIT_meta.csv</em>). The zip file named FR-L-MIGR-TWIT-2011-2022_meta contains yearly Tweets of each and every component of political figures and parties.</p> <p>Detailed information of the&nbsp;FR-L-MIGR-TWIT-2011-2022 CORPUS&nbsp;is illustrated below.</p> <ul> <li><strong>Created at:</strong>&nbsp;2023-04-18</li> <li><strong>Language:</strong>&nbsp;FR</li> <li><strong>Coverage:</strong>&nbsp;<strong>23</strong> <strong>user accounts</strong> ; <strong>5,636 Tweets</strong>&nbsp;; <strong>169,818 words</strong></li> <li><strong>Time of data collection:</strong>&nbsp;start=2011-01-01&nbsp;; end=2022-06-30</li> <li><strong>Keywords:&nbsp;</strong>words derived from a&nbsp;latine root &ldquo;<strong><em>migr</em></strong>&rdquo; of&nbsp;<em>migrare</em></li> <li><strong>Corpus composition:</strong></li> </ul> <table> <tbody> <tr> <td> <p>&nbsp;</p> </td> <td> <p><strong>Political Figure/party</strong></p> </td> <td> <p><strong>Type of representative</strong></p> </td> <td> <p><strong>Username</strong></p> </td> <td> <p><strong><em>migr</em>-Tweets</strong></p> </td> </tr> <tr> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Adrien Quatennens</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@AQuatennens</strong></p> </td> <td> <p><strong>315</strong></p> </td> </tr> <tr> <td> <p><strong>2</strong></p> </td> <td> <p><strong>Alexis Corbi&egrave;re</strong></p> </td> <td> <p><strong>PERSON(M)</strong></p> </td> <td> <p><strong>@Alexiscorbiere</strong></p> </td> <td> <p><strong>209</strong></p> </td> </tr> <tr> <td> <p><strong>3</strong></p> </td> <td> <p><strong>Anne Hidalgo</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Anne_Hidalgo</strong></p> </td> <td> <p><strong>801</strong></p> </td> </tr> <tr> <td> <p><strong>4</strong></p> </td> <td> <p><strong>Arnaud Montebourg*</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@montebourg</strong></p> </td> <td> <p><strong>7</strong></p> </td> </tr> <tr> <td> <p><strong>5</strong></p> </td> <td> <p><strong>Beno&icirc;t Hamon</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@benoithamon</strong></p> </td> <td> <p><strong>172</strong></p> </td> </tr> <tr> <td> <p><strong>6</strong></p> </td> <td> <p><strong>Christiane Taubira</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@ChTaubira</strong></p> </td> <td> <p><strong>11</strong></p> </td> </tr> <tr> <td> <p><strong>7</strong></p> </td> <td> <p><strong>Cl&eacute;mentine Autain</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Clem_Autain</strong></p> </td> <td> <p><strong>102</strong></p> </td> </tr> <tr> <td> <p><strong>8</strong></p> </td> <td> <p><strong>Dani&egrave;le Obono</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@Deputee_Obono</strong></p> </td> <td> <p><strong>415</strong></p> </td> </tr> <tr> <td> <p><strong>9</strong></p> </td> <td> <p><strong>Esther Benbassa**</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@EstherBenbassa</strong></p> </td> <td> <p><strong>936</strong></p> </td> </tr> <tr> <td> <p><strong>10</strong></p> </td> <td> <p><strong>Fran&ccedil;ois Hollande</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@fhollande</strong></p> </td> <td> <p><strong>28</strong></p> </td> </tr> <tr> <td> <p><strong>11</strong></p> </td> <td> <p><strong>Fran&ccedil;ois_Ruffin</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@Francois_Ruffin</strong></p> </td> <td> <p><strong>19</strong></p> </td> </tr> <tr> <td> <p><strong>12</strong></p> </td> <td> <p><strong>Jean-Luc M&eacute;lenchon</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@JLMelenchon</strong></p> </td> <td> <p><strong>240</strong></p> </td> </tr> <tr> <td> <p><strong>13</strong></p> </td> <td> <p><strong>Manon Aubry</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@ManonAubryFr</strong></p> </td> <td> <p><strong>182</strong></p> </td> </tr> <tr> <td> <p><strong>14</strong></p> </td> <td> <p><strong>Natalie Arthaud</strong></p> </td> <td> <p><strong>PERSON (F)</strong></p> </td> <td> <p><strong>@n_arthaud</strong></p> </td> <td> <p><strong>165</strong></p> </td> </tr> <tr> <td> <p><strong>15</strong></p> </td> <td> <p><strong>Philippe Poutou</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@PhilippePoutou</strong></p> </td> <td> <p><strong>83</strong></p> </td> </tr> <tr> <td> <p><strong>16</strong></p> </td> <td> <p><strong>Raphael Glucksmann</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@rglucks1</strong></p> </td> <td> <p><strong>142</strong></p> </td> </tr> <tr> <td> <p><strong>17</strong></p> </td> <td> <p><strong>Yannick Jadot</strong></p> </td> <td> <p><strong>PERSON (M)</strong></p> </td> <td> <p><strong>@yjadot</strong></p> </td> <td> <p><strong>374</strong></p> </td> </tr> <tr> <td> <p><strong>18</strong></p> </td> <td> <p><strong>Europe &Eacute;cologie-Les Verts</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@EELV</strong></p> </td> <td> <p><strong>484</strong></p> </td> </tr> <tr> <td> <p><strong>19</strong></p> </td> <td> <p><strong>Gauche R&eacute;publicaine et Socialiste</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@Gauche_RS</strong></p> </td> <td> <p><strong>73</strong></p> </td> </tr> <tr> <td> <p><strong>20</strong></p> </td> <td> <p><strong>G&eacute;n&eacute;ration.s</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@GenerationsMvt</strong></p> </td> <td> <p><strong>165</strong></p> </td> </tr> <tr> <td> <p><strong>21</strong></p> </td> <td> <p><strong>La France Insoumise</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@FranceInsoumise</strong></p> </td> <td> <p><strong>300</strong></p> </td> </tr> <tr> <td> <p><strong>22</strong></p> </td> <td> <p><strong>Parti Radical Gauche</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@PartiRadicalG</strong></p> </td> <td> <p><strong>37</strong></p> </td> </tr> <tr> <td> <p><strong>23</strong></p> </td> <td> <p><strong>Parti Socialiste</strong></p> </td> <td> <p><strong>ORGANIZATION</strong></p> </td> <td> <p><strong>@partisocialiste</strong></p> </td> <td> <p><strong>376</strong></p> </td> </tr> </tbody> </table> <ul> <li>Political figures and parties, listed in alphabetical order, are selected according to the four criteria: (1) the high number of&nbsp;<em>migr</em>-tweets, (2) the political affiliation, (3) the political careers, that is, the Member of the European Parliament or (4) the presidential candidate during the period between 2011 and 2022. These four criteria are not mutually exclusive.</li> <li>As part of a doctoral thesis (<a href="https://theses.fr/s360032">Jeon, 2025</a>), the FR-L-MIGR-TWIT and FR-R-MIGR-TWIT corpora are compiled, annotated and analyzed through a comparative discourse analysis approach, with the aim to study the semantic construction of <em>migr</em>-lexicon over the period between 2011 and 2022.</li> <li>*One migration Tweet retrieved from the user account @montebourg for the year of 2019 was removed and is not included in his 7&nbsp;<em>migr</em>-tweets because it refers to the issue of the migration of honey bees.</li> <li>**We later added the user account @EstherBenbassa represented by Esther Benbassa, senator and former member of political party Europe &Eacute;cologie-Les Verts (representative of the user account @EELV), because of the high number of her&nbsp;<em>migr</em>-tweets that were retweeted by @EELV.</li> </ul> <p>The&nbsp;<strong>MIGR-TWIT</strong>&nbsp;<strong>Corpus</strong>&nbsp;consists of three subcorpora for a total amount of&nbsp;<strong>23,869&nbsp;Tweets </strong>and&nbsp;<strong>703,016</strong>&nbsp;<strong>words</strong>:</p> <ul> <li>FR-R-MIGR-TWIT-2011-2022 Corpus:&nbsp;<em>French Right-wing</em>&nbsp;politics'&nbsp;<em>migr</em>-tweets</li> <li>UK-R-MIGR-RA-TWIT-2011-2022 Corpus:&nbsp;<em>British&nbsp;Right-wing</em>&nbsp;politics'&nbsp;<em>migr</em>-tweets</li> <li>FR-L-MIGR-TWIT-2011-2022 Corpus:&nbsp;<em>French Left-wing</em>&nbsp;politics'&nbsp;<em>migr</em>-tweets&nbsp;</li> </ul> <p>&nbsp;</p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

Measuring coselectional constraint in learner corpora: A graph-based approach

<p>All data from the thesis, plots, scripts, rough draft of annotation guidelines (more is included in thesis). not all svgs are included yet, but can be computed from scripts + data. more data will be added in next version (after defense).</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Files and code for English dictionaries, gold and silver standard corpora for biomedical natural language processing related to SARS-CoV-2 and COVID-19

<p><span lang="EN-GB">Automated information extraction with natural language processing (NLP) tools is required to gain systematic insights from the large number of COVID-19 publications, reports and social media posts, which far exceed human processing capabilities.&nbsp;</span></p> <p><span lang="EN-GB">Here we present an NLP toolbox comprising COVID-19-related dictionaries and annotated corpora in English as well as useful code and workflows for their update and use. The dictionaries contain terms referring to the COVID-19 disease, the SARS-CoV-2 virus, its variants and common mutations, respectively. They were used together with the EasyNER NLP tool to extract and annotate all 764&nbsp;398 abstracts in the CORD-19 dataset, creating a very large silver standard corpus (named Lund-Annotated-CORD-19 corpus). This was complemented with a small gold standard corpus consisting of PubMed abstracts manually annotated for key entity classes such as disease, virus, symptom, protein/gene, cell type, chemical and species terms. </span></p> <p><span lang="EN-GB">The toolbox can support various text analysis tasks related to COVID-19 such as named entity recognition and co-mention analysis. A preliminary version of the toolbox, which was released early in the pandemic, was</span><span lang="EN-GB"> for example already used to create a COVID-19 knowledge graph and study the evolution and variation of COVID-19-related terminology. In addition, the toolbox can be applied in the development of other NLP tools, for example to train and evaluate large language models.</span></p> <p><span lang="EN-GB">When using the toolbox, please cite this record and the associated article.</span></p> <p>&nbsp;</p> <p>&nbsp;</p>

openJun 2022View details →
zenodo40/100

Accelerating Digital Skills for Music Researchers - Processing Text-Based Corpora for Musical Discourse Analysis - Episode 2

<p>Dataset containing three subgenre-specific .xlsx files for the exercises in Episode 2 of the <a href="https://acceleratingdigitalskills.github.io/Processing-Text-Based-Corpora/">Processing Text-Based Corpora for Musical Discourse Analysis</a> lesson of the <a href="https://acceleratingdigitalskills.org/">Accelerating Digital Skills for Music Researchers</a> project. The original data was collected from <a href="https://boomkat.com/">Boomkat.com</a> with permission.</p>

opencc-by-4.0Mar 2024View details →
zenodo40/100

Annotated Corpora of Historical Catalan (HisCat) - Llibre dels Fets

<p>This repository is part of the Annotated Corpora of Historical Catalan (HisCat). It contains the first POS-tagged text that is partially manually corrected and used to train Old Catalan POS taggers, described in the following paper:</p> <p>Meelen, Marieke &amp; Pujol i Campeny, Afra, (2021) &#39;Old Catalan Morphosyntax: developing an annotated corpus&#39; in <em>Journal of Open Humanities Data</em>.</p> <p>This POS-tagged text is the 13th century <em>Llibre dels Fets</em>, a historical chronicle. The version of the text used for this project is</p> <p>Bruguera, J. (1991). <em>El Llibre dels Fets del Rei en Jaume</em>. Barcelona: Barcino.</p> <p>as prepared for the <em>Corpus Informatitzat del Catal&agrave; Antic</em></p> <p>Torruella, J., P&eacute;rez Saldanya, M., &amp; Martines, J. (2009). <em>Corpus Informatitzat del Catal&agrave; Antic</em>. URL: <a href="http://cica.cat/">http://cica.cat/</a>.</p> <p>The subcorpus counts with 164,096 POS-annotated tokens (165,538 tokens including punctuation and folio markers), of which 60,000 have been manually corrected. This subcorpus contains a total of and 4,506 main clauses. POS tagging of this text was done with the Memory-Based Tagger by TiMBL (<a href="https://languagemachines.github.io/mbt/">https://languagemachines.github.io/mbt/</a>). The code accompanying the paper can be found on GitHub: <a href="https://github.com/lothelanor/catalancorpora">https://github.com/lothelanor/catalancorpora</a>). In addition to memory-based tagging, have tried neural-based tagging with TARGER (<a href="https://github.com/achernodub/targer">https://github.com/achernodub/targer</a>) for which we created word embeddings that can be found on <a href="https://doi.org/10.5281/zenodo.5615556">Zenodo</a>. Results for memory-based tagging were better, however, which is why this version is uploaded here.</p> <pre>&nbsp;</pre>

opencc-by-4.0Oct 2021View details →
zenodo40/100

BiodivBERT: Pre-training Corpora DOIs

<p>This repository contains the DOIs we used to construct the&nbsp;pre-training corpora for BiodivBERT model.</p> <p>BiodivBert<sub>Abs<sup>&nbsp;</sup></sub>uses the abstracts DOIs&nbsp;while BiodivBERT<sub>Abs+Full</sub> uses both of them.&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo40/100

IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022

<p>IsiZulu News (articles and headlines) and Siswati News (headlines) Corpora - za-isizulu-siswati-news-2022</p> <p>Reference paper</p> <p>Madodonga, A., Marivate, V., &amp; Adendorff, M. (2023). Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati.&nbsp;<em>Journal of the Digital Humanities Association of Southern Africa</em>,&nbsp;<em>4</em>(01). https://doi.org/10.55492/dhasa.v4i01.4449</p> <p>&nbsp;</p> <p>&gt;&nbsp;@article{Madodonga_Marivate_Adendorff_2023, title={Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati}, volume={4}, url={https://upjournals.up.ac.za/index.php/dhasa/article/view/4449}, DOI={10.55492/dhasa.v4i01.4449}, author={Madodonga, Andani and Marivate, Vukosi and Adendorff, Matthew}, year={2023}, month={Jan.} }</p>

opencc-by-sa-4.0Oct 2022View details →
zenodo40/100

AQL queries and benchmark results from PhD thesis "ANNIS: A graph-based query system for deeply annotated text corpora"

<p>These are the queries, the benchmark results and the evaluation scripts of the thesis &quot;ANNIS: A graph-based query system for deeply annotated text corpora&quot; (Thomas Krause 2018, Humboldt-Universit&auml;t zu Berlin)</p> <p><strong>diss_2018-01-12_v0.5.0.csv </strong><br> Results of all configurations of executed benchmarks for graphANNIS and also the baseline times of relANNIS.</p> <p><strong>queries.zip</strong><br> Contains folders for each corpus containing all queries used for the benchmark. Each file-name begins with the ID of the query. The extension denotes the type, which can be one of the following:</p> <ul> <li><em>&quot;.</em>aql&quot; contains the original AQL (ANNIS query language) query which was collected</li> <li>&quot;.json&quot; is JSON representation of the parsed AQL query</li> <li>&quot;.count&quot; is the number of matches a query should have</li> <li>&quot;.time&quot; is the average time in milliseconds that was needed to execute the query in relANNIS on the benchmark system</li> <li>&quot;.corpora&quot; contains the name of the corpus the query belongs to (should be only one corpus and the same as the folder name in the selection of queries in this data set)</li> <li>&quot;.relplan&quot; contains the PostgreSQL plan for the query</li> <li>&quot;.graphplan&quot; contains the graphANNIS plan for the query</li> </ul> <p><strong>evaluation-scripts.py/evaluation-scripts.ipynb</strong><br> Python scripts to perform the evaluation and generate the images. This are both a Python-file and the original notebook file that can be used with the Jupyter Notebook application.</p> <p><strong>relannis_benchmark_scripts.zip </strong><br> The files in this zip-file can be used to execute the benchmarks in the relANNIS system by piping the into the &quot;annis.sh&quot; command line tool of relANNIS</p>

opencc-by-4.0Jan 2018View details →
zenodo40/100

Common Sounds in Bedrooms (CSIBE) Corpora

<p>These audio corpora can be used for domestic sound event recognition. The current sound events in the datasets: baby cry, bell, cat meow, cicada, dog bark, fart, guitar, laugh, parrot, piano, speech, traffic, vacuum cleaner and the background sounds. This latter class&nbsp;can model the&nbsp;ambient sounds which are probably&nbsp;not important for a robot or their short audition is not sufficient for a human without visual cue: door, drawer, keys, knock, pen, chair, cup, keyboard, breathing, throat, cough, microwave, steps and zip.</p> <p>Currently, there are two parts:</p> <p>-&nbsp;<strong>&nbsp;CSIBE-RAW:</strong>&nbsp;Human&nbsp;speech&nbsp;and&nbsp;other&nbsp;events&nbsp;were&nbsp;collected&nbsp;from&nbsp;the&nbsp;internet&nbsp;in&nbsp;this&nbsp;dataset,&nbsp;complemented&nbsp;with&nbsp;new&nbsp; recordings.The&nbsp;samples&nbsp;had&nbsp;excellent&nbsp;and&nbsp;clear&nbsp;sound&nbsp;quality,&nbsp;they&nbsp;were&nbsp;stored&nbsp;in&nbsp;mono&nbsp;WAV&nbsp;format&nbsp;with&nbsp;16&nbsp;bit&nbsp;depth, 44.1&nbsp;kHz&nbsp;sampling&nbsp;rate.&nbsp;All&nbsp;files&nbsp;were&nbsp;labeled&nbsp;according&nbsp;to&nbsp;the&nbsp;sound&nbsp;event&nbsp;type.</p> <p>- <strong>CSIBE-AIBO:&nbsp;</strong>CSIBE-RAW was recorded again with a robot. The original sounds were played back on a high-quality speaker and the result was recorded with the stereo microphones of a Sony ERS-7 robot in a silent room. Four configurations were used&nbsp;relative to the robot:</p> <p>1. Reverberant room, speaker was 1 meter away, 30⁰ counterclockwise to the head.<br> 2. Non-reverberant room, speaker was 1 meter away, 30⁰ counterclockwise to the head.<br> 3. Reverberant room, speaker was 3 meters away, 1 meter high, 180⁰ clockwise to the head.<br> 4. Non-reverberant room, speaker was 3 meters away, 1 meter high, 180⁰ clockwise to the head.</p> <p>These conditions were varied to get such sound recordings which contains dynamics of the robot microphones, different reverberant conditions and&nbsp;various positions of the sound sources.</p> <p>All licenses of the original sounds&nbsp;are included in the dataset.</p>

opencc-by-4.0May 2018View details →
zenodo40/100

waleghwa/low-resource-language-data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin and Dholuo

<p><strong>Description</strong>: The dataset consists of three parallel corpora: Kidaw'ida-Kiswahili; Kalenjin-Kiswahili; Dholuo-Kiswahili. On averate, each corpus has thirty thousand sentence pairs. This dataset is also available on GitHub where it will continue to be grown and its quality improved. Future releases will be uploaded here on Zenodo as new versions.</p> <p><strong>Purpose of the dataset</strong>: The dataset was created for use in training machine translation models. This is to enable translation from Kiswahili, which is the national language in Kenya, into indigenous languages. Three indigenous Kenyan languages were targeted, namely, Kidaw'ida, Kalenjin, and Dholuo.</p> <p><strong>Principal Investigator</strong>: Audrey Mbogho, United States International University - Africa</p> <p><strong>Co-Investigators</strong>:</p> <ol> <li>Andrew Kipkebut, Kabarak University</li> <li>Quin Awuor, United States International University - Africa</li> <li>Rose Lugano, University of Florida</li> <li>Lilian Wanzare, Maseno University</li> <li>Vivian Oloo, Maseno University</li> </ol> <p><strong>Funding</strong>: This dataset was collected with funding from Lacuna Fund.</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Spanish Legal Domain Corpora

<p><strong>Spanish Legal Domain Corpora</strong></p> <p>A collection of corpora of Spanish legal domain.</p> <p>More legal domain resources:&nbsp;https://github.com/PlanTL-GOB-ES/lm-legal-es</p> <p><strong>Citation</strong></p> <pre><code>@misc{gutierrezfandino2021legal, title={Spanish Legalese Language Model and Corpora}, author={Asier Gutiérrez-Fandiño and Jordi Armengol-Estapé and Aitor Gonzalez-Agirre and Marta Villegas}, year={2021}, eprint={2110.12201}, archivePrefix={arXiv}, primaryClass={cs.CL} }</code></pre> <p><strong>Copyright </strong></p> <p>Copyright (c) 2021 Secretar&iacute;a de Estado de Digitalizaci&oacute;n e Inteligencia Artificial</p>

opencc-by-4.0Sep 2021View details →
zenodo40/100

Hierarchical Text Classification corpora

<p>A set of 3 datasets for Hierarchical Text Classification (HTC), with samples divided into training and testing splits. The hierarchies of labels within all datasets have depth 2.</p> <ul> <li>The <strong>Amazon5x5</strong> dataset contains 500,000 user reviews tagged with the reviewed product's categories. There are 5 product categories with 100,000 examples each, and each category has 5 sub-categories.</li> <li>The <strong>Bugs</strong> dataset contains 30,050 bugs of the Linux kernel, labeled with exactly two categories identifying the affected component.</li> <li>Finally, the <strong>Web Of Science</strong> dataset contains 46,960 abstracts of scientific papers, labeled the article's domain (see <a href="https://data.mendeley.com/datasets/9rw3vkcfy4/6">original repo</a> for more details).</li> </ul> <p>Datasets are published in JSONL format, where each line is a string formatted as a JSON, like in the example below.</p> <pre><code>{ "text": &lt;article text&gt;, "labels": [&lt;label1&gt;, &lt;label2&gt;, ...] }</code></pre> <p>The <em>hierarchical structure</em> of labels in each dataset is documented in <a href="https://gitlab.com/distration/dsi-nlp-publib/-/tree/main/htc-survey-24/data/taxonomies">this repository</a>.</p> <p>&nbsp;</p> <p>These datasets have been presented in this paper:</p> <ul> <li>"Hierarchical Text Classification and its Foundations: a Review of Current Research" - DOI: <a href="https://doi.org/10.3390/electronics13071199">10.3390/electronics13071199</a></li> </ul> <p>Some of these datasets have also been used in:</p> <ul> <li>"Ticket Automation: an Insight into Current Research with Applications to Multi-level Classification Scenarios" - DOI: <a href="https://doi.org/10.1016/j.eswa.2023.119984">10.1016/j.eswa.2023.119984</a></li> <li>"A multi-level approach for hierarchical Ticket Classification", accepted at WNUT 2022 - <a href="https://aclanthology.org/2022.wnut-1.22/">link</a></li> </ul> <p>&nbsp;</p> <p>These datasets are partially derived from previous work, namely:</p> <ul> <li>[Amazon] J. Ni, J. Li, J. McAuley, "Justifying Recommendations using Distantly-Labeled Reviews and Fine-Grained Aspects", EMNLP 2019, doi: <a href="http://dx.doi.org/10.18653/v1/D19-1018">10.18653/v1/D19-1018</a></li> <li>[WOS] K. Kowsari, D. E. Brown, M. Heidarysafa, K. Jafari Meimandi, M. S. Gerber and L. E. Barnes, "HDLTex: Hierarchical Deep Learning for Text Classification," 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), 2017, pp. 364-371, doi: <a href="http://doi.org/10.1109/ICMLA.2017.0-134">10.1109/ICMLA.2017.0-134</a></li> <li>[Linux Bugs] V. Lyubinets, T. Boiko and D. Nicholas, "Automated Labeling of Bugs and Tickets Using Attention-Based Mechanisms in Recurrent Neural Networks," <em>2018 IEEE Second International Conference on Data Stream Mining &amp; Processing (DSMP)</em>, 2018, pp. 271-275, doi: <a href="http://doi.org/10.1109/DSMP.2018.8478511">10.1109/DSMP.2018.8478511</a></li> </ul>

opencc-by-4.0Dec 2022View details →
zenodo36/100

LatinISE test data for SemEval 2020 task 1 with additional token versions of the corpora

<p>This data collection contains the Latin test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:&nbsp;</p> <ul> <li>a Latin text corpus pair (`corpus1/lemma`, `corpus2/lemma`)</li> <li>40 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)</li> <li>the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</li> </ul> <p>The corpus data have been automatically lemmatized and part-of-speech tagged, and have been partially corrected by hand. For homonyms, the lemmas are followed by the &#39;\#&#39; symbol and the number of the homonym according to the Lewis-Short dictionary of Latin when this number is greater than 1. For example, the lemma &#39;dico&#39; corresponds to the first homonym in the Lewis-Short dictionary and &#39;dico\#2&#39; corresponds to the second homonym, cf. Lewis-Short dictionary.</p> <p>__Corpus 1__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a>&nbsp;(McGillivray and Kilgarriff 2013), <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the second century before Christ (BC) to the end of the first century BC</li> <li>size: ~1.7 million tokens</li> <li>format: lemmatized, sentence length &gt;= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>__Corpus 2__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a>&nbsp;(McGillivray and Kilgarriff 2013) , <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the first century after Christ (AD) to the end of the twenty-first century AD</li> <li>size: ~9.4 million tokens</li> <li>format: lemmatized, sentence length &gt;= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>Find more information on the data in the papers referenced below.</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (<code>corpus1/token/</code>,&nbsp;<code>corpus2/token/</code>). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>The creation of the data was supported by the CRETA center and the CLARIN-D grant funded by the German Ministry for Education and Research (BMBF).</p> <p><strong>References</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p> <p>McGillivray, B. and Kilgarriff, A. (2013). <a href="https://www.sketchengine.co.uk/wp-content/uploads/2015/05/Latin_historical_corpus_2013.pdf">Tools for historical corpus research, and a corpus of Latin</a>. In Paul Bennett, Martin Durrell, Silke Scheible, Richard J. Whitt (eds.), New Methods in Historical Corpus Linguistics, T&uuml;bingen: Narr.<br> &nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

BiodivNERE: Gold Standard Corpora for Named Entity Recognition and Relation Extraction in Biodiversity Domain

<p>BiodivNER+RE are two gold-standard manually annotated corpora that are meant to be for Named Entity Recognition (NER) and Relation Extraction (RE) tasks based on abstracts and metadata files from the biodiversity domain</p> <p>Such corpora are designed for machine learning techniques,&nbsp;for example, NER as&nbsp;TokenClassification technique, and RE as&nbsp;SequenceClassification technique.</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record