Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

135

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

135 results for “multilingual”

Learn how ShareScore rates datasets ↗
zenodo52/100

Multilingual news article similarity dataset

<p>This dataset contains the extended version of the authors' earlier work:&nbsp;<a href="../records/6507872">https://zenodo.org/records/6507872,</a> where pairs of news articles drawn from the first half of 2020 are annotated for seven aspects of similarity in the original version as well as an additional FRAME aspect:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> <li><strong>FRAME</strong> Do the articles have similar framing and express similar opinions?</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo52/100

SemEval-2022 Task 8: Multilingual news article similarity

<p>This dataset contains pairs of news articles drawn from the first half of 2020 and annotated for seven aspects of similarity:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong>&nbsp;Do the articles have similar tones?</li> </ul> <p>Further details are provided in</p> <blockquote> <p>Chen et al. (2022). SemEval-2022 Task 8: Multilingual news article similarity. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022).&nbsp;<a href="https://aclanthology.org/2022.semeval-1.155/">https://aclanthology.org/2022.semeval-1.155/</a></p> </blockquote> <p>The data in this repository includes pairs of URLs and annotations. The text of webpages is generally&nbsp;via the Internet Archive in this special collection: https://archive.org/details/2020-multilingual-news-article-similarity . A script to download and process the webpages is available at&nbsp;https://github.com/euagendas/semeval_8_2022_ia_downloader .&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo48/100

Curlie Enhanced with LLM Annotations: Two Datasets for Advancing Homepage2Vec's Multilingual Website Classification

<h3>Advancing Homepage2Vec with LLM-Generated Datasets for Multilingual Website Classification</h3> <p>This dataset contains two subsets of labeled website data, specifically created to enhance the performance of Homepage2Vec, a multi-label model for website classification. The datasets were generated using Large Language Models (LLMs) to provide more accurate and diverse topic annotations for websites, addressing a limitation of existing Homepage2Vec training data.</p> <p><strong>Key Features:</strong></p> <ul> <li><strong>LLM-generated annotations:</strong>&nbsp;Both datasets feature website topic labels generated using LLMs,&nbsp;a novel approach to creating high-quality training data for website classification models.</li> <li><strong>Improved multi-label classification:</strong> Fine-tuning Homepage2Vec with these datasets has been shown to improve its macro F1 score from 38% to 43% evaluated on a human-labeled dataset, demonstrating their effectiveness in capturing a broader range of website topics.</li> <li><strong>Multilingual applicability:</strong> The datasets facilitate classification of websites in multiple languages, reflecting the inherent multilingual nature of Homepage2Vec.</li> </ul> <p><strong>Dataset Composition:</strong></p> <ul> <li><strong>curlie-gpt3.5-10k:</strong> 10,000 websites labeled using GPT-3.5, context 2 and 1-shot</li> <li><strong>curlie-gpt4-10k:</strong> 10,000 websites labeled using GPT-4, context 2 and zero-shot</li> </ul> <p><strong>Intended Use:</strong></p> <ul> <li>Fine-tuning and advancing Homepage2Vec or similar website classification models</li> <li>Research on LLM-generated datasets for text classification tasks</li> <li>Exploration of multilingual website classification</li> </ul> <p><strong>Additional Information:</strong></p> <ul> <li><strong>Project and report repository:</strong> https://github.com/CS-433/ml-project-2-mlp</li> </ul> <p><strong>Acknowledgments:</strong></p> <p>This dataset was created as part of a project at EPFL's Data Science Lab (DLab) in collaboration with <a href="https://people.epfl.ch/robert.west">Prof. Robert West</a> and <a href="https://tizianopiccardi.github.io/" rel="nofollow">Tiziano Piccardi.</a></p>

opencc-by-4.0Dec 2023View details →
zenodo48/100

Multilingual Wikidata Property Translation Flow Dataset

<p>Multilingual Wikidata property translation Flow dataset contains the translation flow of Wikidata properties as collected on July 7, 2019. It contains four columns: timestamp, property, language, type. Every line in the dataset corresponds to the <strong>first-time action</strong> related to a language, i.e., the time at which the first translation of a label, description or an alias was made in a given language for a given property.</p> <p>Taking an example line from this dataset,</p> <p><em>2013-09-10T22:43:54Z,P856,en,label </em></p> <p>corresponds to the action that an <em>English</em> <em>label</em> of Property <em>P856</em> was added for the first time at <em>2013-09-10T22:43:54Z</em>.</p> <p>Following is a description of each column.</p> <p>1. timestamp: the time at which an action was made. For example, <em>2013-09-10T22:43:54Z</em></p> <p>2. property: Wikidata property identifier. It uses the P-number, For example, <em>P856</em></p> <p>3. language: the language in which a label/description/alias was first translated</p> <p>4. type: It could be one of the following values: label, description and alias</p>

opencc-by-4.0Jul 2019View details →
zenodo48/100

Wikipedia Multilingual Vandalism Detection Dataset

<p>This dataset accompanies a research paper that introduces a novel system designed to support the Wikipedia community in combating vandalism on the platform. The dataset has been prepared to enhance the accuracy and efficiency of Wikipedia patrolling in multiple languages.</p> <p>The release of this comprehensive dataset aims to encourage further research and development in vandalism detection techniques, fostering a safer and more inclusive environment for the Wikipedia community. Researchers and practitioners can utilize this dataset to train and validate their models for vandalism detection and contribute to improving online platforms' content moderation strategies.</p> <p><strong>Dataset Details:</strong></p> <ul> <li><strong>Number of Languages:</strong>&nbsp;47</li> <li><strong>Observation period:&nbsp;</strong>6 months training, one week hold-out testing</li> <li><strong>Use Case:</strong>&nbsp;The dataset is primarily intended for training and evaluating vandalism detection systems.</li> <li><strong>Features:</strong>&nbsp;Each record characterizes the corresponding revision of the Wikipedia page, including revision metadata, user details, text inserted, removed, or changed, and corresponding MLMs-based features.&nbsp;</li> <li><strong>Data Filtering and Feature Engineering:</strong>&nbsp;Advanced filtering and feature engineering techniques were applied to ensure the dataset's quality and relevance for effectively training the vandalism detection system.</li> <li><strong>Files:&nbsp;</strong>Training and hold-out testing datasets of anonymous and all users.&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>Related paper citation:</strong></p> <pre><code>@inproceedings{10.1145/3580305.3599823, author = {Trokhymovych, Mykola and Aslam, Muniza and Chou, Ai-Jou and Baeza-Yates, Ricardo and Saez-Trumper, Diego}, title = {Fair Multilingual Vandalism Detection System for Wikipedia}, year = {2023}, isbn = {9798400701030}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3580305.3599823}, doi = {10.1145/3580305.3599823}, abstract = {This paper presents a novel design of the system aimed at supporting the Wikipedia community in addressing vandalism on the platform. To achieve this, we collected a massive dataset of 47 languages, and applied advanced filtering and feature engineering techniques, including multilingual masked language modeling to build the training dataset from human-generated data. The performance of the system was evaluated through comparison with the one used in production in Wikipedia, known as ORES. Our research results in a significant increase in the number of languages covered, making Wikipedia patrolling more efficient to a wider range of communities. Furthermore, our model outperforms ORES, ensuring that the results provided are not only more accurate but also less biased against certain groups of contributors.}, booktitle = {Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining}, pages = {4981&ndash;4990}, numpages = {10}, location = {Long Beach, CA, USA}, series = {KDD '23} }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2023View details →
zenodo44/100

CLDF dataset accompanying Chechuro et al.'s "Small-scale multilingualism through the prism of lexical borrowing" from 2021

<p>Cite the source of the dataset as:</p> <blockquote> <p>Chechuro et al. (2021) Small-scale multilingualism through the prism of lexical borrowing. International Journal of Bilingualism.</p> </blockquote>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Gado2: multilingual newspapers from the Netherlands Indies

<p>This Handwritten Text Recognition (HTR) xml-page file dataset contains the ground truths of the Gado2 named entity processing application for newspapers from the Netherlands Indies and Indonesia, see: https://github.com/KBNLresearch/gado2. Optical Character Recognition (OCR) resulted in high Character Error Rates (CER) due to the inferior quality of many scans. In contrast, HTR led to CERs below 0.5 percent thus increasing the efficiency of the NER engine. All uploaded files are free of errors and fully tagged. A relevant knowledge base of Indonesian persons, places and organisations is attached in json format for entity linking.</p>

opencc-by-4.0May 2021View details →
zenodo44/100

A Multilingual Dataset of COVID-19 Vaccination Attitudes on Twitter

<p>This dataset consists of the IDs of 2,198,090 tweets collected from Western Europe,&nbsp;of which 17,934 are annotated with&nbsp;labels indicating the originators&#39; affective vaccination stances,&nbsp;including Positive (PO), Negative (NG), Positive but dissatisfaction (PD), Neutral (NE) and Off-topic (OT).</p> <p>all_tweets.txt contains all the ids of the collected tweets, annotated_tweets.txt&nbsp;contains the ids of the annotated tweets and the categories they are annotated to.</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Common Phone: A Multilingual Dataset for Robust Acoustic Modelling

<p><em>Release Date: 17.01.22</em></p> <p><strong>Welcome to&nbsp;Common Phone 1.0</strong></p> <p><strong>Legal Information</strong></p> <p><em>Common Phone</em>&nbsp;is a subset of the&nbsp;<em>Common Voice</em>&nbsp;corpus collected by&nbsp;<em>Mozilla Corporation</em>. By using&nbsp;<em>Common Phone</em>, you agree to the&nbsp;<a href="https://commonvoice.mozilla.org/en/terms">Common Voice Legal Terms</a>.&nbsp;<em>Common Phone</em>&nbsp;is maintained and distributed by speech researchers at the&nbsp;<a href="https://lme.tf.fau.de/">Pattern Recognition Lab</a>&nbsp;of Friedrich-Alexander-University Erlangen-Nuremberg (<a href="https://www.fau.de/">FAU</a>) under the&nbsp;<a href="https://creativecommons.org/publicdomain/zero/1.0/">CC0 license</a>.</p> <p>Like for&nbsp;<em>Common Voice</em>, you must not make any attempt to identify speakers that contributed to&nbsp;<em>Common Phone</em>.</p> <p><strong>About&nbsp;<em>Common Phone</em></strong></p> <p>This corpus aims to provide a basis for Machine Learning (ML) researchers and enthusiasts to train and test their models against a wide variety of speakers, hardware/software ecosystems and acoustic conditions to improve generalization and availability of ML in real-world speech applications.<br> The current version of&nbsp;<em>Common Phone</em>&nbsp;comprises 116,5 hours of speech samples, collected from 11.246 speakers in 6 languages:</p> <table align="center"> <thead> <tr> <th> <p><strong>Language</strong></p> </th> <th> <p><strong>Speakers</strong></p> </th> <th> <p><strong>Hours</strong></p> </th> </tr> </thead> <tbody> <tr> <td>&nbsp;</td> <td> <p><code>train</code>&nbsp;/&nbsp;<code>dev</code>&nbsp;/&nbsp;<code>test</code></p> </td> <td> <p><code>train</code>&nbsp;/&nbsp;<code>dev</code>&nbsp;/&nbsp;<code>test</code></p> </td> </tr> <tr> <td> <p>English</p> </td> <td> <p>4716&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;771&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;774</p> </td> <td> <p>14.1&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.3&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.3</p> </td> </tr> <tr> <td> <p>French</p> </td> <td> <p>796&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;138&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;135</p> </td> <td> <p>13.6&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.3&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.2</p> </td> </tr> <tr> <td> <p>German</p> </td> <td> <p>1176&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;202&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;206</p> </td> <td> <p>14.5&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.5&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.6</p> </td> </tr> <tr> <td> <p>Italian</p> </td> <td> <p>1031&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;176&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;178</p> </td> <td> <p>14.6&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.5&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.5</p> </td> </tr> <tr> <td> <p>Spanish</p> </td> <td> <p>508&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;88&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;91</p> </td> <td> <p>16.5&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;3.0&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;3.1</p> </td> </tr> <tr> <td> <p>Russian</p> </td> <td> <p>190&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;34&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;36</p> </td> <td> <p>12.7&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.6&nbsp;&nbsp;&nbsp;/&nbsp;&nbsp;&nbsp;2.8</p> </td> </tr> <tr> <td> <p><strong>Total</strong></p> </td> <td> <p>8417&nbsp;&nbsp;/&nbsp;&nbsp;1409&nbsp;&nbsp;/&nbsp;&nbsp;1420</p> </td> <td> <p>85.8&nbsp;&nbsp;/&nbsp;&nbsp;15.2&nbsp;&nbsp;/&nbsp;&nbsp;15.5</p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>Presented&nbsp;<code>train</code>,&nbsp;<code>dev</code>&nbsp;and&nbsp;<code>test</code>&nbsp;splits are&nbsp;<strong>not identical</strong>&nbsp;to those shipped with&nbsp;<em>Common Voice</em>. Speaker separation among splits was realized by only using those speakers that had provided age and gender information. This information can only be provided as a registered user on the website. When logged in, the session ID of contributed recordings is always linked to your user, thus we could easily link recordings to individual speakers. Keep in mind this would not be possible for unregistered users, as their session ID changes if they decide to contribute more than once.<br> During speaker selection, we considered that some speakers had contributed to more than one of the six&nbsp;<em>Common Voice</em>&nbsp;datasets (one for each language). In&nbsp;<em>Common Phone</em>, a speaker will only appear in one language.<br> The dataset is structured as follows:</p> <ul> <li>Six top-level directories, one for each language.</li> <li>Each language folder contains: <ul> <li>[train|dev|test].csv files listing audio files, respective speaker ID and plain text transcript.</li> <li>meta.csv provides speaker information: age group, gender, language, accent (if available) and which of the three splits this speaker was assigned to. File names match corresponding audio file names except their extension.</li> <li>/grids/ contains phonetic transcription for every audio file in Praat TextGrid format.</li> <li>/mp3/ contains audio files in mp3, identical to those of&nbsp;<em>Common Voice</em>, e.g., sampling rates have been preserved and may vary for different files.</li> <li>/wav/ contains raw audio files in 16 bits/sample, 16 kHz single channel. They had been created from the original mp3 audios. We provide them for convenience, keep in mind that their source had undergone MP3-compression.</li> </ul> </li> </ul> <p><strong>Where does the phonetic annotation come from?</strong></p> <p>Phonetic annotation was computed via&nbsp;<a href="https://clarin.phonetik.uni-muenchen.de/BASWebServices/interface/Pipeline">BAS Web Services</a>. We used the regular Pipeline (G2P-MAUS) without ASR to create an alignment of text transcripts with audio signals. We chose International Phonetic Alphabet (IPA) output symbols as they work well even in a multi-lingual setup.&nbsp;<em>Common Phone</em>&nbsp;annotation comprises 101 phonetic symbols, including silence.</p> <p><strong>Why&nbsp;<em>Common Phone</em>?</strong></p> <ul> <li>Large number of speakers and varying acoustic conditions to improve robustness of ML models</li> <li>Time-aligned IPA phonetic transcription for every audio sample</li> <li>Gender-balanced and age-group-matched (equal number of female/male speakers in every age group)</li> <li>Support for six different languages to leverage multi-lingual approaches</li> <li>Original MP3 files plus standard WAVE files</li> </ul> <p><strong>Is there any publication available?</strong></p> <p><em>Yes, a paper describing Common Phone in detail is currently under revision for LREC </em><em>2022. You can access a pre-print version on arXiv entitled &ldquo;<a href="https://arxiv.org/abs/2201.05912">Common Phone: A Multilingual Dataset for Robust Acoustic Modelling</a>&rdquo;.</em></p>

opencc-zeroJan 2022View details →
zenodo44/100

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

<p>We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria&mdash;Hausa, Igbo, Nigerian-Pidgin, and Yor&ugrave;b&aacute;&mdash;consisting of around 30,000 annotated tweets per language (except for Nigerian-Pidgin), including a significant fraction of code-mixed tweets.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

A Dataset of Multilingual Facebook Comments on Moros and Armed Conflict in the Southern Philippines

<p>This dataset is a collection of 12,478 social media comments found on the official Facebook pages of ten Philippine newspapers, T<em>he Philippine Daily Inquirer, Manila Bulletin, The Philippine Star, The Manila Times, Sunstar Cebu, Sunstar Davao, Cebu Daily News, The Freeman, Sunstar Davao, MindaNews, </em>and <em>The Mindanao Times</em>, spanning the years 2015, 2017 and 2019. The comments contain terms related to the Moro identity and the Mamasapano Clash, the Marawi Siege and the establishment of BARMM in the southern Philippines, allowing researchers to study semantic fields with regard to Muslims and the relationship between the texts and the source newspaper, their region of origin, and political administration, among other variables. All comments in the dataset were downloaded through Facebook's Graph API via Facepager (J&uuml;nger &amp; Keyling, 2019).</p> <p>One CSV file (MMB151719SOCMED_v2.csv) is provided, along with a codebook that contains descriptions of the variables and codes used in the CSV file, and a Readme document with a changelog.&nbsp;</p> <p>Each social media comment is annotated with the following metadata:&nbsp;</p> <ol> <li><strong>object_id</strong>: identifier associated with the comment;</li> <li><strong>message</strong>: the textual string of the comment;</li> <li><strong>message_proc:</strong> the textual string of the comment after pre-processing;</li> <li><strong>lang_label:</strong> categorical value for the language of the comment (Tagalog (Filipino), Cebuano, English, Taglish, Bislog, Bislish, Trilingual or Other);</li> <li><strong>from_name</strong>:&nbsp; identifier of public pages (not profiles of individuals) leaving comments (NaN for profiles of individuals, 'NAME' for public pages besides the newspapers, otherwise, the page name of the newspaper);</li> <li><strong>created_time</strong>: Facebook Graph API's-generated string for the date and time the comment was posted;&nbsp; &nbsp;</li> <li><strong>month_year</strong>: categorical value in the form string+YY (e.g. Jun-15) of the month and year when the comment was posted;</li> <li><strong>year</strong>: numerical value in the form YY;&nbsp;</li> <li><strong>newspaper</strong>: categorical value for the newspaper Facebook page under which the comment was found;</li> <li><strong>corpus</strong>: categorical value for comments from the main corpus or the side (control) corpus;</li> <li><strong>administration</strong>: categorical value for political administration (pbsa = President Benigno Aquino III, prrd = President Rodrigo Roa Duterte);</li> <li><strong>count</strong>: numerical value referring to the number of string sequences without spaces;</li> </ol> <p>The dataset may only be used for non-commercial purposes and is licensed under the&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/" target="_blank" rel="noopener">CC BY-NC-SA 4.0 DEED.</a></p> <p>_____________________________________________________________________________________</p> <p>V2 - 05/06/2024</p> <p>Corrections</p> <ul> <li>Corrections made to region to include Luzon, Visayas and Mindanao (as opposed to Mindanao, non-Mindanao);</li> <li>Corrections made to administration coding.&nbsp;</li> </ul> <p>&nbsp;</p> <p>This dataset is described by:&nbsp;</p> <p><span>Cruz, F. A. (2024). </span><span>A Multilingual Collection of Facebook Comments on the Moro Identity and Armed Conflict in the Southern Philippines<em>. Journal of Open Humanities Data, 10</em>(1), 41. </span><span>DOI: <a href="https://doi.org/10.5334/johd.219" target="_blank" rel="noopener">https://doi.org/10.5334/johd.219</a></span></p> <p>&nbsp;</p> <p><strong>Bibiliography</strong></p> <div> <div>J&uuml;nger, J., &amp; Keyling, T. (2019). <em>Facepager: An application for automated data retrieval on the web</em> (4.5.3) [Computer software]. <a href="https://github.com/strohne/Facepager/">https://github.com/strohne/Facepager/</a></div> <div>&nbsp;</div> <div>&nbsp;</div> </div>

openDec 2023View details →
zenodo44/100

Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse

<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon.&nbsp;</p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts.&nbsp;</p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website:&nbsp;<a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a>&nbsp;on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint:&nbsp;<a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO:&nbsp;<a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2:&nbsp;<a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3:&nbsp;<a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH:&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a>&nbsp;should not be used as a relation. We have replaced it with&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use&nbsp;<a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of&nbsp;<a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future:&nbsp;<a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesj&ouml;, QLIT, G&ouml;teborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kristr&ouml;m, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p>&nbsp;</p> <p>Thank you very much for your interest in our project!</p>

opengpl-3.0-or-laterJul 2024View details →
zenodo44/100

MaSS - Multilingual corpus of Sentence-aligned Spoken utterances

<p><strong>Abstract</strong></p> <p>The CMU Wilderness Multilingual Speech Dataset is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) models for potentially 700 languages. However, the fact that the source content (the Bible), is the same for all the languages is not exploited to date. Therefore, this article proposes to add multilingual links between speech segments in different languages, and shares a large and clean dataset of 8,130 para-lel spoken utterances across 8 languages (56 language pairs).We name this corpus MaSS (Multilingual corpus of Sentence-aligned Spoken utterances). The covered languages (Basque, English, Finnish, French, Hungarian, Romanian, Russian and Spanish) allow researches on speech-to-speech alignment as well as on translation for syntactically divergent language pairs. The quality of the final corpus is attested by human evaluation performed on a corpus subset (100 utterances, 8 language pairs).</p> <p><a href="https://arxiv.org/pdf/1907.12895.pdf">Paper </a>| <a href="https://github.com/getalp/mass-dataset">GitHub Repository</a>&nbsp;containing&nbsp;the scripts needed to build the data set from scratch (if needed)</p> <p><strong>Project structure</strong></p> <p>This repository contains 8 Numpy files, one for each featured language, pickled with Python 3.6. Each line corresponds to the spectrogram of the file mentioned in the file <em>verses.csv</em>. There is a direct mapping between the ID of the verse and its index in the list (thus verse with ID 5634 is located at index 5634 in the Numpy file). Verses not available for a given language (as stated by the value &quot;Not Available&quot; in the CSV file) are represented by empty lists in the Numpy files, thus ensuring a perfect verse-to-verse alignement between each file.</p> <p>Spectrogram were extracted using Librosa with the following parameters:</p> <pre><code>Pre-emphasis = 0.97 Sample rate = 16000 Window size = 0.025 Window stride = 0.01 Window type = 'hamming' Mel coefficients = 40 Min frequency = 20</code></pre> <p>&nbsp;</p>

openmit-licenseJul 2019View details →
zenodo44/100

Dataset for the paper "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia"

<p>Dataset for the EMNLP'24 Main conference paper titled "Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia".</p>

opencc-by-sa-4.0Oct 2024View details →
zenodo44/100

MultiSubs: A Large-scale Multimodal and Multilingual Dataset

<p>MultiSubs&nbsp;is a dataset of multilingual subtitles gathered from&nbsp;<a href="https://opus.nlpl.eu/OpenSubtitles.php">the OPUS OpenSubtitles dataset</a>,&nbsp;which in turn was sourced from <a href="http://www.opensubtitles.org/">opensubtitles.org</a>. We have supplemented some text fragments (visually salient nouns in this release) within the subtitles with web images, where the word sense of the fragment has been disambiguated using a cross-lingual approach.&nbsp;</p> <p>Please refer to our&nbsp;paper for a more detailed description of the dataset:</p> <p>Josiah Wang, Pranava Madhyastha, Josiel Figueiredo, Chiraag Lala, Lucia Specia (2021). <a href="https://arxiv.org/abs/2103.01910">MultiSubs: A Large-scale Multimodal and Multilingual Dataset</a>. CoRR, abs/2103.01910. Available at: <a href="https://arxiv.org/abs/2103.01910">https://arxiv.org/abs/2103.01910</a></p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

CLDF dataset accompanying List and Forkel's "Borrowing Detection in Multilingual Wordlists" from 2021

<p>Cite the source of the dataset as:</p> <blockquote> <p>List, Johann-Mattis and Forkel, Robert (2021): Automated identification of borrowings in multilingual wordlists. Leipzig: Max Planck Institute for Evolutionary Anthropology.</p> </blockquote>

opencc-by-4.0Jun 2021View details →
zenodo44/100

I-MSV 2022: Indic-Multilingual and Multi-sensor Speaker Verification Challenge

<p><strong>Dear Users,</strong></p> <p><strong>Data is password protected, to get password all you need to do is register using below link. Note that data is free of Cost&nbsp;</strong></p> <p><a href="https://forms.gle/1gsVhJaJYT4mBp83A">Click here for Registration</a></p> <p>Speaker Verification (SV) is a task to verify the claimed identity of the claimant using his/her voice sample. Though there exists an ample amount of research in SV technologies, the development concerning a multilingual conversation is limited. In a country like India, almost all the speakers are polyglot in nature. Consequently, the development of a Multilingual SV (MSV) system on the data collected in the Indian scenario is more challenging. With this motivation, the Indic- Multilingual Speaker Verification (I-MSV) Challenge 2022 has been designed for understanding and comparing the state of-the-art SV techniques. For the challenge, approximately 100 hours of data spoken by 100 speakers has been collected using 5 different sensors in 13 Indian languages. The data is divided into development, training, and testing sets and has been made publicly available for further research. The goal of this challenge is to make the SV system robust to language and sensor variations between enrollment and testing. In the challenge, participants were asked to develop the SV system in two scenarios, viz. constrained and unconstrained. The best system in the constrained and unconstrained scenario achieved a performance of 2.12% and 0.26% in terms of Equal Error Rate (EER), respectively.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

MOL - Multilingual Offensive Lexicon

<p>MOL - Multilingual Offensive Lexicon is a specialized lexicon for abusive language detection for low-resource languages. It consists of 1,000 explicit and implicit terms and expressions with pejorative connotations, manually identified by a specialist and annotated by three different experts with contextual information archiving high annotator human agreement (73% Kappa).</p> <p>Each term and expression from MOL contains a binary class: context-dependent offensiveness and context-independent offensiveness. For example, the term ``hypocrite'' is classified as context-independent offensiveness, since it is mostly found in the pejorative contexts of use. On the other hand, the term &nbsp;(``worm'') is classified as context-dependent offensiveness because it also may be found in both pejorative and non-pejorative contexts of use such as ``Politicians are like society worms'' and ``Reduces virus, worm, and unwanted access threats''.</p> <p>Finally, terms that showed strong potential to indicate hate speech targets were also annotated. For instance, ``slut'' and ``Jews from hell'' may indicate sexist and antisemitism comments. Finally, these terms and expressions, originally written in Portuguese, were manually translated by native speakers in English, Spanish, German, French, and Turkish.&nbsp;</p>

openother-openMar 2023View details →
zenodo44/100

Subset of 'MLSUM: The Multilingual Summarization Corpus' for constraints annotation experiment

<p><strong>[EN] Subset of &#39;MLSUM: The Multilingual Summarization Corpus&#39; for constraints annotation experiment.</strong></p> <ul> <li><strong>Description</strong>: MLSUM is a dataset of newspappers articles aimed at training summaring model. We use it for a constraints annotation experiment on newspapper titles according to their topic classification.</li> <li><strong>Content</strong>: For constraints annotation experiment based on data similarity, this dataset have been subsetted (randomly pick 75 articles in the following 14 most used topics: &#39;economie&#39;, &#39;politique&#39;, &#39;sport&#39;, &#39;planete&#39; (renamed in &#39;ecologie&#39;), &#39;sciences&#39;, &#39;police-justice&#39;, &#39;disparitions&#39;, &#39;emploi&#39;, &#39;sante&#39;, &#39;musiques&#39;, &#39;arts&#39;, &#39;educations&#39;, &#39;climat&#39; (renamed in &#39;meteo&#39;), &#39;immobilier&#39;) and filtered (keep articles that have an obvious topics regarding their titles, without their bodies). Two reviewers have working on this task in order to limit the subjectivity of the filtering. This subsetted dataset is used (1) to estimate needed time to annotate titles similarity with constraints (MUST-LINK, CANNOT-LINK) and (2) to test interactive clustering methodology (constraints annotation and constrained clustering).</li> <li><strong>Origin</strong>: The dataset is bassed on the original &#39;MLSUM: The Multilingual Summarization Corpus&#39; dataset (https://doi.org/10.48550/arXiv.2004.14900).</li> </ul> <p><br> <strong>[FR] Echantillon de &#39;MLSUM: The Multilingual Summarization Corpus&#39; pour une exp&eacute;rience&nbsp;d&#39;annotation de contraintes.</strong></p> <ul> <li><strong>Description </strong>: MLSUM est un ensemble de donn&eacute;es d&#39;articles de journaux destin&eacute;s &agrave; l&#39;entra&icirc;nement d&#39;un mod&egrave;le de r&eacute;sum&eacute; automatique. Nous l&#39;utilisons pour une exp&eacute;rience d&#39;annotation de contraintes sur des titres de journaux en fonction de leur classification th&eacute;matique.</li> <li><strong>Contenu </strong>: Pour une exp&eacute;rience d&#39;annotation de contraintes bas&eacute;e sur la similarit&eacute; des donn&eacute;es, cet ensemble de donn&eacute;es a &eacute;t&eacute; &eacute;chantillonn&eacute; (s&eacute;lectionner au hasard de 75 articles dans les 14 sujets les plus utilis&eacute;s&nbsp;: &#39;&eacute;conomie&#39;, &#39;politique&#39;, &#39;sport&#39;, &#39;plan&egrave;te&#39; (renomm&eacute; en &laquo; &eacute;cologie &raquo;). ), &#39;sciences&#39;, &#39;police-justice&#39;, &#39;disparitions&#39;, &#39;emploi&#39;, &#39;sante&#39;, &#39;musiques&#39;, &#39;arts&#39;, &#39;&eacute;ducations&#39;, &#39;climat&#39; (renomm&eacute; en &#39;meteo&#39;), &#39;immobilier&#39; ) et filtr&eacute; (conserver les articles qui ont un sujet &eacute;vident par rapport &agrave; leur titre, sans leur corps). Deux relecteurs ont travaill&eacute; sur cette t&acirc;che afin de limiter la subjectivit&eacute; du filtrage. Ce sous-ensemble de donn&eacute;es est utilis&eacute; (1) pour estimer le temps n&eacute;cessaire pour annoter la similarit&eacute; des titres avec des contraintes (MUST-LINK, CANNOT-LINK) et (2) pour tester la m&eacute;thodologie de clustering interactif (annotation de contraintes et clustering contraint).</li> <li><strong>Origine </strong>: L&#39;ensemble de donn&eacute;es est bas&eacute; sur l&#39;ensemble de donn&eacute;es original &#39;MLSUM : The Multilingual Summarization Corpus&#39; (https://doi.org/10.48550/arXiv.2004.1490).</li> </ul>

openmit-licenseOct 2023View details →
zenodo44/100

Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper

<p>Corpora used in the publication:</p> <ul> <li>Cristina Espa&ntilde;a-Bonet. 2023. <strong>Multilingual Coarse Political Stance Classification of Media. The Editorial Line of a ChatGPT and Bard Newspaper. </strong>In <em>Findings of the Association for Computational Linguistics: EMNLP 2023</em>, Singapore. Pages 11757&ndash;11777. Association for Computational Linguistics.</li> </ul> <p>Three corpora are included:</p> <ol> <li>Newspaper articles extracted from the OSCAR corpus in English, German, Spanish and Catalan automatically annotated for political stance (left vs right) and topic</li> <li>Newspaper-like article generations by different versions of ChatGPT for 101 topics in the 4 languages</li> <li>Newspaper-like article generations by Bard for 101 topics in the 4 languages</li> </ol> <p>See the README file and the original article for further details.</p>

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record