Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
135
datasets available to search
ShareScore release 0.7.1
Dataset results
135 results for “multilingual”
GeoCoV19: A Dataset of Hundreds of Millions of Multilingual COVID-19 Tweets with Location Information
<p>We present GeoCoV19, a large-scale Twitter dataset related to the ongoing COVID-19 pandemic. The dataset has been collected over a period of 90 days from February 1 to May 1, 2020 and consists of more than 524 million multilingual tweets. As the geolocation information is essential for many tasks such as disease tracking and surveillance, we employed a gazetteer-based approach to extract toponyms from user location and tweet content to derive their geolocation information using the Nominatim (Open Street Maps) data at different geolocation granularity levels. In terms of geographical coverage, the dataset spans over 218 countries and 47K cities in the world. The tweets in the dataset are from more than 43 million Twitter users, including around 209K verified accounts. These users posted tweets in 62 different languages.</p>
Multilingual MigrationsKB: A Mulitlingual Knowledge Base of Migration related annotated Tweets
<p><strong>Multilingual MigrationskB (MGKB) </strong>is a mulitlingual extended version of English <a href="https://zenodo.org/record/5206820#.YRqF1nUza0o">MGKB</a>. The tweets geotagged with Geo location from 32 European Countries (<em><strong>Austria, Belgium, Bulgaria, Croatia, Cyprus, Czech, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, Netherlands, Poland, Portugal, Romania, Slovakia, Slovenia, Spain, Sweden, Iceland, Liechtenstein, Norway, Switzerland, the United Kingdom</strong></em>) are extracted and filtered by 11 languages (<em><strong>English, French, Finnish, German, Greek, Dutch, Hungarian, Italian, Polish, Spain, Swedish</strong></em>). Metadata information about the tweets, such as <strong>Geo information (place name, coordinates, country code)</strong> are included. <strong>MGKB</strong> contains <strong>sentiments, offensive and hate speeches, topics, hashtags, user mentions</strong> in RDF format. The schema of <strong>MGKB</strong> is an extension of TweetsKB for migration related information. Moreover, to associate and represent the potential economic and social factors driving the migration flows, the data from <a href="https://ec.europa.eu/eurostat/web/main/home">Eurostat</a> and <a href="https://spec.edmcouncil.org/fibo/ontology/">FIBO</a> ontology was used. To represent multilinguality, the<a href="https://www.cidoc-crm.org/"> CIDOC Conceptual Reference Model (CIDOC-CRM)</a> is used. The extracted economic indicators, i.e., GDP Growth Rate, Total Unemployment Rate, Youth Unemployment Rate, Long-term Unemployment Rate and Income per househould, are connected with each tweet in RDF using geographical and temporal dimensions. </p> <p>For this version, the Multilingual MGKB is delivered separated by year. The extracted topic words are also published.</p> <p>Code: <a href="https://github.com/migrationsKB/MRL">https://github.com/migrationsKB/MRL</a></p> <p>Please contact Yiyi Chen (yiyi.chen@partner.kit.edu) for pretrained models (Sentiment analysis/hate speech detection/ETM) if necessary.</p> <p> </p> <p> </p>
Multilingual Persuasion Dataset
<p>This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1. Some of the dialogue lines are marked as persuasive (which is when the player character is attempting a Persuade skill check.) </p> <p>If you use this data, please cite:</p> <p>Pöyhönen, T., Hämäläinen, M., Alnajjar, K. (2022) "Multilingual Persuasion Detection: Video Games as an Invaluable Data Source for NLP" DiGRA '22 - Proceedings of the 2022 DiGRA International Conference</p> <p>Code:</p> <p>https://github.com/Teemursu/multilingual_persuasion_dataset</p>
VivesDebate: A New Annotated Multilingual Corpus of Argumentation in a Debate Tournament
<p>The application of the latest Natural Language Processing breakthroughs in computational argumentation has shown promising results which have raised the interest in this area of research. However, the available corpora with argumentative annotations are often limited to a very specific purpose or are not of adequate size to take advantage of state-of-the-art deep learning techniques (e.g., deep neural networks). In this paper, we present VivesDebate, a large, richly annotated, and versatile professional debate corpus for computational argumentation research. The corpus has been created from 29 transcripts of a debate tournament in Catalan and has been machine-translated into Spanish and English. The annotation contains argumentative propositions, argumentative relations, debate interactions, and professional evaluations of the arguments and argumentation. The presented corpus can be useful for research on a heterogeneous set of computational argumentation underlying tasks such as argument mining, argument analysis, argument evaluation, or argument generation among others. All this makes VivesDebate a valuable resource for computational argumentation research within the context of massive corpora aimed at Natural Language Processing tasks.</p>
WAGEINDICATOR MULTILINGUAL TASKS DATABASE
<p>The International Standard Classification of Occupations (ISCO–08) is maintained by the International Labour Organisation (ILO). ISCO-08 is increasingly being adopted worldwide. ISCO-08 defines a job as a bundle of tasks and duties performed by one person. Homogeneity of tasks define an occupation. ILO's coding index (2012) provides lists of task sets for 427 occupational titles at 4-digit level, varying between 2 and 14 tasks per title. In total 3,264 tasks are available, on average 7.6 tasks per occupational title. Almost all tasks are occupation-specific, and only a limited number of tasks are similar across more than one occupation.</p> <p>The ISCO-08 lists of tasks per 4-digit occupation offer a great opportunity to investigate the homogeneity of tasks within an occupation across survey respondents, either in one country or in multiple countries. However, due to technical limitations this assumption has hardly ever been tested empirically, but recently web surveys offer possibilities to do so. For this purpose a WAGEINDICATOR MULTILINGUAL TASKS DATABASE has been developed. The starting point for this database were the task descriptions provided by ILO's English ISCO-08 coding index for occupations at the most detailed ISCO-08 4-digit occupational groups (ILO, 2012). Task descriptions for other languages have been added from two sources, namely the task descriptions as included in coding indexes of national statistical offices, and from translations of ILO's task descriptions provided by WageIndicator Foundation. For the descriptions in the coding indexes, the tasks have been checked for correspondence with the ILO index. Please note that in these countries the TASKS DATABASE does not contain translations of the tasks in the ILO coding index, but the ‘translation’ provided by the national statistical offices. In most cases these comprise of a literal translation, but in other cases tasks were dropped if the office considered these not to be part of an occupation, new tasks were added if applicable, or task descriptions were adapted to national practices. In the latter cases, the ILO task and the national task may deviate slightly from each other.</p> <p>The resulting WAGEINDICATOR MULTILINGUAL TASKS DATABASE consists of 3,264 tasks for 427 ISCO-08 4-digit occupations in 19 languages: Albanian, Bosnian, Bulgarian, Czech, Dutch, English (British), Estonian, Finnish, French, German, Indonesian, Lithuanian, Mongolian, Norwegian, Portuguese, Russian, Slovakian, Spanish and Turkish. Note that in the Finnish and the Norwegian columns approx. 20% of the tasks have no content, because the statistical office did not provide these in their coding indexes. Note that the Norwegian language is of moderate quality. For Mongolian 6% is not translated, for Indonesia and Albanian 1% is not translated, and for Bosnian, Bulgarian, Lithuanian, and Slovakian 1-19 tasks are not translated. The table hereafter details the sources used.</p> <table> <tbody> <tr> <td> <p><strong>COUNTRY</strong></p> </td> <td> <p><strong>SOURCE</strong></p> </td> </tr> <tr> <td> <p>Albania</p> </td> <td> <p>PËR MIRATIMIN E LISTËS KOMBËTARE TË PROFESIONEVE (LKP), TË RISHIKUAR, VENDIM Nr. 514, datë 20.9.2017</p> </td> </tr> <tr> <td> <p>Austria</p> </td> <td> <p>Übersicht über die Gliederung der ÖISCO, OEISCO08_Erlaeuterungen</p> </td> </tr> <tr> <td> <p>Bosnian</p> </td> <td> <p>Translations</p> </td> </tr> <tr> <td> <p>Bulgaria</p> </td> <td> <p>Национална Класификация На Професиите И Длъжностите, 2011 г.</p> </td> </tr> <tr> <td> <p>Czech Republic</p> </td> <td> <p>Klasifikace zaměstnání - systematická část_1_20110909</p> </td> </tr> <tr> <td> <p>English (British)</p> </td> <td> <p>ILO. (2012). International Standard Classification of Occupations ISCO-08 Volume 1 Structure, Group Definitions And Correspondence Tables. Geneva: International Labour Office.</p> </td> </tr> <tr> <td> <p>Estonia</p> </td> <td> <p>Ametite klassifikaator 2008v1.5b</p> </td> </tr> <tr> <td> <p>Finland</p> </td> <td> <p>2011 Tilastokeskus – Statistikcentralen – Statistics Finland</p> </td> </tr> <tr> <td> <p>French</p> </td> <td> <p>Translations</p> </td> </tr> <tr> <td> <p>Indonesia</p> </td> <td> <p>Translations</p> </td> </tr> <tr> <td> <p>Lithuania</p> </td> <td> <p>LIETUVOS RESPUBLIKOS PROFESIJŲ KLASIFIKATORIAUS ATNAUJINTA ŠEŠIAŽENKLĖ STRUKTŪRA PAGAL ISCO-08, LPK pagal ISCO-08 kodas, PATVIRTINTA Lietuvos darbo rinkos mokymo tarnybos prie Socialinės apsaugos ir darbo ministerijos direktoriaus 2010 m. rugsėjo 30 d. įsakymu Nr. V(5)-167</p> </td> </tr> <tr> <td> <p>Mongolia</p> </td> <td> <p>Монгол Улсын Їндэсний, Ажил, Мэргэжлийн, Ангилал ба Тодорхойлолт ЇАМАТ-08 (ISCO - 08). Нийгмийн хамгаалал, хєдєлмєрийн яам Улаанбаатар. 2010 он</p> </td> </tr> <tr> <td> <p>Netherlands</p> </td> <td> <p>Translations</p> </td> </tr> <tr> <td> <p>Norway</p> </td> <td> <p>Notater 17/2011, Standard for yrkesklassifisering (STYRK-08) Statistisk sentralbyrå • Statistics Norway, Oslo–Kongsvinger</p> </td> </tr> <tr> <td> <p>Portugal</p> </td> <td> <p>Classificação Portuguesa das Profissões 2010, Edição 2011. Instituto Nacional de Estatística, Lisboa, Portugal</p> </td> </tr> <tr> <td> <p>Russian</p> </td> <td> <p>Translations</p> </td> </tr> <tr> <td> <p>Slovak Republic</p> </td> <td> <p>Štatistická Klasifikácia Zamestnaní Sk ISCO-08 Číselný kód Názov zamestnania</p> </td> </tr> <tr> <td> <p>Spain</p> </td> <td> <p>Clasificación Nacional de Ocupaciones 2011 (CNO2011), Notas Explicativas. INE (Statistics Spain), Enero 2012</p> </td> </tr> <tr> <td> <p>Turkey</p> </td> <td> <p>Uluslararası Standart Meslek Sınıflaması (ISCO-08)</p> </td> </tr> </tbody> </table> <p> </p> <p> ****</p>
Multilingual Speaker Anonymization Trials for CommonVoice and Multilingual LibriSpeech
<p>This dataset contains the speaker verification trial files of the evaluation data splits for Multilingual LibriSpeech (MLS) and CommonVoice (CV) that we propose in our paper "Probing the Feasibility of Multilingual Speaker Anonymization". <strong>The actual audio files are not included and have to be obtained separately</strong>, following the licenses of the respective corpus creators. The files in this dataset only contain the audio file IDs that can be used to prepare the data for the evaluation. </p> <h2>Data</h2> <p>All files contain the utterance IDs of the original MLS and CV corpora which makes it possible to align them to the audio files as provided by the dataset creators. If you use these datasets, you need to cite the original sources. We do not claim any rights to the audios.</p> <p>The dataset does not include all IDs of the corpora. For MLS, "dev" and "test" correspond to the dev and test splits as provided in the MLS corpus. For CV, the corpus was divided randomly into dev and test while ensuring no speaker overlap between both splits. We use the CV 16.1 version of the corpus and restrict it to validated audios where the user specified their gender as either female or male. Further information about the data restrictions can be found in the paper.</p> <p>Please note that we use the "client ID" of CV to distinguish between speakers. We acknowledge that this can lead to having the same speaker under different name multiple times in the corpus if they were assigned different client IDs. Based on our results for the original (non-anonymized) data, we believe this to have only little effect on our evaluation data.</p> <p>The following languages are included in this dataset: <strong>English (en), German (de), Dutch (nl), French (fr), Spanish (es), Italian (it), Portuguese (pt), Polish (pl), and Russian (ru).</strong></p> <h2>Structure</h2> <p>The directory contains separate folders for MLS and CV, with subfolders for each language. Each language subfolder contains 7 files:</p> <ul> <li>dev_enrolls</li> <li>dev_trials_f</li> <li>dev_trials_m</li> <li>test_enrolls</li> <li>test_trials_f</li> <li>test_trials_m</li> <li>utt2spk</li> </ul> <p>The file structures follow the evaluation data of the Voice Privacy Challenges (<a href="https://www.voiceprivacychallenge.org" rel="nofollow">https://www.voiceprivacychallenge.org</a>). The "enrolls" file contain the list of utterance IDs (i.e., audio files) that are used for enrollment of the speaker verification model. "trials_f" and "trials_m" correspond to the trial files for female and male speakers, respectively. Each line in a trial file consists of three constituents, separated by space: "enrollment speaker" "trial utterance" "target/nontarget". The last constituent signals whether the trial utterance was originally (i.e., before the anonymization) spoken by the enrollment speaker (target) or not (nontarget). The utt2spk file contains the true mapping between utterance and original speaker. This file is especially important for the CV corpus where we created new speaker names to replace the long client IDs, and where the speaker assignment is not visible in the file name.</p> <h2>Creation Process</h2> <p>For the preparation of the data into enrollment and trial subset, we tried to follow the dev and test files of the Voice Privacy Challenge 2022. This results in far more nontarget than target trials, and additional speakers in the trial set that are not contained in the enrollment set.</p> <h3>MLS</h3> <p>The MLS corpus (<a href="https://www.openslr.org/94/" rel="nofollow">available here</a>) already comes with a split into train / dev / test, which we reuse here. The data is further divided into 8 languages: en, de, nl, fr, es, it, pt and pl. We only take the dev and test sets for 6 languages (de, nl, fr, es, it, and pt). For en, we already have an alternative from the Voice Privacy Challenges based on the monolingual English LibriSpeech, so there is no need for another part from MLS. For pl, the MLS part only contains 2 speakers per gender and dev / test split, which is too small for effective speaker verification. <em>(Sidenote: Dutch is not significantly bigger with only 3 speakers per gender and split, but we decided to keep it anyway).</em> We do not further restrict the number of utterances or speaker in each language and dev / test set, which results in a large inbalance across languages. However, as given in the original MLS splits, the languages itself are balanced in terms of gender.</p> <h3>CV</h3> <p>Mozilla's CommonVoice data collection is significantly bigger than MLS, so we could select more speakers per language and make sure that the datasets per language were more or less balanced. We take the data from CV Version 16.1 (<a href="https://commonvoice.mozilla.org/en/datasets" rel="nofollow">available here</a>). Please note that users who <em>donated</em> their voice to the data collection can opt out of being included in the data at any point. <strong>This means that some speakers or utterances contained in our trial data might be missing in future downloads of the CV corpus.</strong> We further want to mention that we had to use the field <code>client ID</code> in the CV corpus for speaker assignment which is not fully accurate. The same speaker might end up with several client IDs if they are recording the utterances in different sessions. However, for the purpose of speaker anonymization, this issue is not as relevant as for pure speaker recognition.</p> <p>We use CV for all of our 9 languages. CV does not come with a division into train / dev / test splits, so we randomly sample speakers and utterances from it. For this sampling, we consider only speakers that have gender annotated as either female or male, and have recorded at least 50 validated audios. We further make sure that we have the same number of female as male speakers which leads for several languages to a large reduction in size, with most languages having far more male speakers in CV than female speakers. This leads especially to a smaller dataset for pl, for which only 14 speakers per gender and dev / test split are available. We randomly select at most 20 speakers per gender for each dev and test in each language, and randomly choose up to 70 utterances per speaker. As our lower bound for speaker selection was originally 50 utterances per speaker, this results in 50-70 utterances per speaker.</p> <h3>Separation into Enrollment and Trials</h3> <p>In MLS, we use all speakers for the enrollment and trial. The only exception is de, for which more speakers are available. In the MLS-de data, we reserve 5 speakers per gender and split for trial only, which creates some unseen distraction speakers in the trial data. In CV, we use 15 speakers per gender and split for enrollment (except for the smaller pl part, for which it is only 10), and also reserve up to 5 speakers per gender and split only for trial. Naturally, all enrollment speakers are also used in trial.</p> <p>15% of all utterances of a speaker (at least 5 utterances) are used as enrollment utterances, the rest for trial. All trial utterances are paired with each enrollment speaker of the respective gender. If an enrollment speaker is the actual speaker of that utterance, this is denoted as <em>target</em>, otherwise as <em>nontarget</em>.</p> <p>During the trials, the enrollment speaker is modeled as an average of the speaker embeddings of all enrollment utterances of that speaker.</p> <h2>Statistics</h2> <p>The following section displays the statistics for each dataset, language and dev / test split. In these statistics, female and male speakers are not distinguished, but the numbers are balanced for each subset.</p> <p>The following information is given for each dataset and language:</p> <ul> <li># speakers: number of speakers used for both enrollment and trial (50% female / 50% male)</li> <li># add.trial speakers: number of speakers additionally used only in trial</li> <li># enroll utts: total number of utterances used in enrollment (across all speakers)</li> <li># trial utts: total number of utterances used in trials (across all speakers)</li> <li># target trials: number of target trials (enrollment speaker == trial speaker)</li> <li># nontarget trials: number of nontarget trials (enrollment speaker != trial speaker)</li> <li># words: total number of words across all trial utterances (the WER is computed based on them)</li> <li># avg. utt length: average length of all utterances in the dataset, in seconds</li> </ul> <div> <h3>Development Data</h3> <p><strong>Total dataset statistics:</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>20</td> <td>10</td> <td>333</td> <td>3,136</td> <td>1,936</td> <td>29,424</td> <td>111,245</td> <td>14.90</td> </tr> <tr> <td> </td> <td>fr</td> <td>18</td> <td>0</td> <td>354</td> <td>2,062</td> <td>2,062</td> <td>16,496</td> <td>73,007</td> <td>15.03</td> </tr> <tr> <td> </td> <td>it</td> <td>10</td> <td>0</td> <td>183</td> <td>1,065</td> <td>1,065</td> <td>4,260</td> <td>34,636</td> <td>14.88</td> </tr> <tr> <td> </td> <td>es</td> <td>20</td> <td>0</td> <td>349</td> <td>2,059</td> <td>2,059</td> <td>18,531</td> <td>74,782</td> <td>14.95</td> </tr> <tr> <td> </td> <td>pt</td> <td>10</td> <td>0</td> <td>119</td> <td>707</td> <td>707</td> <td>2,828</td> <td>24,733</td> <td>15.90</td> </tr> <tr> <td> </td> <td>nl</td> <td>6</td> <td>0</td> <td>461</td> <td>2,634</td> <td>2,634</td> <td>5,268</td> <td>11,0384</td> <td>14.83</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>10</td> <td>279</td> <td>2,306</td> <td>1,691</td> <td>32,899</td> <td>22,394</td> <td>4.96</td> </tr> <tr> <td> </td> <td>de</td> <td>30</td> <td>9</td> <td>289</td> <td>2,396</td> <td>1,738</td> <td>34,202</td> <td>21,421</td> <td>5.16</td> </tr> <tr> <td> </td> <td>fr</td> <td>30</td> <td>10</td> <td>293</td> <td>2,432</td> <td>1,761</td> <td>34,719</td> <td>23,240</td> <td>4.93</td> </tr> <tr> <td> </td> <td>it</td> <td>30</td> <td>10</td> <td>286</td> <td>2,421</td> <td>1,722</td> <td>34,593</td> <td>23,745</td> <td>5.69</td> </tr> <tr> <td> </td> <td>es</td> <td>30</td> <td>10</td> <td>290</td> <td>2,401</td> <td>1,735</td> <td>34,280</td> <td>22,569</td> <td>5.23</td> </tr> <tr> <td> </td> <td>pt</td> <td>30</td> <td>10</td> <td>284</td> <td>2,398</td> <td>1,708</td> <td>34,262</td> <td>16,411</td> <td>3.97</td> </tr> <tr> <td> </td> <td>nl</td> <td>30</td> <td>10</td> <td>288</td> <td>2,401</td> <td>1,725</td> <td>34,290</td> <td>22,108</td> <td>4.51</td> </tr> <tr> <td> </td> <td>pl</td> <td>20</td> <td>8</td> <td>199</td> <td>1,739</td> <td>1,196</td> <td>16,194</td> <td>12,876</td> <td>4.39</td> </tr> <tr> <td> </td> <td>ru</td> <td>30</td> <td>10</td> <td>289</td> <td>2,429</td> <td>1,737</td> <td>34,698</td> <td>20,599</td> <td>5.24</td> </tr> </tbody> </table> <div> <p><strong>Dataset statistics per speaker (average)</strong></p> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.6</td> <td>104.5</td> <td>96.8</td> <td>1,471.2</td> <td>3,553.0</td> </tr> <tr> <td> </td> <td>fr</td> <td>19.7</td> <td>114.6</td> <td>114.6</td> <td>916.4</td> <td>4,350.0</td> </tr> <tr> <td> </td> <td>it</td> <td>18.3</td> <td>106.5</td> <td>106.5</td> <td>426.0</td> <td>4,405.5</td> </tr> <tr> <td> </td> <td>es</td> <td>17.4</td> <td>103.0</td> <td>103.0</td> <td>926.6</td> <td>4,300.5</td> </tr> <tr> <td> </td> <td>pt</td> <td>11.9</td> <td>70.7</td> <td>70.7</td> <td>282.8</td> <td>2,084.0</td> </tr> <tr> <td> </td> <td>nl</td> <td>76.8</td> <td>439.0</td> <td>439.0</td> <td>878.0</td> <td>45,515.0</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.3</td> <td>59.1</td> <td>56.4</td> <td>1,096.6</td> <td>627.0</td> </tr> <tr> <td> </td> <td>de</td> <td>9.6</td> <td>59.9</td> <td>57.9</td> <td>1,140.1</td> <td>492.5</td> </tr> <tr> <td> </td> <td>fr</td> <td>9.8</td> <td>60.8</td> <td>58.7</td> <td>1,157.3</td> <td>606.0</td> </tr> <tr> <td> </td> <td>it</td> <td>9.5</td> <td>60.5</td> <td>57.4</td> <td>1,153.1</td> <td>594.5</td> </tr> <tr> <td> </td> <td>es</td> <td>9.7</td> <td>60.0</td> <td>57.8</td> <td>1,142.7</td> <td>494.0</td> </tr> <tr> <td> </td> <td>pt</td> <td>9.5</td> <td>60.0</td> <td>56.9</td> <td>1,142.1</td> <td>308.0</td> </tr> <tr> <td> </td> <td>nl</td> <td>9.6</td> <td>60.0</td> <td>57.5</td> <td>1,143.0</td> <td>520.5</td> </tr> <tr> <td> </td> <td>pl</td> <td>10.0</td> <td>62.1</td> <td>59.8</td> <td>809.7</td> <td>514.5</td> </tr> <tr> <td> </td> <td>ru</td> <td>9.6</td> <td>60.7</td> <td>57.9</td> <td>1,156.6</td> <td>502.0</td> </tr> </tbody> </table> <h3>Test Data</h3> <p><strong>Total dataset statistics:</strong></p> <p> </p> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># speakers</th> <th># add. trial speakers</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> <th>avg. utt length</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>30</td> <td>0</td> <td>329</td> <td>3,065</td> <td>1,906</td> <td>28,744</td> <td>110,202</td> <td>15.18</td> </tr> <tr> <td> </td> <td>fr</td> <td>18</td> <td>0</td> <td>357</td> <td>2,069</td> <td>2,069</td> <td>16,552</td> <td>79,524</td> <td>14.94</td> </tr> <tr> <td> </td> <td>it</td> <td>10</td> <td>0</td> <td>185</td> <td>1,077</td> <td>1,077</td> <td>4,308</td> <td>34,796</td> <td>15.07</td> </tr> <tr> <td> </td> <td>es</td> <td>20</td> <td>0</td> <td>348</td> <td>2,037</td> <td>2,037</td> <td>18,333</td> <td>75,536</td> <td>15.11</td> </tr> <tr> <td> </td> <td>pt</td> <td>10</td> <td>0</td> <td>125</td> <td>746</td> <td>746</td> <td>2,984</td> <td>26,769</td> <td>15.47</td> </tr> <tr> <td> </td> <td>nl</td> <td>6</td> <td>0</td> <td>458</td> <td>2,617</td> <td>2,617</td> <td>5,234</td> <td>108,489</td> <td>14.96</td> </tr> <tr> <td>CV</td> <td>en</td> <td>30</td> <td>9</td> <td>289</td> <td>2,344</td> <td>1,733</td> <td>33,427</td> <td>22,560</td> <td>5.09</td> </tr> <tr> <td> </td> <td>de</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,713</td> <td>33,942</td> <td>21,492</td> <td>5.24</td> </tr> <tr> <td> </td> <td>fr</td> <td>30</td> <td>10</td> <td>291</td> <td>2,408</td> <td>1,745</td> <td>34,275</td> <td>22,887</td> <td>4.74</td> </tr> <tr> <td> </td> <td>it</td> <td>30</td> <td>10</td> <td>298</td> <td>2,444</td> <td>1,786</td> <td>34,874</td> <td>24,091</td> <td>5.37</td> </tr> <tr> <td> </td> <td>es</td> <td>30</td> <td>10</td> <td>284</td> <td>2,377</td> <td>1,703</td> <td>33,952</td> <td>22,655</td> <td>5.25</td> </tr> <tr> <td> </td> <td>pt</td> <td>30</td> <td>7</td> <td>290</td> <td>2,213</td> <td>1,743</td> <td>31,452</td> <td>15,727</td> <td>4.23</td> </tr> <tr> <td> </td> <td>nl</td> <td>30</td> <td>10</td> <td>283</td> <td>2,254</td> <td>1,704</td> <td>32,106</td> <td>19,913</td> <td>4.32</td> </tr> <tr> <td> </td> <td>pl</td> <td>20</td> <td>8</td> <td>196</td> <td>1,712</td> <td>1,181</td> <td>15,939</td> <td>13,809</td> <td>4.82</td> </tr> <tr> <td> </td> <td>ru</td> <td>30</td> <td>10</td> <td>294</td> <td>2,435</td> <td>1,759</td> <td>34,766</td> <td>20,509</td> <td>5.13</td> </tr> </tbody> </table> <div> <h5>Dataset statistics per speaker (average)</h5> </div> <table> <tbody> <tr> <th>Dataset</th> <th>Lang</th> <th># enroll utts</th> <th># trial utts</th> <th># target trials</th> <th># nontarget trials</th> <th># words</th> </tr> </tbody> <tbody> <tr> <td>MLS</td> <td>de</td> <td>16.4</td> <td>102.2</td> <td>95.3</td> <td>1,437.2</td> <td>4,047.0</td> </tr> <tr> <td> </td> <td>fr</td> <td>19.8</td> <td>114.9</td> <td>114.9</td> <td>919.6</td> <td>3,945.5</td> </tr> <tr> <td> </td> <td>it</td> <td>18.5</td> <td>107.7</td> <td>107.7</td> <td>430.8</td> <td>3,658.0</td> </tr> <tr> <td> </td> <td>es</td> <td>17.4</td> <td>101.8</td> <td>101.8</td> <td>916.6</td> <td>4,870.5</td> </tr> <tr> <td> </td> <td>pt</td> <td>12.5</td> <td>74.6</td> <td>74.6</td> <td>298.4</td> <td>3,141.5</td> </tr> <tr> <td> </td> <td>nl</td> <td>76.3</td> <td>436.2</td> <td>436.2</td> <td>872.3</td> <td>24,892.5</td> </tr> <tr> <td>CV</td> <td>en</td> <td>9.6</td> <td>60.1</td> <td>57.8</td> <td>1,114.2</td> <td>551.5</td> </tr> <tr> <td> </td> <td>de</td> <td>9.5</td> <td>59.4</td> <td>57.1</td> <td>1,131.4</td> <td>568.5</td> </tr> <tr> <td> </td> <td>fr</td> <td>9.7</td> <td>60.2</td> <td>58.2</td> <td>1,145.8</td> <td>618.0</td> </tr> <tr> <td> </td> <td>it</td> <td>9.9</td> <td>61.1</td> <td>59.5</td> <td>1,162.5</td> <td>612.0</td> </tr> <tr> <td> </td> <td>es</td> <td>9.5</td> <td>59.4</td> <td>56.8</td> <td>1,131.7</td> <td>610.0</td> </tr> <tr> <td> </td> <td>pt</td> <td>9.7</td> <td>59.7</td> <td>58.1</td> <td>1,048.4</td> <td>409.5</td> </tr> <tr> <td> </td> <td>nl</td> <td>9.4</td> <td>59.3</td> <td>56.8</td> <td>1,070.2</td> <td>564.0</td> </tr> <tr> <td> </td> <td>pl</td> <td>9.8</td> <td>61.1</td> <td>59.0</td> <td>797.0</td> <td>515.0</td> </tr> <tr> <td> </td> <td>ru</td> <td>9.8</td> <td>60.9</td> <td>58.6</td> <td>1,158.9</td> <td>392.5</td> </tr> </tbody> </table> <h2>More Information</h2> <h3>Paper</h3> <p>The paper in which this dataset is proposed will be published at Interspeech 2024.</p> <p><a href="https://arxiv.org/abs/2407.02937">The preprint is available on arXiv</a>: Meyer, Sarina, Florian Lux, and Ngoc Thang Vu. "Probing the Feasibility of Multilingual Speaker Anonymization." <em>arXiv preprint arXiv:2407.02937</em> (2024). </p> <h3>Code</h3> <p>All code related to this data, including the data and descriptions, as well as preparation scripts to use this data for speaker anonymization, can be found in our Github repository:<a href="https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual"> https://github.com/DigitalPhonetics/speaker-anonymization/tree/multilingual</a></p>
Multilingual Linguistic Landscapes of Mallorca
<p>The dataset contains photographs of the linguistic landscape of Mallorca and was collected during a field trip in May 2023. The photographs were taken in Palma de Mallorca, Port de Pollença, Valledemosa and Alcúdia. A special focus is placed on the interplay of Romanian and Catalan with Spanish in the linguistic landscape. However, semiotic elements, the presence of the languages of tourism and transgressive signage were also part of the explorative fieldwork and are therefore included in this dataset too.</p>
BVS Corpus: A Multilingual Parallel Corpus and Translation Experiments of Biomedical Scientific Texts
<p>The BVS database (Health Virtual Library) is a centralized source of biomedical information for Latin America and Carib, created in 1998 and coordinated by BIREME in agreement with the Pan American Health Organization (OPAS). Abstracts are available in English, Spanish, and Portuguese, with a subset in more than one language, thus being a possible source of parallel corpora. In this article, we present the development of parallel corpora from BVS in three languages: English, Portuguese, and Spanish. Sentences were automatically aligned using the Hunalign algorithm for EN/ES and EN/PT language pairs, and for a subset of trilingual articles also. We demonstrate the capabilities of our corpus by training a Neural Machine Translation (OpenNMT) system for each language pair, which outperformed related works on scientific biomedical articles. Sentence alignment was also manually evaluated, presenting an average 96\% of correctly aligned sentences across all languages. Our parallel corpus is freely available, with complementary information regarding article metadata.</p> <p> </p> <p>Copyright (c) 2019 Secretaría de Estado para el Avance Digital</p>
FONA corpus: Food & Nutrition Abstracts Multilingual corpus
<p>The FONA corpus is a collection of case reports specifically selected to foster the development of Language Technologies, Text Mining and NLP for applications in the domain of food & nutrition.</p> <p> </p> <p>It contains a large collection of documents (titles and abstracts) with metadata information on their MeSH terms. In addition, a subset of the collection contains automatically recognized entities of the following categories:</p> <ul> <li>medical procedures</li> <li>symptoms</li> <li>diseases</li> <li>medications</li> <li>occupational and demographic information</li> <li>species (pathogens)</li> <li>cancer morphology</li> </ul>
Survey on Speech-and-Language Therapists' attitudes and approaches towards multilingualism across four European countries
<p>The file contains all the responses of 300 Speech-and-Language Therapists (SLTs) from Germany, Austria, Italy and Switwerland to a series of questions concerning the provision of speech and language therapy to multilingual children with Developmental Language Disorders. Both the original responses and numerically coded versions are included. Two excel sheets are provided: in the first sheet all responses to the general questionnaire are reported, whereas in the second sheet a subset of the responses obtained from 154 German-speaking SLT respondents are included (more in-depth analyses will be performed on an extended dataset concerning German-speaking SLTs only). These responses have been analysed and presented in the paper titled "Speech and Language Therapy service for multilingual children: Attitudes and approaches across four European countries".</p>
Multilingualism (TRIPLE Video Tutorial Series)
<p>In this tutorial series, researcher Agnieszka Szulińska and Research Infrastructure Coordinator Edward Gray discuss gotriple.eu. This episode explores the opportunities GoTriple offers for multilingual research. </p> <p>Filmed and edited by Fabian Fess.</p>
GLAMI-1M: A Multilingual Image-Text Fashion Dataset
<p>We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text. The dataset, source code and model checkpoints are published at: https://github.com/glami/glami-1m.</p>
GLAMI-1M: A Multilingual Image-Text Fashion Dataset - 800px
<p>We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has high-quality annotations: all 100k images in the test set and 75% of the 1M training set were human-labeled. The paper presents baselines for image-text classification showing that the dataset presents a challenging fine-grained classification problem: The best scoring EmbraceNet model using both visual and textual features achieves 69.7% accuracy. Experiments with a modified Imagen model show the dataset is also suitable for image generation conditioned on text. The dataset, source code and model checkpoints are published at: https://github.com/glami/glami-1m</p>
Migration Reframed - Multilingual stance annotated Twitter news replies on migration in Europe in the context of the Ukrainian crisis
<p><em>The corresponding paper for this dataset "Migration Reframed? A multilingual analysis on the stance shift in Europe during the Ukrainian crisis" has been published in the ACM Web Conference 2023 (WWW'23), and can be accessed here: </em><a href="https://doi.org/10.1145/3543507.3583442">https://doi.org/10.1145/3543507.3583442</a> <em>. Please cite this when using the dataset.</em></p> <p>Twitter dataset of European news and replies to investigate public stance on refugees/migrants around the Ukrainian Crisis.</p> <p>September 2021 to August 2022.</p> <p>Countries:</p> <ul> <li>France</li> <li>Germany</li> <li>Italy</li> <li>Poland</li> <li>Spain</li> </ul> <p>Dataset contains:</p> <ul> <li>Usernames of news outlet accounts on Twitter</li> <li>Tweet IDs of these news accounts during the mentioned period filtered for the migration topic + respective replies from the public</li> <li>Tweet IDs of stance annotated replies</li> <li>8,242 tweet/reply pairs labeled with the stance (positive / negative / neutral) on migrants/refugees (on request)</li> </ul> <table> <caption>Dataset overview by the numbers</caption> <thead> <tr> <th scope="col">Country</th> <th scope="col">News Outlets</th> <th scope="col">News Tweets</th> <th scope="col">Replies</th> <th scope="col">Stance Annotated</th> </tr> </thead> <tbody> <tr> <td>France</td> <td>37</td> <td>2,020</td> <td>32,839</td> <td>500</td> </tr> <tr> <td>Germany</td> <td>72</td> <td>3,752</td> <td>55,317</td> <td>500</td> </tr> <tr> <td>Italy</td> <td>21</td> <td>1,305</td> <td>9,892</td> <td>500</td> </tr> <tr> <td>Poland</td> <td>35</td> <td>3,138</td> <td>27,892</td> <td>6,242</td> </tr> <tr> <td>Spain</td> <td>35</td> <td>1,263</td> <td>20,771</td> <td>500</td> </tr> <tr> <td> </td> <td>200</td> <td>11,478</td> <td>146,711</td> <td>8,242</td> </tr> </tbody> </table> <p>Please note: Stance labels are not included and are only available on request.</p>
The Vuk'uzenzele South African Multilingual Corpus
<pre># The Vuk'uzenzele South African Multilingual Corpus [](https://doi.org/10.5281/zenodo.7598539) Github: <a href="https://github.com/dsfsi/vukuzenzele-nlp">https://github.com/dsfsi/vukuzenzele-nlp</a> ## About dataset The dataset contains editions from the South African government magazine Vuk'uzenzele. Data was scraped from PDFs that have been placed in the [data/raw](data/raw/) folder. The PDFS were obtained from the [Vuk'uzenzele website](https://www.vukuzenzele.gov.za/). The datasets contain government magazine editions in 11 languages, namely: | Language | Code | Language | Code | |------------|-------|------------|-------| | English | (eng) | Sepedi | (sep) | | Afrikaans | (afr) | Setswana | (tsn) | | isiNdebele | (nbl) | Siswati | (ssw) | | isiXhosa | (xho) | Tshivenda | (ven) | | isiZulu | (zul) | Xitstonga | (tso) | | Sesotho | (nso) | ### Number of Aligned Pairs with Cosine Similarity Score >= 0.65 | src_lang | trg_lang | num_aligned_pairs | |----------|----------|-------------------| | ven | zul | 186 | | ssw | xho | 1965 | | sep | xho | 279 | | nbl | zul | 227 | | nso | tsn | 1279 | | nso | tso | 1491 | | tsn | zul | 1346 | | afr | eng | 1369 | | eng | ssw | 1601 | | afr | ssw | 1496 | | nbl | ssw | 264 | | tso | zul | 1758 | | afr | zul | 1384 | | eng | zul | 1888 | | ssw | tsn | 1263 | | sep | tsn | 302 | | nso | xho | 1248 | | sep | tso | 324 | | ssw | tso | 1657 | | tsn | ven | 235 | | eng | nbl | 153 | | nso | sep | 349 | | afr | nbl | 359 | | nbl | ven | 657 | | eng | ven | 243 | | afr | ven | 281 | | tso | ven | 256 | | ven | xho | 215 | | eng | tsn | 1380 | | afr | tsn | 1076 | | nso | ssw | 1132 | | eng | tso | 2016 | | afr | tso | 1139 | | xho | zul | 1895 | | tsn | xho | 1209 | | sep | zul | 223 | | nbl | xho | 204 | | ssw | zul | 2161 | | afr | xho | 1363 | | eng | xho | 1354 | | tso | xho | 1485 | | sep | ssw | 219 | | nbl | tso | 215 | | tsn | tso | 1570 | | nso | zul | 1247 | | nbl | tsn | 140 | | eng | sep | 276 | | afr | sep | 394 | | ssw | ven | 217 | | sep | ven | 1140 | | afr | nso | 962 | | eng | nso | 1721 | | nbl | nso | 151 | | nbl | sep | 843 | | nso | ven | 262 | The dataset is present in several forms on the repo. Generally the dataset is split by edition, eg. `2020-01-ed1` The data directory is broken down as follows ``` ./data ├── external # Data external to this repo ├── interim # I am not really sure - looks like interim in regards to processed. ├── processed # The data from scraping the raw pdfs ├── raw # The raw pdfs of the Vuk'uzenzele magazine ├── sentence_align_output # The output (csv) of the sentence alignment with LASER language encoders └── simple_align_output # The output (csv) of a simple one to one sentence alignment ``` The dataset is split by edition in the [data/processed](data/processed/) folder. Authors ------- - Vukosi Marivate - [@vukosi](https://twitter.com/vukosi) - Andani Madodonga - Daniel Njini - Richard Lastrucci Citation -------- Vukosi Marivate, Andani Madodonga, Daniel Njini, Richard Lastrucci, Isheanesu Dzingirai . **The Vuk'uzenzele South African Multilingual Corpus**, 2023 > @inproceedings{lastrucci-etal-2023-preparing, title = "Preparing the Vuk{'}uzenzele and {ZA}-gov-multilingual {S}outh {A}frican multilingual corpora", author = "Richard Lastrucci and Isheanesu Dzingirai and Jenalea Rajab and Andani Madodonga and Matimba Shingange and Daniel Njini and Vukosi Marivate", booktitle = "Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)", month = may, year = "2023", address = "Dubrovnik, Croatia", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2023.rail-1.3", pages = "18--25" } > @dataset{marivate_vukosi_2023_7598540, author = {Marivate, Vukosi and Njini, Daniel and Madodonga, Andani and Lastrucci, Richard and Dzingirai, Isheanesu}, title = {The Vuk'uzenzele South African Multilingual Corpus}, month = feb, year = 2023, publisher = {Zenodo}, doi = {10.5281/zenodo.7598539}, url = {https://doi.org/10.5281/zenodo.7598539} } Licences ------- * License for Data - [CC 4.0 BY SA](LICENSE.data.md) * Licence for Code - [MIT License](LICENSE.md)</pre>
Six User Personas for the Multilingual DH Community
<p>This document contains six fictional user personas. The personas are based on data collected during a status quo survey carried out in June 2020 as part of the “Linguistic and geocultural diversity in digital knowledge infrastructures” thematic group at the Disrupting Digital Monolingualism Symposium hosted by King’s College London. The personas were originally appended to a paper prepared for an alternative session at the DH Unbound 2022 conference (ACH/CSDH-SCHN).</p>
NLAS-multi: A Multilingual Corpus of Automatically Generated Natural Language Argumentation Schemes
<p>The multilingual corpus of natural language argumentation schemes (NLAS-multi) consists of 3,810 natural language argumentation schemes of which 1,893 are in English and 1,917 in Spanish. It has a total of 253,516 words distributed in 118,493 words in English and 135,023 words in Spanish. In terms of inferences, our corpus has a total of 7,964 (3,949 in English and 4,015 in Spanish). Furthermore, the NLAS-multi corpus contains a total of 23,781 conflict relations between arguments in the same topic.</p>
Data for Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake
<p>Data to go with our submission to Climatic Change titled "Progress on Climate Action: a Multilingual Machine Learning Analysis of the Global Stocktake".</p> <p>Dataset contains the embeddings (.zip with pickles) as well as the associated document items (idem), the most-closely associated keywords and paragraphs per topic in the final model (.xlsx), the reduced 2d embeddings with all selected paragraphs (.csv utf-8 encoded), as well as an overview with the meta-data per source (.csv utf-8 encoded).</p>
SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)
<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab: <a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is <a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>
FakeCovid- A Multilingual Cross domain Fact Check Dataset for COVID-19
<p>FakeCovid is the first multilingual cross-domain dataset of 7623 fact-checked news articles for COVID-19, collected from 04/01/2020 to 01/07/2020. We have collected the fact-checked articles from 92 fact-checking websites after obtaining references from Poynter and Snopes. We have manually annotated the collected articles into 11 categories of the fact-checked news according to their content. We ultimately generated dataset is in 40 languages from 105 countries. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.