Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

14

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

14 results for “Cross-Lingual”

Learn how ShareScore rates datasets ↗
zenodo48/100

Cross-Lingual Dataset of Crisis-Related Social Media

<p>The cross-lingual natural disaster dataset includes public tweets collected using Twitter&rsquo;s public API, filtering by location-related keywords and date, without using any additional filtering (e.g., we did not restrict the query to specific languages). We considered two&nbsp;disaster events and two long-term natural disasters across Europe (floods and wildfires)&nbsp;that received substantial news coverage internationally.</p> <p>Three of the top languages were common to all the studied events: English (ISO 639-1 code: en), Spanish (es), and French (fr). Additionally, we found hundreds of messages for each event in other five languages, including Arabic (ar), German (de), Japanese (ja), Indonesian (id), Italian (it) and Portuguese (pt).&nbsp;</p> <p>After collecting the data, we labelled tweets that contained potentially informative factual information. We name this group of tweets &ldquo;informative messages.&rdquo; Next, we used crowdsourcing to further categorize the messages into various informational categories. We asked three different workers to label each&nbsp;informative messages across languages. The target categories were based on an ontology from TREC-IS 2018, where we grouped some low level ontology categories into higher-level ones.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Cross-Lingual Dataset of Crisis-Related Social Media

<p>The cross-lingual natural disaster dataset includes public tweets collected using Twitter&rsquo;s public API, filtering by location-related keywords and date, without using any additional filtering (e.g., we did not restrict the query to specific languages). We considered five disaster events between January 2020 and February 2021 that received substantial news coverage internationally.</p> <p>All messages include a &ldquo;language&rdquo; field computed by Twitter us ing a language detection model developed specifically for tweets. We counted the number of messages per language in each event. Three of the top languages were common to all the studied events: English (ISO 639-1 code: en), Spanish (es), and French (fr). Additionally, we found several hundred messages for each event in other languages, including Catalan (ca), Tagalog (tl), Croatian (hr), German (de), Japanese (ja), Indonesian (id), and Portuguese (pt).&nbsp;</p> <p>After collecting the data, we labelled tweets or their translation to English that contained potentially informative factual information. We name this group of tweets &ldquo;informative messages.&rdquo; Next, we used crowdsourcing to further categorize the messages into various informational categories. We asked three different workers to label each of the approximately 5,700 informative messages across languages. The target categories were based on an ontology from TREC-IS 2018, where we grouped some low level ontology categories into higher-level ones.</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

Sentiment Analysis and Cross-lingual Word Embeddings for Endangered Languages

<p>A sentiment analyzer and cross-lingual word embeddings for endangered languages (e.g., Erzya, Moksha, Skolt Sami, Komi-Zyrian).</p>

opencc-by-4.0Mar 2021View details →
zenodo36/100

Dataset from the ISMIR2020 article "Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Item Annotation"

<p>We release the data required to reproduce the cross-lingual music genre translation experiments from the article&nbsp;<strong><em>Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Item Annotation</em></strong>&nbsp;presented at the&nbsp;<a href="https://ismir.github.io/ISMIR2020/">ISMIR 2020</a>&nbsp;conference.</p> <p>More information about this&nbsp;data&nbsp;and how it should be used in the experiments can be found&nbsp;in the GitHub repository <a href="http://github.com/deezer/MultilingualMusicGenreEmbedding">deezer/MultilingualMusicGenreEmbedding</a>.</p> <p>Please cite our paper if you use the code or data in your work.</p>

opencc-by-4.0Oct 2020View details →
zenodo36/100

SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages

<div> <div>SpokeN-100 is a novel, entirely artificially generated benchmarking dataset tailored for speech recognition, representing a core challenge in the field of tiny deep learning. SpokeN-100 consists of spoken numbers from 0 to 99 spoken by 32 different speakers in four different languages, namely English, Mandarin, German and French, resulting in 12,800 audio samples.</div> </div>

opencc-by-4.0Mar 2024View details →
zenodo36/100

DEepfake CROss-lingual evaluation dataset (DECRO)

<p>Deepfake cross-lingual evaluation dataset (DECRO) is constructed to evaluate the influence of language differences on deepfake detection.&nbsp;</p> <p><strong>If you use DECRO dataset for deepfake detection, please&nbsp;cite the paper &quot;<em>Transferring Audio Deepfake Detection Capability across Languages</em>&quot; published in <em>www&#39;23</em>.</strong></p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

xLiD-Lexica: Cross-lingual Linked Data Lexica

<p>We provide <strong>cross-lingual linked data lexica called xLiD-Lexica</strong>. The data set contains the <strong>cross-lingual groundings of linked data resources from the Linked Open Data cloud as RDF data</strong>, which can be easily integrated into the LOD data sources.&nbsp; In addition, we created a SPARQL endpoint over ourxLiD-Lexica to allow users to easily access them using SPARQL query language. Multilingual and cross-lingual information access can be facilitated by the availability of such lexica, e.g., allowing for an easy mapping of natural language expressions in different languages to linked data resources from LOD. Many tasks in natural language processing, such as natural language generation, cross-lingual entity linking, text annotation and question answering, can benefit from our xLiD-Lexica.</p> <p>More information can be found in the <a href="http://dbis.informatik.uni-freiburg.de/content/team/faerber/papers/xLiD_LREC2014.pdf"><strong>LREC&#39;14 paper <em>xLiD-Lexica: Cross-lingual Linked Data Lexica</em></strong></a> and on our website <strong><a href="https://km.aifb.kit.edu/sites/xlid-lexica/">https://km.aifb.kit.edu/sites/xlid-lexica/</a></strong>.</p> <p>Please cite this data set as follows (see also <a href="https://dblp.org/rec/bibtex/conf/lrec/ZhangFR14">DBLP</a>):</p> <pre><code>Lei Zhang, Michael Färber, Achim Rettinger. "xLiD-Lexica: Cross-lingual Linked Data Lexica". In: Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC 2014). Reykjavik, Iceland, 2014, pp. 2101–2105.</code></pre> <p>&nbsp;</p> <p><strong>Example queries:</strong></p> <p>1. Retrieve all entities with surface form which contain &quot;iPhone&quot;:</p> <pre><code class="language-sql">Select ?resource, ?label, ?probability from &lt;http://www.xlid-lexica.org&gt; where { ?resource &lt;http://www.xlid-lexica.org/block&gt; ?b1 . ?b1 &lt;http://www.xlid-lexica.org/res#sf&gt; ?sf . ?b1 &lt;http://www.xlid-lexica.org/res#priorProbability&gt; ?probability . ?sf &lt;http://www.xlid-lexica.org/block&gt; ?b2. ?b2 &lt;http://www.xlid-lexica.org/sf#label&gt; ?label . ?label bif:contains "iPhone" . } order by DESC(?probability) limit 100</code></pre> <p>2. Retrieve the top 100 resources for a given surface form (&quot;iphone&quot;):</p> <pre><code class="language-sql">Select ?resource, ?probability from &lt;http://www.xlid-lexica.org&gt; where { ?resource &lt;http://www.xlid-lexica.org/block&gt; ?b1 . ?b1 &lt;http://www.xlid-lexica.org/res#sf&gt; ?sf . ?b1 &lt;http://www.xlid-lexica.org/res#priorProbability&gt; ?probability . ?sf &lt;http://www.xlid-lexica.org/block&gt; ?b2. ?b2 &lt;http://www.xlid-lexica.org/sf#label&gt; "iphone"@en . } order by DESC(?probability) limit 100</code></pre> <p>3. Retrieve the top 100 resources for a given surface form (&quot;iphone&quot;, case-insensitive):</p> <pre><code class="language-sql">Select ?resource, ?probability from &lt;http://www.xlid-lexica.org&gt; where { ?resource &lt;http://www.xlid-lexica.org/block&gt; ?b1 . ?b1 &lt;http://www.xlid-lexica.org/res#sf&gt; ?sf . ?b1 &lt;http://www.xlid-lexica.org/res#priorProbability&gt; ?probability . ?sf &lt;http://www.xlid-lexica.org/block&gt; ?b2. ?b2 &lt;http://www.xlid-lexica.org/sf#label&gt; ?surfaceform . filter(regex(?surfaceform, "^iphone$", "i")) } limit 100</code></pre> <p>4. Retrieve the top 100 surface forms per entity:</p> <pre><code class="language-sql">Select ?label ?probability from &lt;http://www.xlid-lexica.org&gt; where { &lt;http://dbpedia.org/resource/IPhone_5&gt; &lt;http://www.xlid-lexica.org/block&gt; ?b1. ?b1 &lt;http://www.xlid-lexica.org/res#sf&gt; ?sf. ?b1 &lt;http://www.xlid-lexica.org/res#priorProbability&gt; ?probability. ?sf &lt;http://www.xlid-lexica.org/block&gt; ?b2. ?b2 &lt;http://www.xlid-lexica.org/sf#label&gt; ?label. ?b2 &lt;http://www.xlid-lexica.org/block#lang&gt; "en". } order by DESC(?probability) limit 100</code></pre> <p>&nbsp;</p>

opencc-by-4.0May 2014View details →
zenodo32/100

MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer

<p>The dataset is published with:<br> <br> <em>MultiEURLEX - A multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer. Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. Proceedings of the&nbsp;2021&nbsp;Conference on&nbsp;Empirical Methods in Natural Language Processing. 2021. Punta Cana, Dominican Republic.</em><br> <br> <strong>Documents: </strong>MultiEURLEX&nbsp;comprises 65k EU&nbsp;in 23 official EU languages. Each EU&nbsp;law has been annotated with EUROVOC&nbsp;concepts (labels) by the Publication Office of EU. Each EUROVOC&nbsp;label ID is associated with a Label descriptor, e.g., [60, `agri-foodstuffs&#39;], &nbsp;[6006, `plant product&#39;], [1115, `fruit&#39;]. The descriptors are also available in 23 languages. Chalkidis et al. (2019) published a&nbsp;monolingual&nbsp;(English) version of this dataset, called EURLEX57K, comprising 57k EU&nbsp;laws with the originally assigned gold labels.</p> <p><strong>Languages: </strong>MultiEURLEX&nbsp;covers 23 languages from 7 families. EU&nbsp;laws are published in all official EU&nbsp;languages, except for Irish for resource-related reasons&nbsp;(Read more:&nbsp;https://europa.eu/european-union/about-eu/eu-languages_en).&nbsp;This wide coverage makes the dataset a valuable testbed for cross-lingual transfer. All languages use the Latin script, except for Bulgarian (Cyrillic script) and Greek.</p> <p><strong>Multi-granular Labeling: </strong>EUROVOC<strong>&nbsp;</strong>has eight levels of concepts. Each document is assigned one or more concepts (labels). If a document is assigned a concept, the ancestors and descendants of that concept are typically not assigned to the same document. The documents were originally annotated with concepts from levels 3 to 8. &nbsp;We created three alternative sets of labels per document, by replacing each assigned concept by its ancestor from levels 1, 2, or 3, respectively. Thus, we provide four sets of gold labels per document, one for each of the first three levels of the hierarchy, plus the original sparse label assignment.</p> <p><strong>Supported Tasks:&nbsp;</strong>Similarly to EURLEX&nbsp;(Chalkidis et al., 2019), MultiEURLEX&nbsp;can be used for legal topic classification, a multi-label classification task where legal documents need to be assigned concepts (in our case, from EUROVOC) reflecting their topics. Unlike EURLEX57K, however, MultiEURLEX&nbsp;supports labels from three different granularities (EUROVOC&nbsp;levels). More importantly, apart from monolingual (one-to-one) experiments, it can be used to study cross-lingual transfer scenarios, including one-to-many&nbsp;(systems trained in one language and used in other languages with no training data), and many-to-one&nbsp;or many-to-many&nbsp;(systems jointly trained in multiple languages and used in one or more other languages).</p> <p><strong>Data Split and Concept Drift:&nbsp;</strong>MultiEURLEX&nbsp;is chronologically&nbsp;split in training (55k, 1958-2010), development (5k, 2010-2012), test (5k, 2012-2016) subsets, using the English documents. The test subset contains the same 5k documents in all 23 languages. The development subset also contains the same 5k documents in 23 languages, except Croatian. Croatia is the most recent EU&nbsp;member (2013); older laws are gradually translated.&nbsp;For the official languages of the seven oldest member countries, the same 55k training documents are available; for the other languages, only a subset of the 55k training documents is available.&nbsp;Compared to EURLEX57K&nbsp;(Chalkidis et al., 2019), MultiEURLEX&nbsp;is not only larger (8k more documents) and multilingual; it is also more challenging, as the chronological split leads to temporal real-world concept drift&nbsp;across the training, development, test subsets, i.e., differences in label distribution and phrasing, representing a realistic temporal generalization&nbsp;problem (Huang and Paul, 2019; Lazaridou et al., 2021). Recently, S&oslash;gaard et al. (2021) showed this setup is more realistic, as it does not overestimate real performance, contrary to random splits (Gorman and Bedrick, 2019).</p>

opencc-by-4.0Aug 2021View details →
zenodo28/100

Dataset construction method of cross-lingual summarization based on filtering and text augmentation

<p>The NCLS dataset is provided by its authors (Zhu et al.): https://drive.google.com/file/d/1GZpKkHnTH_1Wxiti0BrrxPm18y9rTQRL/view. We work on the train set, validation set, and manually corrected test set.</p>

opencc-by-4.0Mar 2023View details →
dryad24/100

Data from: A cross-lingual similarity measure for detecting biomedical term translations

Bilingual dictionaries for technical terms such as biomedical terms are an important resource for machine translation systems as well as for humans who would like to understand a concept described in a foreign language. Often a biomedical term is first proposed in English and later it is manually translated to other languages. Despite the fact that there are large monolingual lexicons of biomedical terms, only a fraction of those term lexicons are translated to other languages. Manually compiling large-scale bilingual dictionaries for technical domains is a challenging task because it is difficult to find a sufficiently large number of bilingual experts. We propose a cross-lingual similarity measure for detecting most similar translation candidates for a biomedical term specified in one language (source) from another language (target). Specifically, a biomedical term in a language is represented using two types of features: (a) intrinsic features that consist of character n-grams extracted from the term under consideration, and (b) extrinsic features that consist of unigrams and bigrams extracted from the contextual windows surrounding the term under consideration. We propose a cross-lingual similarity measure using each of those feature types. First, to reduce the dimensionality of the feature space in each language, we propose prototype vector projection (PVP)—a non-negative lower-dimensional vector projection method. Second, we propose a method to learn a mapping between the feature spaces in the source and target language using partial least squares regression (PLSR). The proposed method requires only a small number of training instances to learn a cross-lingual similarity measure. The proposed PVP method outperforms popular dimensionality reduction methods such as the singular value decomposition (SVD) and non-negative matrix factorization (NMF) in a nearest neighbor prediction task. Moreover, our experimental results covering several language pairs such as English–French, English–Spanish, English–Greek, and English–Japanese show that the proposed method outperforms several other feature projection methods in biomedical term translation prediction tasks.

opencc-zeroDec 2014View details →
dryad24/100

Data from: A cross-lingual similarity measure for detecting biomedical term translations

Open the record for dataset details and reuse information.

publicMay 2016View details →
zenodo16/100

Webis Cross-Lingual Sentiment Dataset 2010 (Webis-CLS-10)

<p>The Cross-Lingual Sentiment (CLS) dataset comprises about 800.000 Amazon product reviews in the four languages English, German, French, and Japanese.</p> <p>For more information on the construction of the dataset see (Prettenhofer and Stein, 2010) or the enclosed readme files. If you have a question after reading the paper and the readme files, please contact <a href="https://weimar.webis.de/people#prettenhofer">Peter Prettenhofer</a>.</p> <p>We provide the dataset in two formats: 1) a processed format which corresponds to the preprocessing (tokenization, etc.) in (Prettenhofer and Stein, 2010); 2) an unprocessed format which contains the full text of the reviews (e.g., for machine translation or feature engineering).</p> <p>The dataset was first used by (Prettenhofer and Stein, 2010). It consists of Amazon product reviews for three product categories---books, dvds and music---written in four different languages: English, German, French, and Japanese. The German, French, and Japanese reviews were crawled from Amazon in November, 2009. The English reviews were sampled from the <a href="http://www.cs.jhu.edu/~mdredze/datasets/sentiment/">Multi-Domain Sentiment Dataset</a> (Blitzer et. al., 2007). For each language-category pair there exist three sets of training documents, test documents, and unlabeled documents. The training and test sets comprise 2.000 documents each, whereas the number of unlabeled documents varies from 9.000 - 170.000.</p>

restrictedJul 2010View details →
zenodo12/100

PaCL: A Multi-Domain and Bilingual Benchmark for Cross-Lingual Patent Ranking

Open the record for dataset details and reuse information.

restrictedcc-by-4.0Jun 2024View details →
zenodo12/100

xMP: A Cross-lingual Multi-label Media Profiling Dataset

<p>We propose xMP, a zero-shot cross-linguistic assessment benchmark for factual and political bias in media coverage. We gathered data to create a multilingual test set of media profiling tasks in 34 languages and 18 language families. This allowed us to evaluate model performance and expand work based on English resources to other languages.</p>

restrictedAug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record