Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
135
datasets available to search
ShareScore release 0.7.1
Dataset results
135 results for “multilingual”
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
<p>We present a dataset that collects tweets from news media channels worldwide that pertain to the Russo-Ukrainian war. This dataset spans a period of February 2022-May 2023. The dataset is unique in its global scope, encompassing tweets in various languages and from different parts of the world. Additionally, we extracted information about the stance, sentiment, prominent entities & concepts that occur in tweets to be able to answer questions about the discourse: who says what (prominent entities), who stands (stance) where on what aspect (prominent concepts), how are the aspects portrayed (sentiment). We also downloaded the images attached to the post and classified them to extract image tags for each image. The dataset includes 1,524,826 tweets, out of which 306,295 tweets have images, for 60 languages.<br><br>The source code for the collection and processing of tweets can be found on here: <a href="https://github.com/sherzod-hakimov/ru-ua-news-discourse-twitter"><em>https://github.com/sherzod-hakimov/ru-ua-news-discourse-twitter</em></a></p> <p>Each entry in the dataset is a single JSON line and has the following entries:</p> <pre><code>{ 'tweet_id': 'lang': 'stanza_output': 'stanza_named_entities': 'sentiment': 'stance': 'channel': 'country': 'verified':<br>'image_tags': }</code></pre> <pre> </pre> <p><em><strong>If you need access to the full text of the dataset, please</strong> <strong>contact us via an email: <a href="mailto:sherzodhakimov@gmail.com">sherzodhakimov (at sign) gmail.com</a></strong></em><br><br>If you find the resources useful, please cite us:<br><br>```</p> <p>@inproceedings{hakimov2023unveiling,<br> title={Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict}, <br> author={Sherzod Hakimov and Gullal S. Cheema},<br> booktitle={Proceedings of the 2024 {ACM} International Conference on Multimedia Retrieval, {ICMR} 2024},<br> year={2024}<br>}<br>```</p>
Supplementary materials for article entitled 'Do languages spoken in multilingual communities converge? A case study of reflexivity marking in Mano and Kpelle', published in Linguistics
<p>See the publication for details</p>
Multilingual Stylometry: Data and Code
<p>Data and code accompanying the Multilingual Stylometry Showcase as well as the research paper describing the showcase. </p>
FIGURES 173–181 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 173–181. Genera of New World Scarabaeinae. 173. Uroxys cuprescens Westwood. 174. Uroxys dilaticollis Blanchard. 175. Uroxys sp., ventral view (arrows indicate trochanteral pits). 176. Uroxys sp., lateral view (arrow indicates lateral pronotal sulcus). 177. Vulcanocanthon seminulum (Harold). 178. Xenocanthon sericans (Schimdt). 179. Zonocopris gibbicollis (Harold). 180. Zonocopris gibbicollis, metatibia and tarsus (arrow indicates apical spine of last tarsomere). 181. Zonocopris gibbicollis, ventral view (arrows indicate sternal foveae).
FIGURES 113–120 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 113–120. Genera of New World Scarabaeinae. 113. Megathopa villosa Eschscholtz. 114. Megathopa villosa, tip of metatibia and tarsus (arrow indicates apical spine). 115. Megathoposoma candezei Balthasar. 116. Melanocanthon nigricornis (Say). 117. Melanocanthon nigricornis, hind leg (arrows indicate tibial spurs). 118. Nunoidium argentinum (Arrow). 119. Onitis alexis Klug. 120. Onoreidium howdeni Ferreira & Galileo
FIGURES 11–20 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 11–20. Genera of New World Scarabaeinae. 11. Ateuchus viduus (Blanchard). 12. A. viduus, ventral view. 13. A. texanus (Robinson) 14. A. texanus, hind leg. 15. Attavicinus monstrosus (Bates). 16. Bdelyropsis bowditchi (Paulian). 17. B. bowditchi, posterior view. 18. Bdelyrus seminudus (Bates). 19. B. seminudus, ventral view posterior portion body. 20. B. seminudus, hind leg (a - dorsal view; b - lateral view).
FIGURES 21–29 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 21–29. Genera of New World Scarabaeinae. 21. Besourenga horacioi (Martínez) 22. Bolbites onitoides Harold. 23. Bradypodidium adisi (Ratcliffe) 24. Canthidium (C.) barbacenicum Preudhomme de Borre. 25. Canthidium (C.) barbacenicum, dorsal view of pronotum (arrows indicate basal row punctures). 26. Canthidium (Eucanthidium) cupreum (Blanchard) 27. Canthidium (E.) cupreum, ventral view (arrows indicate mesosternum). 28. Canthidium (C.) haroldi Preudhomme de Borre, detail left elytral margin. 29. Canthochilum oakleyi Chapin
FIGURES 128–136 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 128–136. Genera of New World Scarabaeinae. 128. Onthophagus (O.) sp. 129. Onthophagus (O.) xanthomerus Bates. 130. Onthophagus (Palaeonthophagus) nuchicornis (Linnaeus). 131. Onthophagus (O.) hoepfneri Harold. 132. Onthophagus sp., hind leg. 133. Onthophagus (O.) chevrolati Harold. 134. Oruscatus davus (Erichson). 135. Oruscatus davus, antenna. 136. Oruscatus davus, tip of abdomen (arrow indicates basal pygidial carina).
Publishing, linking and translating news in multilingual communities: a mirror of cultural differences? (auxiliary material)
<p>Dataset with snapshots of published news from SwissInfo in English, French, German and Italian homepages.</p> <p>Used for the analyses presented in the cited paper.</p>
CT-FAN-22 corpus: A Multilingual dataset for Fake News Detection
<p><strong>Data Access: </strong>The data in the research collection provided may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use it only for research purposes. Due to these restrictions, the collection is not open data. Please download the Agreement at <a href="https://drive.google.com/file/d/1QU-rw4D26r3F04FB63hTvToxOvDKaKdv/view?usp=sharing">Data Sharing Agreement</a> and send the signed form to <a href="mailto:fakenewstask@gmail.com">fakenewstask@gmail.com</a> .</p> <p><strong>Citation</strong></p> <p>Please cite our work as</p> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Subtask 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong> Sub-task A would detect fake news designed as a four-class classification problem. The training data will be released in batches and roughly about 900 articles with the respective label. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <p><strong>Task 3</strong></p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <p><strong>Task 3</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p>Sample file</p> <pre><code>public_id, predicted_domain 1, health 2, crime</code></pre> <p><strong>Additional data for Training</strong></p> <p>To train your model, the participant can use additional data with a similar format; some datasets are available over the web. We don't provide the background truth for those datasets. For testing, we will not use any articles from other datasets. Some of the possible sources:</p> <ul> <li><a href="https://www.kaggle.com/liberoliber/onion-notonion-datasets">Fakenews Classification Datasets</a></li> <li><a href="https://www.kaggle.com/c/fakenewskdd2020/overview">Fake News Detection Challenge KDD 2020</a></li> <li><a href="https://www.kaggle.com/mdepak/fakenewsnet?select=PolitiFact_real_news_content.csv">FakeNewsNet</a></li> </ul> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2021, and the content of fake news is mixed up with several topics like election, COVID-19 etc.</li> </ol> <p><strong>Evaluation Metrics</strong></p> <p>This task is evaluated as a classification task. We will use the F1-macro measure for the ranking of teams. There is a limit of 5 runs (total and not per day), and only one person from a team is allowed to submit runs.</p> <p><strong>Submission Link: </strong><a href="https://codalab.org/">Coming soon</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi, G. K., Struß, J. M., & Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection. <em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeño, A., Míguez, R., Shaar, S., ... & Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In <em>European Conference on Information Retrieval</em> (pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeño, A., Míguez, R., Shaar, S., ... & Kartal, Y. S. (2021, September). Overview of the CLEF–2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In <em>International Conference of the Cross-Language Evaluation Forum for European Languages</em> (pp. 264-291). Springer, Cham.</li> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, “FakeCovid – a multilingualcross-domain fact check news dataset for covid-19,” inWorkshop Proceedings of the 14th International AAAIConference on Web and Social Media, 2020. <a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., & Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter. <em>Online Social Networks and Media</em>, <em>22</em>, 100104. doi: <a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> </ul>
FIGURE 184 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURE 184. Morfología externa básica de los Scarabaeinae.
DAMP-MVP: Digital Archive of Mobile Performances - Smule Multilingual Vocal Performance 300x30x2
<p>The Smule 300x30x2 dataset contains recordings of sung karaoke tracks, lyrics text files, and some metadata describing the songs being performed by each singer. This dataset was collected from performances on Smule by selecting the most popular singers, female and male, of the 300 most popular arrangements in 30 countries.</p> <p>The most popular arrangements were determined by counting song starts or joins (duet/group) of recordings for each arrangement, within the country of interest. The term "arrangement" is used, because there might be multiple arrangements of the same song.</p> <p>The most popular performances and singers of those arrangements were determined by counting Listens (and/or Loves, if no complete Listens) for all performances of each arrangement.</p> <p>Users of this dataset must read and accept Smule's Research Data License Agreement (LICENSE.txt).</p>
EmoFilm - A multilingual emotional speech corpus
<p><strong>EmoFilm</strong> is a multilingual emotional speech corpus comprising 1115 audio instances produced in English, Italian, and Spanish languages. The audio clips (with a mean length of 3.5 sec. and std 1.2 sec.) were extracted in wave format (uncompressed, mono, 48 kHz sample rate and 16-bit) from 43 films (original in English and their over-dubbed Italian and Spanish versions). Genres including comedy, drama, horror, and thriller were considered; anger, contempt, happiness, fear, and sadness emotional states were taken into account. EmoFilm has been presented at Interspeech 2018:</p> <p>Emilia Parada-Cabaleiro, Giovanni Costantini, Anton Batliner, Alice Baird, and Björn Schuller (2018), <em>Categorical vs Dimensional Perception of Italian Emotional Speech</em>, in Proc. of Interspeech, Hyderabad, India, pp. 3638-3642 .</p> <p>We would like to thank Linda Ratz for her contribution in the generation of the transcriptions.</p> <p> </p> <p><strong>How to access EmoFilm</strong></p> <p>To get access to the dataset, please send the signed End User License Agreement (EULA) when making the request. The EULA <strong>must be signed by somebody from a university holding a permanent position</strong>, typically a full professor. Note that requests without an EULA appropriately filled out, as well as those performed from a non-institutional e-mail address, will be automatically rejected. Please download the EULA from the following link:</p> <p>https://drive.google.com/file/d/1pFHfsqk7snF_EVqq0WAC0Dz8FcTD3s9_/view?usp=share_link</p>
Multilingual fine-grained sentiment analysis corpus
<p>A sentiment annotated corpus based on Fallout New Vegas. The corpus has the following sentiments: <em>neutral, anger, disgust, fear, happy, pained, sad, surprised</em> in the following languages: <em>English, German, Italian, Spanish and French</em>.</p> <p>Please cite the following paper: Mika Hämäläinen, Khalid Alnajjar, and Thierry Poibeau. 2022. Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog. In <em>FDG’22: Proceedings of the 17th International Conference on the Foundations of Digital Games (FDG ’22)</em></p> <p> </p>
Navigating Power in the Multilingual Classroom
<p>Interview Data</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.