Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
135
datasets available to search
ShareScore release 0.7.1
Dataset results
135 results for “multilingual”
The South African Gov-ZA multilingual corpus
<pre>The South African Gov-ZA multilingual corpus ============================== Github: https://github.com/dsfsi/gov-za-multilingual Zenodo: About Dataset --------------------- The data set contains cabinet statements from the South African government. Data was scraped from the governments website: https://www.gov.za/cabinet-statements The datasets contain government cabinet statements in 11 languages, namely: | Language | Code | Language | Code | |------------|------|------------|------| | English | (eng) | Sepedi | (nso) | | Afrikaans | (afr) | Setswana | (tsn) | | isiNdebele | (nbl) | Siswati | (ssw) | | isiXhosa | (xho) | Tshivenda | (ven) | | isiZulu | (zul) | Xitstonga | (tso) | | Sesotho | (sot) | The dataset contains the full data in a JSON file (/data/govza-cabinet-statements.json), as well as CSV’s split by each language, eg: “govza-cabinet-statements-en.csv” for english. The dataset does not contain special characters like unicode or ascii. Please see the [data-statement.md](/data_statement.md) for full dataset information. *(TODO)* Number of Aligned Pairs with Cosine Similarity Score >= 0.65 ------------------------------------------------------------ | src_lang | trg_lang | num_aligned_pairs | |----------|----------|-------------------| | afr | eng | 14549 | | afr | nbl | 6621 | | afr | nso | 15388 | | afr | sot | 8834 | | afr | ssw | 15610 | | afr | tsn | 12605 | | afr | tso | 14936 | | afr | ven | 5776 | | afr | xho | 16065 | | afr | zul | 14998 | | nbl | eng | 3616 | | nbl | nso | 6342 | | nbl | sot | 16163 | | nbl | ssw | 4655 | | nbl | tsn | 3369 | | nbl | tso | 4465 | | nbl | ven | 18984 | | nbl | xho | 5213 | | nbl | zul | 3868 | | nso | eng | 15257 | | nso | ssw | 18697 | | nso | tsn | 16179 | | nso | tso | 17617 | | nso | ven | 6367 | | sot | eng | 5212 | | sot | nso | 8077 | | sot | ssw | 5811 | | sot | tsn | 5450 | | sot | tso | 6586 | | sot | ven | 14098 | | ssw | eng | 15721 | | ssw | tso | 17880 | | ssw | ven | 4588 | | tsn | eng | 14544 | | tsn | ssw | 16386 | | tsn | tso | 16681 | | tsn | ven | 3267 | | tso | eng | 16068 | | ven | eng | 3670 | | ven | tso | 4578 | | xho | eng | 16537 | | xho | nso | 18110 | | xho | sot | 7489 | | xho | ssw | 18387 | | xho | tsn | 16571 | | xho | tso | 17954 | | xho | ven | 4559 | | xho | zul | 18145 | | zul | eng | 16149 | | zul | nso | 17630 | | zul | sot | 5975 | | zul | ssw | 18563 | | zul | tsn | 16482 | | zul | tso | 17789 | | zul | ven | 3606 | Authors ------- - Vukosi Marivate - [@vukosi](https://twitter.com/vukosi) - Matimba Shingange - Richard Lastrucci - Isheanesu Joseph Dzingirai - Jenalea Rajab Publications -------</pre> <p>> @inproceedings{lastrucci-etal-2023-preparing,<br> title = "Preparing the Vuk{'}uzenzele and {ZA}-gov-multilingual {S}outh {A}frican multilingual corpora",<br> author = "Richard Lastrucci and Isheanesu Dzingirai and Jenalea Rajab and Andani Madodonga and Matimba Shingange and Daniel Njini and Vukosi Marivate",<br> booktitle = "Proceedings of the Fourth workshop on Resources for African Indigenous Languages (RAIL 2023)",<br> month = may,<br> year = "2023",<br> address = "Dubrovnik, Croatia",<br> publisher = "Association for Computational Linguistics",<br> url = "https://aclanthology.org/2023.rail-1.3",<br> pages = "18--25"<br> }</p>
Multilingual publishing in the SSH in Poland and attitudes towards English
<p>This dataset contains responses to a self-constructed questionnaire that was designed to provide the following information: What languages are used for research dissemination in various SSH disciplines in Poland? What are the main languages of the research cited by authors in these disciplines? When languages other than Polish are used for research dissemination, are the results published in national or international venues? What are the prevailing reasons for language choices? What is the position of English in SSH disciplines in Poland? What is the attitude of Polish SSH scholars towards the dominance of English as the international language of science?</p> <p>The questionnaire was written in Polish and consisted of 52 items arranged in three thematic lines: multilingual publication practices, the role of English in research dissemination, and attitudes to English as the global language of science.</p> <p>The data were collected in an online survey based on Google Forms. The link to the questionnaire was distributed via email among scholars affiliated with the social sciences and humanities units of 20 Polish universities which took part in the Excellence Initiative – Research University competition, a funding programme launched in 2019 by the Ministry of Science and Higher Education (<a href="https://www.gov.pl/web/science/the-excellence-initiative---research-university-programme">https://www.gov.pl/web/science/the-excellence-initiative---research-university-programme</a>).</p> <p>Time of data collection: 12 October 2020–16 November 2020.</p> <p>Volume of data and the response rate: 12,100 emails sent; 1,575 completed forms received (response rate about 13%); 50 forms removed (contradictory or random responses; this dataset is limited to speakers of Polish as the first language).</p> <p>The classification of fields and disciplines follows Polish regulations in force at the time of the study (Regulation of the Polish Minister of Science and Higher Education of 20 September 2018 on Classification of fields and disciplines of science and disciplines of the arts; Journal of Laws 2018, item 1818). Compared to the OECD classification, the main points of difference involve the status of history and archaeology, linguistics and literary studies, and economics and management as separate disciplines.</p> <p>The data collection was not funded by any external source.</p>
Public database of multilingual map reading test
<p>The Excel file contains the filtered data records of the map-reading study of the Research Group on Experimental Cartography at the Eötvös Loránd University (ktk.elte.hu). The data collection started in the autumn of 2015 and lasted until April 2022. The file contains three sheets: demographic_questions; correct_answers; map_reading_database. The first two sheets contain the questions asked, the answer codes, and the correct answers. The third one has 511 records, which is the result of a filtering of the original 805 fills. The filtering excluded the unfinished tests, and the ones with fill time below 2.5 minutes and above 15 minutes.</p>
Multilingual bottle-neck feature learning from untranscribed speech for track 1 in zerospeech2017 (system 2 -- with VTLN)
<p>We investigate the extraction of bottle-neck features (BNFs) for multiple languages without access to manual transcription. Multilingual BNFs are derived from a multi-task learning deep neural network which is trained with unsupervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models separately trained on untranscribed speech of multiple languages.</p> <blockquote> <p>In this version, the input MFCC for DPGMM is processed with VTLN.</p> </blockquote> <p> </p>
Challenges in providing multilingual information access
Open the record for dataset details and reuse information.
Multilingual paired code and comment changes
<p>Dataset used for the master's thesis "LLMs for Code Comment Consistency." Covers the languages Go, Java, JavaScript, TypeScript, and Python. All data is mined from permissively-licensed GitHub public projects.</p><p>This dataset consists of pairs of function/method code blocks and their documentation comments, before and after commits.<br>Examples are labeled 0 if the comment was not changed before and after, and 1 if the comment was changed. For the purpose of comment consistency, that means a 1-labeled example has an <i>old comment</i> that is inconsistent with the <i>new code</i>.<br>If you're training a <strong>code summarization</strong> or <strong>comment generation</strong> task, then of course ignore the classification label.</p><ul><li>All-22k contains the training, validation, and test set used in the models trained in the paper. The examples are balanced by language and between the positive and negative classes. Any code repository is only present in one of these sets.</li></ul>
Data of Paper "Turning a Multilingual Historical Archive into an Information System through Post-OCR Correction and Content-Based Indexation"
<p>We evaluated our approach on a collection of 946 historical documents belonging to the Biblioteca Nacional de Catalunya (BNC), spanning from 1914 to 1951. Each document is the issue of a magazine, comprising different articles by different authors. This implies that, despite the thematic nature of magazines and specific issues, there is a certain degree of heterogeneity in each document. Magazines were selected based on their relevance w.r.t. art in general and, more specifically, early 20th century avant-garde movements (e.g., Dadaism, Cubism, etc.). For each document, we have the scanning of the original artifact and the plain raw text extracted through ABBYY FineReader OCR tool. To the best of our knowledge, this is the first Catalan-dominated OCR corpus ever released. </p>
Interview Corpora for the Study of Multilingual Repertoires in South-South Migration Dynamics I: Haitians in Chapecó (SC, Brazil)
<p>Corpus of 19 multilingual interviews with Haitian migrants in Chapecó (Santa Catarina, Brazil), conducted in various languages: in order from the most to the least documented in the interviews, Portuguese, French, Spanish, and Haitian Creole. This corpus is part of broader research on the evolution of multilingual repertoires in South-South migration dynamics. The interviews were conducted in March 2023 by the author in collaboration with Leonie Ette (University of Augsburg) and the two coordinators of the research group <em>Atlas das Línguas em Contato na Fronteira</em>, Professors Cristiane Horst and Marcelo Krug (UFFS, Campus Chapecó).</p>
Resources for the paper "Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training"
Open the record for dataset details and reuse information.
Multilingual CoNaLa Datset, train data
<p>Training datasets used in the Multilingual CoNaLa benchmark experimentation. </p>
Multilingual Transformations Data
<p>Data for "Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models" (Findings of ACL 2022).</p>
Implementation and evaluation of a multilingual search pilot in the Europeana digital library (dataset)
<p>The dataset contains the data required to reproduce the experiments done in the paper "Implementation and evaluation of a multilingual search pilot in the Europeana digital library", published in the 26th International Conference on Theory and Practice of Digital Libraries (<a href="http://tpdl2022.dei.unipd.it/">TPDL'22</a>). In that work we implemented a pilot applying query translation to English from the Spanish version of the website in order to surface results that have English metadata associated with them. The dataset is also available at <a href="https://rnd-2.eanadev.org/share/crosslingual_SpanishPilot/">https://rnd-2.eanadev.org/share/crosslingual_SpanishPilot/</a>, and it is organized in three main folders:</p> <ul> <li><strong>sample</strong>: stratified sample of 300 queries queries issued from the Europeana Spanish portal from 1st<br> December 2020 to 28th February 2021.</li> <li><strong>evaluation.translations: </strong>manual annotation of the quality of the identification of the language of the queries using Google Cloud Translation API, and the quality of the translation obtained using Google plus the CEF translation service (eTranslation).</li> <li><strong>evaluation.search_retrieval: </strong>manual annotation of the relevancy of the (binary) relevance of the documents that are retrieved by one system but not by the other (current monolingual version vs pilot) in their top ten.</li> </ul>
TUNDRA - A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision,
<p>The corpus is described in:</p> <p><strong>A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King,</strong> <em>TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision</em>, In Proc. Interspeech, Lyon, France, August 2013</p> <p> </p> <p> </p> <pre>############################################################### ## ## ## THE SIMPLE4ALL TUNDRA CORPUS ## ## version 1.0 ## ## ## ############################################################### Simple4All Tundra (version 1.0) is the first release of a standardised multilingual corpus designed for text-to-speech research with imperfect or found data. The corpus consists of approximately 60 hours of speech data from audiobooks in 14 languages, as well as utterance-level alignments obtained with a lightly-supervised process. Most audiobooks are from the public domain and allow redistribution. However, some have restricted use, and in those cases the segmented and aligned data cannot be downloaded from our website. --------------------------------------------------------------- LICENCE --------------------------------------------------------------- This work is licensed under a Creative Commons Attribution 3.0 Unported License http://creativecommons.org/licenses/by/3.0/ This licence applies to the selection, segmentation and alignment of the speech and text data. The underlying audio and text are licensed under their specific datasource terms. Please refer to the links below for a full description of them. If you use any part of the corpus in your work, please cite the following paper: A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 --------------------------------------------------------------- SPEECH AND TEXT SOURCES --------------------------------------------------------------- 1) Bulgarian - "Zhetvariat" by Yordan Yovkov audio: http://librivox.org/zhetvariat-by-yordan-yovkov text: http://slovo.bg/showwork.php3?AuID=95&WorkID=9610&Level=1 2)Danish - "Grimms eventyr I udvalg" by Grimm Brothers audio: http://librivox.org/grimms-eventyr-i-udvalg-by-br%C3%B8drene-grimm text: http://www.estrup.org/cms/?mod=text&id=392 3)Dutch- "Anna Karenina" by Leo Tolstoy audio: http://librivox.org/anna-karenina-by-leo-tolstoy text: http://www.gutenberg.org/ebooks/13214 4) English - "Living Alone" by Stella Benson audio: http://librivox.org/living-alone-by-stella-benson text: http://www.gutenberg.org/ebooks/14907 5) Finnish- "Rautatie" by Juhani Aho audio: http://librivox.org/rautatie-by-juhani-aho text: http://www.gutenberg.org/ebooks/10481 6) French - "Candide" by Voltaire audio: http://librivox.org/candide-by-voltaire text: http://www.gutenberg.org/cache/epub/4650/pg4650.txt 7) German - "Das Bildnis des Dorian Gray" by Oscar Wilde audio: http://librivox.org/das-bildnis-des-dorian-gray-by-oscar-wilde text: http://gutenberg.spiegel.de/buch/1836/1 8) Hungarian - "Egri csillagok" by Geza Gardonyi audio: http://gutenberg.spiegel.de/buch/1836/1 text: http://mek.oszk.hu/00600/00656/index.phtml 9) Italian - "Galatea" by Anton Giulio Barrili audio: http://librivox.org/galatea-by-anton-giulio-barrili/ text: http://www.gutenberg.org/ebooks/19427 10) Polish - "Siedem wybranyc opowiadan" by Wladyslaw Orkan audio: http://librivox.org/siedem-wybranych-opowiadan-by-wladyslaw-orkan/ text: http://pl.wikisource.org/wiki/Autor:W%C5%82adys%C5%82aw_Orkan 11) Portuguese - "Senhora" by Jose de Alencar audio: http://librivox.org/senhora-by-jose-de-alencar/ text: http://stat.correioweb.com.br/arquivos/educacao/arquivos/JosdeAlencar-Senhora0.pdf 12) Romanian - "Mara" by Ioan Slavici audio: http://speech.utcluj.ro/corpora/mara.html text: http://ro.wikisource.org/wiki/Mara 13) Russian - "Ucheniye Khrista" by Leo Tolstoy audio: http://librivox.org/teachings-of-christ-rus-by-leo-tolstoy/ text: http://az.lib.ru/t/tolstoj_lew_nikolaewich/text_0520.shtml 14) Spanish - "Don Quijote de la Mancha" by Miguel de Cervantes audio: http://www.quijote.es/IVCentenario_AudioLibro.php text: http://www.gutenberg.org/ebooks/5921 --------------------------------------------------------------- CONTENTS --------------------------------------------------------------- For each audiobook you can download the following information, with the exception of the Spanish audiobook which has a restricted use and the speech data cannot be downloaded from out website: 1) Segmented and aligned data -- http://tundra.simple4all.org/download.html -- an archive containing the results of the lightly supervised segmentation and alignment algorithm; -- folders and files (the following are the same for both training and test data sets): -- wav/ - speech data maintaining the original chapter names, but with additional indexes resulted from the sentence-level segmentation; -- txt/ - raw text files corresponding to each speech file from the wav/ folder; -- txtWithPunctuation/ - text files for each speech segment with punctuation restored from the original book text; -- speech_transcript.txt - a single file for all orthographic transcripts; -- a separate handmade test set data in the handmadeTest/ folder (see below for its description); 2) 1 hour subset of selected data -- http://tundra.simple4all.org/ssw8data.html -- an archive containing approximately 1 hour of selected audio used to train the voices from the Demo section; 3) Synthetic samples -- http://tundra.simple4all.org/ssw8data.html -- an archive containing the synthetic samples obtained with our lightly supervised TTS system for the handmade test set; 4) Chapter-level annotation -- http://tundra.simple4all.org/download.html -- a file with the chapter-level time alignment within the original data and the corresponding text for the confident data. --------------------------------------------------------------- SEGMENTATION AND ALIGNMENT --------------------------------------------------------------- Descriptions of the lightly supervised segmentation and alignment methods can be found in the following papers: 1) A. Stan, O. Watts, Y. Mamiya, M. Giurgiu, R. A. J. Clark, J. Yamagishi, S. King, TUNDRA: A Multilingual Corpus of Found Data for TTS Research Created with Light Supervision, In Proc. Interspeech, Lyon, France, August 2013 2) Adriana STAN, Peter BELL, Simon KING A grapheme-based method for automatic alignment of speech and text data, In Proc. IEEE Workshop on Spoken Language Technology, Miami, Florida, USA, December 2012 3) Yoshitaka Mamiya, Junichi Yamagishi, Oliver Watts, Robert A.J. Clark, Simon King and Adriana STAN Lightly Supervised GMM VAD to use Audiobook for Speech Synthesiser, In Proc. ICASSP, May 2013 The synthetic voice building algorithm is described in detail here: 4) O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from ‘found’ data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 With a similar approach being presented in: 5) O. Watts, A. Stan, A. Suni, M. Burgos, J.M. Montero, The Simple4All entry to the Blizzard Challenge 2013, Blizzard Challenge 2013 --------------------------------------------------------------- TRAIN/TEST DIVISION OF DATA --------------------------------------------------------------- Test material is taken from the ends of books, from enough whole chapters or stories to make up at least 10 min of audio of aligned data (NB more can be harvested from these chapters from the unaligned utterances). The exceptions are the Hungarian and Portuguese audiobooks in which the following chapters have variable recording conditions and are not considered suitable for comparisons: Hungarian: egricsillagok_[19-49] Portuguese: senhora_[14-20] and senhora_[23-41] The following chapters are reserved for testing: Bulgarian: zhetvariat_2{3,4,5}* Danish: eventyr_{08,09,10,11,12}* Dutch: annakarenina_021* German: doriangray_17* English: livingalone_{09,10}* Finnish: rautatie_{7,8}* French: candide_{29,30}* Hungarian: egricsillagok_{17,18}* Italian: galatea_{19,20}* Polish: siedemwybranchopowiadan_7* Portuguese: senhora_{12,13}* Romanian: mara_7{1,2}* Russian: teachingsofchrist_9* Spanish: Parte1_35* For the evaluations published in the following paper: O. Watts, A. Stan, R. Clark, Y. Mamiya, M. Giurgiu, J. Yamagishi, S. King, Unsupervised and lightly-supervised learning for rapid construction of TTS systems in multiple languages from 'found' data: evaluation and analysis, In Proc. SSW8, Barcelona, Spain, August 2013 a hand-segmented test set of about 40 utterances in all languages was prepared from the test chapters, so that various problems with the automatically aligned test utterances (sentence fragments, non-matching transcripts etc.) would not confuse the evaluation results. These hand segmented and aligned utterances are contained in the ./handmadeTest folder. Synthesised samples of this handmade test set are also available for download. --------------------------------------------------------------- CONTRIBUTORS --------------------------------------------------------------- Adriana Stan (Communications Department, Technical University of Cluj-Napoca) Oliver Watts (Centre for Speech Technology Research, University of Edinburgh) Yoshitaka Mamiya (Centre for Speech Technology Research, University of Edinburgh) Junichi Yamagishi (National Institute of Informatics, Tokyo) Mircea Giurgiu (Communications Department, Technical University of Cluj-Napoca) Rob Clark (Centre for Speech Technology Research, University of Edinburgh) Simon King (Centre for Speech Technology Research, University of Edinburgh) --------------------------------------------------------------- CONTACT --------------------------------------------------------------- Please send all you enquires regarding the Tundra Copus to one of the following e-mail addresses: adriana.stan@com.utcluj.ro owatts@inf.ed.ac.uk --------------------------------------------------------------- ACKNOWLEDGEMNTS --------------------------------------------------------------- The research leading to these results has received funding from the European Community's Seventh Framework Programme (FP7/2007-2013) under grant agreement No 287678 (the Simple4All project - http://www.simple4all.org) The research presented here has made use of the resources provided by the Edinburgh Compute and Data Facility (ECDF: http://www.ecdf.ed.ac.uk). The ECDF is partially supported by the eDIKT initiative (http://www.edikt.org.uk). We would like to thank Mihai Nae from Cartea Sonora for releasing the Romanian data, as well as to all the volunteers at Librivox and Gutenberg for dedicating their time to distribute this wide variety of data. </pre>
Multilingual Medical Corpora
<p>The amount of digital data derived from healthcare processes have increased tremendously in the last years. This applies especially to unstructured data, which are often hard to analyze due to the lack of available tools to process and extract information. Natural language processing is often used in medicine, but the majority of tools used by researchers are developed primarily for the English language. For developing and testing natural language processing methods, it is important to have a suitable corpus, specific to the medical domain that covers the intended target language. To improve the potential of natural language processing research, we developed tools to derive language specific medical corpora from publicly available text sources. n order to extract medicine-specific unstructured text data, openly available pub-lications from biomedical journals were used in a four-step process:(1) medical journal databases were scraped to download the articles,(2) the articles were parsed and consolidated into a single repository,(3) the content of the repository was de-scribed, and (4) the text data and the codes were released. In total, 93 969 articles were retrieved, with a word count of 83 868 501 in three different languages (German, English, and Spanish) from two medical journal databases Our results show that unstructured text data extraction from openly available medical journal databases for the construction of unified corpora of medical text data can be achieved through web scraping techniques.</p>
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis”, Proceedings of the 7th International Conference on Machine Learning and Natural Language Processing (MLNLP 2024), Chengdu, China, October 18-20, 2024 (Paper accepted for publication, Preprint available at: https://arxiv.org/abs/2410.03293)</p> <p> </p> <p><strong>Abstract</strong></p> <p>The outbreak of COVID-19 served as a catalyst for content creation and dissemination on social media platforms, as such platforms serve as virtual communities where people can connect and communicate with one another seamlessly. While there have been several works related to the mining and analysis of COVID-19-related posts on social media platforms such as Twitter (or X), YouTube, Facebook, and TikTok, there is still limited research that focuses on the public discourse on Instagram in this context. Furthermore, the prior works in this field have only focused on the development and analysis of datasets of Instagram posts published during the first few months of the outbreak. The work presented in this paper aims to address this research gap and presents a novel multilingual dataset of <strong>500,153 Instagram posts about COVID-19 published between January 2020 and September 2024</strong>. This dataset contains Instagram posts in <strong>161 different languages</strong>. After the development of this dataset, multilingual sentiment analysis was performed using VADER and twitter-xlm-roberta-base-sentiment. This process involved classifying each post as positive, negative, or neutral. The results of sentiment analysis are presented as a separate attribute in this dataset.</p> <p><em><strong>For each of these posts, the Post ID, Post Description, Date of publication, language code, full version of the language, and sentiment label are presented as separate attributes in the dataset.</strong></em></p> <p>The Instagram posts in this dataset are present in <strong>161 different languages</strong> out of which the top 10 languages in terms of frequency are English (343041 posts), Spanish (30220 posts), Hindi (15832 posts), Portuguese (15779 posts), Indonesian (11491 posts), Tamil (9592 posts), Arabic (9416 posts), German (7822 posts), Italian (5162 posts), Turkish (4632 posts)</p> <p>There are <strong>535,021 distinct hashtags in this dataset</strong> with the top 10 hashtags in terms of frequency being #covid19 (169865 posts), #covid (132485 posts), #coronavirus (117518 posts), #covid_19 (104069 posts), #covidtesting (95095 posts), #coronavirusupdates (75439 posts), #corona (39416 posts), #healthcare (38975 posts), #staysafe (36740 posts), #coronavirusoutbreak (34567 posts)</p> <p>The following is a description of the attributes present in this dataset</p> <ul> <li><em><strong>Post ID</strong></em>: Unique ID of each Instagram post</li> <li><em><strong>Post Description</strong></em>: Complete description of each post in the language in which it was originally published</li> <li><em><strong>Date</strong></em>: Date of publication in MM/DD/YYYY format</li> <li><em><strong>Language code</strong></em>: Language code (for example: “en”) that represents the language of the post as detected using the Google Translate API </li> <li><em><strong>Full Language</strong></em>: Full form of the language (for example: “English”) that represents the language of the post as detected using the Google Translate API </li> <li><em><strong>Sentiment</strong></em>: Results of sentiment analysis (using the preprocessed version of each post) where each post was classified as positive, negative, or neutral</li> </ul> <p><strong>Open Research Questions</strong></p> <p>This dataset is expected to be helpful for the investigation of the following research questions and even beyond:</p> <ol> <li>How does sentiment toward COVID-19 vary across different languages?</li> <li>How has public sentiment toward COVID-19 evolved from 2020 to the present?</li> <li>How do cultural differences affect social media discourse about COVID-19 across various languages?</li> <li>How has COVID-19 impacted mental health, as reflected in social media posts across different languages?</li> <li>How effective were public health campaigns in shifting public sentiment in different languages?</li> <li>What patterns of vaccine hesitancy or support are present in different languages?</li> <li>How did geopolitical events influence public sentiment about COVID-19 in multilingual social media discourse?</li> <li>What role does social media discourse play in shaping public behavior toward COVID-19 in different linguistic communities?</li> <li>How does the sentiment of minority or underrepresented languages compare to that of major world languages regarding COVID-19?</li> <li>What insights can be gained by comparing the sentiment of COVID-19 posts in widely spoken languages (e.g., English, Spanish) to those in less common languages?</li> </ol> <p>All the Instagram posts that were collected during this data mining process to develop this dataset were publicly available on Instagram and did not require a user to log in to Instagram to view the same (at the time of writing this paper).</p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p>
FIGURES 164–172 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 164–172. Genera of New World Scarabaeinae. 164. Sylvicanthon sp. 165. Sylvicanthon sp., pygidium (arrow indicates basal pygidial carina). 166. Sylvicanthon sp., dorsal view forebody. 167. Tesserodoniella elguetai Vaz-de-Mello & Halffter 168. Tetraechma sanguineomaculata Blanchard. 169. Tetramereia convexa Klages. 170. Trichillidium quadridens (Robinson). 171. Trichillum externepunctatum Preudhomme de Borre. Fig. 172. Uroxys epipleuralis Arrow
FIGURES 154–163 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 154–163. Genera of New World Scarabaeinae. 154. Scybalocanthon nigriceps (Harold) 155. Scybalocanthon nigriceps, metatarsus. 156. Scybalophagus rugosus (Blanchard) 157. Silvinha unica Vaz-de-Mello 158. Sinapisoma minutum (Laporte). 159. Sinapisoma minutum, hind leg. 160. Sisyphus mexicanus Harold. 161. Streblopus opatroides Lansberge. 162. Sulcophanaeus imperator (Chevrolat). 163. Sulcophanaeus menelas (Laporte).
FIGURES 137–145 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 137–145. Genera of New World Scarabaeinae. 137. Oxysternon (O.) conspicillatum (Weber). 138. Oxysternon (O.) palaemon Laporte. 139. Oxysternon (Mioxysternon) spiniferum Laporte. 140. Oxysternon sp., ventral view (arrow indicates metasternal spine). 141. Paracanthon sp. 142. Paracanthon sp., dorsal view of head. 143. Paracryptocanthon borgmeieri (Vulcano, Pereira & Martínez). 144. Pedaridium hirsutum Harold 145. Pereiraidium almeidai (Pereira)
FIGURES 121–127 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 121–127. Genera of New World Scarabaeinae. 121. Ontherus (Planontherus) bridgesi Waterhouse. 122. Ontherus (P.) bridgesi, ventral view of thorax. 123. Ontherus (Caelontherus) alexis (Blanchard). 124. Ontherus (O.) digitatus Harold, ventral view of thorax. 125. Ontherus (O.) sulcator (Fabricius). 126. Ontherus (O.) sulcator, ventral view of thorax (arrow indicates mesepisternal carina). 127. Ontherus (C.) alexis (Blanchard), ventral view (arrow indicates meso-metasternal suture).
FIGURES 93–102 in A multilingual key to the genera and subgenera of the subfamily Scarabaeinae of the New World (Coleoptera: Scarabaeidae) 2854
FIGURES 93–102. Genera of New World Scarabaeinae. 93. Genieridium cryptops (Arrow). 94. Glyphoderus sterquilinus (Westwood). 95. Gromphas lacordairei Brullé. 96. Hansreia affinis (Fabricius). 97. Holocanthon mateui Martínez & Pereira. 98. Holocanthon mateui (arrow indicates break in mentum). 99. Holocephalus eridanus (Olivier). 100. Holocephalus eridanus, antenna. 101. Homalotarsus impressus Janssens. 102. Homalotarsus impressus, metatarsus.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.