Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
135
datasets available to search
ShareScore release 0.7.1
Dataset results
135 results for “multilingual”
Dataset from the ISMIR2020 article "Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Item Annotation"
<p>We release the data required to reproduce the cross-lingual music genre translation experiments from the article <strong><em>Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Item Annotation</em></strong> presented at the <a href="https://ismir.github.io/ISMIR2020/">ISMIR 2020</a> conference.</p> <p>More information about this data and how it should be used in the experiments can be found in the GitHub repository <a href="http://github.com/deezer/MultilingualMusicGenreEmbedding">deezer/MultilingualMusicGenreEmbedding</a>.</p> <p>Please cite our paper if you use the code or data in your work.</p>
Wikidata: Persistent identifiers as the basis for multilingual and human-machine collaboration
<p>This repository hosts a video recording of a dry-run of the PIDapalooza 2021 <a href="https://www.pidapalooza.org/schedule">https://www.pidapalooza.org/schedule</a> session</p> <p><strong>"Wikidata: Persistent identifiers as the basis for multilingual and human-machine collaboration"</strong><strong> </strong></p> <p>taking place Thursday January 28, 2021 00:00 - 00:30 UTC on Stage 1, as per <a href="https://sched.co/gD2n">https://sched.co/gD2n</a> .</p> <p>The video is also available on YouTube via <a href="https://youtu.be/g5VCr--Q1Ig">https://youtu.be/g5VCr--Q1Ig</a> .</p> <p>See <a href="https://etherpad.wikimedia.org/p/zenodo.4253308">https://etherpad.wikimedia.org/p/zenodo.4253308</a> for notes and <a href="https://github.com/Daniel-Mietchen/events/blob/master/PIDapalooza-2021.md">https://github.com/Daniel-Mietchen/events/blob/master/PIDapalooza-2021.md</a> for background on the session.</p> <p>Resources demoed in the recording:</p> <ul> <li>Wikidata: <a href="https://www.wikidata.org/wiki/Q420330">https://www.wikidata.org/wiki/Q420330</a> - persistent identifier</li> <li>Scholia: <a href="https://scholia.toolforge.org/event/Q47486859">https://scholia.toolforge.org/event/Q47486859</a> - PIDapalooza 2018 </li> <li>Ordia: <a href="https://ordia.toolforge.org/">https://ordia.toolforge.org/</a> <ul> <li><a href="https://ordia.toolforge.org/text-to-languages">https://ordia.toolforge.org/text-to-languages</a> </li> <li><a href="https://ordia.toolforge.org/text-to-lexemes">https://ordia.toolforge.org/text-to-lexemes</a> <ul> <li><a href="https://www.wikidata.org/wiki/Lexeme:L407266">https://www.wikidata.org/wiki/Lexeme:L407266</a> - biocapacity</li> </ul> </li> </ul> </li> </ul> <ul> <li>Lingua Libre: <a href="https://lingualibre.org/">https://lingualibre.org/</a> <ul> <li><a href="https://w.wiki/wBL">https://w.wiki/wBL</a> - audio recording example</li> </ul> </li> <li>Sample content: <ul> <li><a href="https://doi.org/10.3389/fcosc.2020.615419">https://doi.org/10.3389/fcosc.2020.615419</a> - Underestimating the Challenges of Avoiding a Ghastly Future</li> <li><a href="https://en.wikipedia.org/wiki/Persistent_identifier">https://en.wikipedia.org/wiki/Persistent_identifier</a> </li> </ul> </li> <li>Testing: <a href="https://www.wikidata.org/wiki/Q13406268">https://www.wikidata.org/wiki/Q13406268</a> - Wikidata Sandbox 2</li> </ul> <p>Also relevant:</p> <ul> <li><a href="https://w.wiki/hmb">https://w.wiki/hmb</a> - Items with Disease Ontology ID and MeSH Descriptor ID and optional descriptions in multiple languages</li> <li><a href="https://en.wikiquote.org/wiki/Identity">https://en.wikiquote.org/wiki/Identity</a> - quotes around "identity"</li> </ul> <ul> </ul>
Multilingual Knowledge Graph Completion With Joint Relation and Entity Alignment
<p>Code and data accompanying AlignKGC.</p>
Multilingual Bottle-Neck Feature Learning from Untranscribed data for track 1 in zerospeech2017 (system 1 -- without VTLN)
<p>We investigate the extraction of bottle-neck features (BNFs) for multiple languages without access to manual transcription. Multilingual BNFs are derived from a multi-task learning deep neural network which is trained with unsupervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models separately trained on untranscribed speech of multiple languages.</p>
Glosario: A multilingual glossary for computing and data science terms.
<p><code>glosario</code> is an open-source glossary of terms used in data science that is available online and also as a library in both <a href="https://github.com/carpentries/glosario-r/">R</a> and <a href="https://github.com/carpentries/glosario-py/">Python</a>. By adding glossary keys to a lesson’s metadata, authors can indicate what the lesson teaches, what learners ought to know before they start, and where they can go to find that knowledge. Authors can also use the library’s functions to insert consistent hyperlinks for terms and definitions in their lessons in any of several languages. The master copy of the glossary lives in the <code>glossary.yml</code> file. </p>
TRIDIS: HTR model for Multilingual Medieval and Early Modern Documentary Manuscripts (11th-16th)
<p><strong>TRIDIS (Tria Digita Scribunt)</strong> is a Handwriting Text Recognition model trained on semi-diplomatic transcriptions from medieval and Early Modern Manuscripts. It is suitable for work on documentary manuscripts, that is, manuscripts arising from legal, administrative, and memorial practices more commonly from the Late Middle Ages (13th century and onwards). It can also show good performance on documents from other domains, such as literature books, scholarly treatises and cartularies providing a versatile tool for historians and philologists in transforming and analyzing historical texts.</p> <p>A paper presenting the first version of the model is available here: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval Manuscripts. </strong>Journal of Data Mining and Digital Humanities.<strong> </strong>2023. https://hal.science/hal-03892163</p> <p> </p> <h3>Transcriptions rules :</h3> <p>Since the majority of the training documents come from diplomatic editions, the transcriptions were <strong>normalized</strong> to contemporary reading standards, and <strong>abbreviations were expanded</strong> with the aim of facilitating a more fluid reading of the document.</p> <p>The following rules were applied:</p> <ul> <li>The abbreviations have been expanded, both those by suspension (<code>facimꝰ</code> ---> <code>facimus</code>) and by contraction (<code>dñi</code> --> <code>domini</code>). Likewise, those using conventional signs (<code>⁊</code> --> <code>et</code> ; <code>ꝓ</code> --> <code>pro</code>) have been resolved. </li> <li>The named entities (names of persons, places and institutions) have been <code>capitalized</code>. The beginning of a block of text as well as the original capitals used by the scribe are also capitalized.</li> <li>The consonantal <code>i</code> and <code>u</code> characters have been transcribed as <code>j</code> and <code>v</code> in both French and Latin.</li> <li>The punctuation marks used in the manuscript like: <code>.</code> or <code>/</code> or <code>|</code> have not been systematically transcribed as the transcription has been standardized with modern punctuation.</li> <li>Corrections and words that appear cancelled in the manuscript have been transcribed surrounded by the sign <code>$</code> at the beginning and at the end.</li> </ul> <p> </p> <h3>Versions :</h3> <p><strong>Version 1 </strong>of the model was trained on charters and registers dataset from the Late Medieval period (12th-15th centuries). The training and evaluation involved 1855 pages, 120k lines of text, and almost 1M tokens, conducted using three freely available ground-truth corpora:</p> <ul> <li>The Alcar-HOME database: <a href="../record/5600884" target="_new">https://zenodo.org/record/5600884</a></li> <li>The e-NDP corpus: <a href="../record/7575693" target="_new">https://zenodo.org/record/7575693</a></li> <li>The Himanis project: <a href="../record/5535306" target="_new">https://zenodo.org/record/5535306</a></li> </ul> <p><strong>Version 2</strong> of the model has added new datasets from feudal books and legal proceedings (14th-16th centuries), incorporating an additional 115k lines and more than 1.2M tokens to the previous version using other corpora like:</p> <ul> <li>Königsfelden Abbey corpus: <a href="../record/5179361" target="_new">https://zenodo.org/record/5179361</a></li> <li>Monumenta Luxemburgensia.</li> </ul> <p> </p> <h3>Accuracy</h3> <p>TRIDIS was trained using a CNN+RNN+CTC architecture within the Kraken suite (https://kraken.re/). This final model operates in a multilingual environment (Latin, Old French, and Old Spanish) and is capable of recognizing several Latin script families (mostly Textualis and Cursiva) in documents produced circa 11th - 16th centuries. During evaluation, the model showed an accuracy of 93.1% on the validation set and a CER (Character Error Ratio) of about 0.11 to 0.15 on four external unseen datasets. Fine-tuning the model with 10 ground-truth pages can improve these results to a CER of between 0.06 to 0.10, respectively.</p> <h3>Other formats</h3> <p>The ground truth used for version 2 was also employed to train a Transformer HTR model that combines TrOCR as the encoder with a RoBERTa medieval model as the decoder. This model exhibits a slighly better performance in terms of CER metrics to the current TRIDIS version and shows an improved WER by about 25%. The model is available on the Hugging Face Hub: <a href="https://huggingface.co/magistermilitum/tridis_HTR">magistermilitum/tridis_HTR</a></p>
Data set for "Token-Level Multilingual Epidemic Dataset for Event Extraction"
<p>This is the data for the TPDL 2021 paper "<a href="https://zenodo.org/record/5780020">Token-Level Multilingual Epidemic Dataset for Event Extraction</a>". If you use this resource, please cite the paper:</p> <pre><code>@inproceedings{mutuvi2021dataset, title = "Token-level Multilingual Epidemic Dataset for Event Extraction", author = {Mutuvi, Stephen and Boros, Emanuela and Doucet, Antoine, and Lejeune, Gaël and Jatowt, Adam and Odeo, Moses}, booktitle = "Proceedings of the 25th International Conference on Theory and Practice of Digital Libraries, September 13–17, 2021, TPDL 2021", year = "2021", location = "Online" }</code></pre> <p> </p> <p>This work has been supported by the European Union Horizon 2020 research and innovation programme under grants 825153 (Embeddia) and 770299 (NewsEye).</p>
Fairlex: A multilingual benchmark for evaluating fairness in legal text processing
<p>We present a benchmark suite of four datasets for evaluating the fairness of pre-trained legal language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA, Swiss, and Chinese), five languages (English, German, French, Italian, and Chinese), and fairness across five attributes (gender, age, nationality/region, language, and legal area). In our experiments, we evaluate pre-trained language models using several group-robust fine-tuning techniques and show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparities. Furthermore, we provide a quantitative and qualitative analysis of our results, highlighting open challenges in the development of robustness methods in legal NLP.</p>
WikiProject Clinical Trials for multilingual access to information
<p>Watch at <a href="https://www.youtube.com/watch?v=uJbn0dPqAE8">https://www.youtube.com/watch?v=uJbn0dPqAE8</a></p> <p>WikiProject Clinical Trials is a wiki community project to increase access to medical research metadata. Check it out at <a href="https://www.wikidata.org/wiki/Wikidata:WikiProject_Clinical_Trials">https://www.wikidata.org/wiki/Wikidata:WikiProject_Clinical_Trials</a></p> <p>Here I argue that everyone has a right to access medical research metadata and that for public interest, we need to translate basic information from trials into many languages. I piloted this process in Wikidata.</p>
FAIR and Open multilingual clinical trials in Wikidata and Wikipedia
<p>"FAIR and Open multilingual clinical trials in Wikidata and Wikipedia" was a December 2021 presentation by Lane Rasberry and Cherrie Kwok. It was made at the conference "Understanding Wikipedia’s Dark Matter - Translation and Multilingual Practice in the World'’s Largest Online Encyclopaedia" hosted by the Centre for Translation and the Department of Translation, Interpreting and Intercultural Studies, both at Hong Kong Baptist University.</p> <ul> <li>conference page <a href="https://ctn.hkbu.edu.hk/wikiconf2021/">https://ctn.hkbu.edu.hk/wikiconf2021/</a></li> <li>watch video <a href="https://www.youtube.com/watch?v=5yRhCENeezQ">https://www.youtube.com/watch?v=5yRhCENeezQ</a></li> <li>slides archive <a href="https://commons.wikimedia.org/wiki/File:FAIR_and_Open_multilingual_clinical_trials_in_Wikidata_and_Wikipedia.pdf">https://commons.wikimedia.org/wiki/File:FAIR_and_Open_multilingual_clinical_trials_in_Wikidata_and_Wikipedia.pdf</a></li> </ul> <p> </p> <p> </p> <p> </p>
Multilingual named entity recognition for medieval charters. Datasets and models
<p>Annotated dataset for training named entities recognition models for medieval charters in Latin, French and Spanish.</p> <p> </p> <p>The original raw texts for all charters were collected from four charters collections</p> <p>- HOME-ALCAR corpus : <a href="https://zenodo.org/record/5600884">https://zenodo.org/record/5600884</a></p> <p>- CBMA : <a href="https://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd=&ved=2ahUKEwisvNOa3qP3AhULyIUKHenpBoAQFnoECA0QAQ&url=http%3A%2F%2Fwww.cbma-project.eu%2F&usg=AOvVaw0blsASXzOKSNz_EkixwfJT">http://www.cbma-project.eu</a></p> <p>- Diplomata Belgica : <a href="https://www.diplomata-belgica.be">https://www.diplomata-belgica.be</a></p> <p>- CODEA corpus :<a href="http://https://corpuscodea.es/"> https://corpuscodea.es/</a></p> <p> </p> <p>We include (i) the annotated training datasets, (ii) the contextual and static embeddings trained on medieval multilingual texts and (iii) the named entity recognition models trained using two architectures: Bi-LSTM-CRF + stacked embeddings and fine-tuning on Bert-based models (mBert and RoBERTa)</p> <p>Codes, datasets and notebooks used to train models can be consulted in our gitlab repository: <a href="https://gitlab.com/magistermilitum/ner_medieval_multilingual">https://gitlab.com/magistermilitum/ner_medieval_multilingual</a></p> <p>Our best RoBERTa model is also available in the HuggingFace library: <a href="https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner">https://huggingface.co/magistermilitum/roberta-multilingual-medieval-ner</a></p>
CT-FAN: A Multilingual dataset for Fake News Detection
<p><strong>By downloading the data, you agree with the terms & conditions mentioned below:</strong></p> <p><strong>Data Access: </strong>The data in the research collection may only be used for research purposes. Portions of the data are copyrighted and have commercial value as data, so you must be careful to use them only for research purposes. </p> <p>Summaries, analyses and interpretations of the linguistic properties of the information may be derived and published, provided it is impossible to reconstruct the information from these summaries. You may not try identifying the individuals whose texts are included in this dataset. You may not try to identify the original entry on the fact-checking site. You are not permitted to publish any portion of the dataset besides summary statistics or share it with anyone else.</p> <p>We grant you the right to access the collection's content as described in this agreement. You may not otherwise make unauthorised commercial use of, reproduce, prepare derivative works, distribute copies, perform, or publicly display the collection or parts of it. You are responsible for keeping and storing the data in a way that others cannot access. The data is provided free of charge.</p> <p><strong>Citation</strong></p> <p>Please cite our work as</p> <pre>@InProceedings{clef-checkthat:2022:task3, author = {K{\"o}hler, Juliane and Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Wiegand, Michael and Siegel, Melanie and Mandl, Thomas}, title = "Overview of the {CLEF}-2022 {CheckThat}! Lab Task 3 on Fake News Detection", year = {2022}, booktitle = "Working Notes of CLEF 2022---Conference and Labs of the Evaluation Forum", series = {CLEF~'2022}, address = {Bologna, Italy},} </pre> <pre>@article{shahi2021overview, title={Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection}, author={Shahi, Gautam Kishore and Stru{\ss}, Julia Maria and Mandl, Thomas}, journal={Working Notes of CLEF}, year={2021} }</pre> <p><strong>Problem Definition:</strong> Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other (e.g., claims in dispute) and detect the topical domain of the article. This task will run in <strong>English and German.</strong></p> <p><strong>Task 3:</strong> <strong>Multi-class fake news detection of news articles (English)</strong> Sub-task A would detect fake news designed as a four-class classification problem. Given the text of a news article, determine whether the main claim made in the article is true, partially true, false, or other. The training data will be released in batches and roughly about 1264 articles with the respective label in English language. Our definitions for the categories are as follows:</p> <ul> <li> <p>False - The main claim made in an article is untrue.</p> </li> <li> <p>Partially False - The main claim of an article is a mixture of true and false information. The article contains partially true and partially false information but cannot be considered 100% true. It includes all articles in categories like partially false, partially true, mostly true, miscaptioned, misleading etc., as defined by different fact-checking services.</p> </li> <li> <p>True - This rating indicates that the primary elements of the main claim are demonstrably true.</p> </li> <li> <p>Other- An article that cannot be categorised as true, false, or partially false due to a lack of evidence about its claims. This category includes articles in dispute and unproven articles.</p> </li> </ul> <p><strong>Cross-Lingual Task (German)</strong></p> <p>Along with the multi-class task for the English language, we have introduced a task for low-resourced language. We will provide the data for the test in the German language. The idea of the task is to use the English data and the concept of transfer to build a classification model for the German language.</p> <p><strong>Input Data</strong></p> <p>The data will be provided in the format of Id, title, text, rating, the domain; the description of the columns is as follows:</p> <ul> <li>ID- Unique identifier of the news article</li> <li>Title- Title of the news article</li> <li>text- Text mentioned inside the news article</li> <li>our rating - class of the news article as false, partially false, true, other</li> </ul> <p><strong>Output data format</strong></p> <ul> <li>public_id- Unique identifier of the news article</li> <li>predicted_rating- predicted class</li> </ul> <p>Sample File</p> <pre><code>public_id, predicted_rating 1, false 2, true</code></pre> <p><strong>IMPORTANT! </strong></p> <ol> <li>We have used the data from 2010 to 2022, and the content of fake news is mixed up with several topics like elections, COVID-19 etc.</li> </ol> <p><strong>Baseline:</strong> For this task, we have created a baseline system. The baseline system can be found at <a href="https://zenodo.org/record/6362498">https://zenodo.org/record/6362498</a></p> <p><strong>Related Work</strong></p> <ul> <li>Shahi GK. AMUSED: An Annotation Framework of Multi-modal Social Media Data. arXiv preprint arXiv:2010.00502. 2020 Oct 1.<a href="https://arxiv.org/pdf/2010.00502.pdf">https://arxiv.org/pdf/2010.00502.pdf</a></li> <li>G. K. Shahi and D. Nandini, “FakeCovid – a multilingual cross-domain fact check news dataset for covid-19,” in workshop Proceedings of the 14th International AAAI Conference on Web and Social Media, 2020. <a href="http://workshop-proceedings.icwsm.org/abstract?id=2020_14">http://workshop-proceedings.icwsm.org/abstract?id=2020_14</a></li> <li>Shahi, G. K., Dirkson, A., & Majchrzak, T. A. (2021). An exploratory study of covid-19 misinformation on twitter. <em>Online Social Networks and Media</em>, <em>22</em>, 100104. doi: <a href="https://dx.doi.org/10.1016%2Fj.osnem.2020.100104">10.1016/j.osnem.2020.100104</a></li> <li>Shahi, G. K., Struß, J. M., & Mandl, T. (2021). Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection. <em>Working Notes of CLEF</em>.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeno, A., Míguez, R., Shaar, S., ... & Mandl, T. (2021, March). The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news. In <em>European Conference on Information Retrieval</em> (pp. 639-649). Springer, Cham.</li> <li>Nakov, P., Da San Martino, G., Elsayed, T., Barrón-Cedeño, A., Míguez, R., Shaar, S., ... & Kartal, Y. S. (2021, September). Overview of the CLEF–2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In <em>International Conference of the Cross-Language Evaluation Forum for European Languages</em> (pp. 264-291). Springer, Cham.</li> </ul>
Datasets for the paper: Lost in Translation: Using Global Fact-Checks to Measure Multilingual Misinformation Prevalence, Spread, and Evolution
<p>FullData.csv.gz: Contains links to all claims in the data-set.</p> <ul> <li>publishing_date: Date on which the fact-check was published.</li> <li>claim_date: Date that claim was made.</li> <li>verdict: Rating given by the fact-checking organisation.</li> <li>language: Language of the claim.</li> <li>cluster_{threshold}: ID of the cluster that claim belongs to at all given clusters. Entry "0" means that claim is singleton and not clustered with any other claims.</li> </ul> <p>Embeddings.npy: Contains a dictionary linking each claim to it's embedding calculated with LaBSE.</p>
MultiCardioNER Corpus: Multilingual Adaptation of Clinical NER Systems to the Cardiology Domain
<h1><strong>MultiCardioNER</strong></h1> <p><strong>MultiCardioNER</strong> is a shared task about the adaptation of clinical NER systems to the cardiology domain. It uses a combination of two existing datasets (DisTEMIST for diseases and the newly-released DrugTEMIST for medications), as well as a new, smaller dataset of cardiology clinical cases annotated using the same guidelines.</p> <p>Participants are provided DisTEMIST and DrugTEMIST as training data to use as they see fit (1,000 documents, with the original partitions splitting them into 750 for training and 250 for testing). The cardiology clinical cases (cardioccc) are meant to be used as a development or validation set (258 documents), although participants are encourage to experiment with the documents and annotations as they see fit. The evaluation is done using a different collection of cardiology clinical cases (250).</p> <p>MultiCardioNER proposes two tracks:</p> <p>- Track 1: Spanish adaptation of disease recognition systems to the cardiology domain.<br>- Track 2: Multilingual (Spanish, English and Italian) adaptation of medication recognition systems to the cardiology domain.</p> <p>Please read the README file attached for more information on folder structure and file format.</p> <p><strong>MultiCardioNER</strong> was developed by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis and used as part of BioASQ 2024. For more information on the corpus, annotation scheme and task in general, please visit: <a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">https://temu.bsc.es/multicardioner</a>. This task is promoted by Spanish and European projects such as DataTools4Heart, AI4HF, BARITONE and AI4ProfHealth.</p> <p><strong>UPDATE MAY 28th 2024: </strong>The test set annotations are now out! We've also included the original background set files, as well as a file with the mappings from the masked filenames used during the evaluation phase to the original filenames. Please check the README for more information.</p> <h2><strong>Resources</strong></h2> <ul> <li><a href="https://temu.bsc.es/multicardioner" target="_blank" rel="noopener">MultiCardioNER website</a></li> <li><a href="http://bioasq.org/" target="_blank" rel="noopener">BioASQ website</a></li> <li><a href="../doi/10.5281/zenodo.6458078" target="_blank" rel="noopener">DisTEMIST Guidelines</a></li> <li><a href="../doi/10.5281/zenodo.11065432" target="_blank" rel="noopener">DrugTEMIST Guidelines</a></li> </ul> <h2><strong>License</strong></h2> <p>This work is licensed under a <a href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.</p> <h2><strong>Contact</strong></h2> <p>If you have any questions or suggestions, please contact us at:</p> <p>- Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)<br>- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)</p> <h2><strong>Additional resources and corpora</strong></h2> <p>If you are interested in MultiCardioNER, you might want to check out these corpora and resources:</p> <ul> <li><a href="../records/7614764">DisTEMIST</a> (Corpus of disease mentions and normalization to SNOMED CT)</li> <li><a href="../records/8224056">MedProcNER </a>(Corpus of clinical procedure mentions and normalization to SNOMED CT)</li> <li><a href="../records/10635215">SympTEMIST</a> (Corpus of clinical findings and normalization to SNOMED CT)</li> <li><a href="../records/4270158">PharmaCoNER</a> (Corpus of medications, drugs, chemical substances, genes, proteins and vaccine mentions and normalization)</li> <li><a href="../records/7116201">MEDDOPROF</a> (Corpus of mentions of professions, occupations and working status and normalization)</li> <li><a href="../records/8403498">MEDDOPLACE</a> (Corpus of mentions of place-related entity mentions, including departments, nationalities or patient movements etc.. and normalization)</li> <li><a href="../records/4279323">MEDDOCAN</a> (Corpus of mentions of Personal Health Identifiers (PHI))</li> <li><a href="../records/3978041">CANTEMIST</a> (Corpus of cancer tumor morphology mentions and normalization)</li> <li><a href="../records/3837305">CodiESP</a> (Corpus of clinical case reportes with assigned clinical codes from ICD10, Spanish version)</li> <li><a href="../records/7684093">LivingNER</a> (Corpus of mentions of species, including human/family members, pathogens, food, etc.. and normalization to NCBI Taxonomy)</li> <li><a href="../records/2560344">SPACCC-POS</a> (Corpus of clinical case reports in Spanish annotated with POS-tags)</li> <li><a href="../records/2560338">SPACCC-TOKEN</a> (Corpus of clinical case reports in Spanish annotated with token-tags (word mention boundaries))</li> <li><a href="../records/2560338">SPACCC-SPLIT</a> (Corpus of clinical case reports in Spanish annotated with sentence boundary-tags)</li> <li><a href="../records/5602914">MESINESP-2</a> (Corpus of manually indexed records with DeCS /MeSH terms comprising scientific literature abstracts, clinical trials, and patent abstracts)</li> </ul>
Multilingual Fake News Detection Dataset: Gujarati, Hindi, Marathi, and Telugu
<p>This dataset is designed to support research in fake news detection across four major Indian languages: Gujarati, Hindi, Marathi, and Telugu. The dataset includes a diverse set of news articles collected from various sources, each labeled as either 'fake' or 'real'. The primary goal is to provide a resource that helps in the development and evaluation of natural language processing (NLP) models capable of detecting fake news in these regional languages.</p>
CLDF dataset derived from Chin's "Gelong Language in the Multilingual Hub of Hainan" from 2015
<p>Cite the source of the dataset as:</p> <blockquote> <p>Chin, Andy C. (2015): The Gelong Language in the Multilingual Hub of Hainan. Bulletin of Chinese Linguistics. 8. 140-156.</p> </blockquote>
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
<p><strong>Please cite the following paper when using this dataset</strong>:</p> <p>N. Thakur, “Mpox narrative on Instagram: A labeled multilingual dataset of Instagram posts on mpox for sentiment, hate speech, and anxiety analysis,” arXiv [cs.LG], 2024, URL: https://arxiv.org/abs/2409.05292</p> <p><strong>Abstract</strong></p> <p>The world is currently experiencing an outbreak of mpox, which has been declared a Public Health Emergency of International Concern by WHO. During recent virus outbreaks, social media platforms have played a crucial role in keeping the global population informed and updated regarding various aspects of the outbreaks. As a result, in the last few years, researchers from different disciplines have focused on the development of social media datasets focusing on different virus outbreaks. No prior work in this field has focused on the development of a dataset of Instagram posts about the mpox outbreak. The work presented in this paper (stated above) aims to address this research gap. It presents this <strong>multilingual dataset of</strong> <strong>60,127 Instagram posts</strong> about mpox, published between <strong>July 23, 2022, and September 5, 2024</strong>. This dataset contains Instagram posts about mpox in <strong>52 languages</strong>. For each of these posts, the Post ID, Post Description, Date of publication, language, and translated version of the post (translation to English was performed using the Google Translate API) are presented as separate attributes in the dataset.</p> <p>After developing this dataset, sentiment analysis, hate speech detection, and anxiety or stress detection were also performed. This process included classifying each post into</p> <ul> <li>one of the fine-grain sentiment classes, i.e., <strong>fear, surprise, joy, sadness, anger, disgust, or neutral</strong>, </li> <li><strong>hate or not hate</strong></li> <li><strong>anxiety/stress detected or no anxiety/stress detected</strong>.</li> </ul> <p>These results are presented as separate attributes in the dataset for the training and testing of machine learning algorithms for sentiment, hate speech, and anxiety or stress detection, as well as for other applications. </p> <p><strong>The 52 distinct languages in which Instagram posts are present in the dataset </strong><strong>are </strong>English, Portuguese, Indonesian, Spanish, Korean, French, Hindi, Finnish, Turkish, Italian, German, Tamil, Urdu, Thai, Arabic, Persian, Tagalog, Dutch, Catalan, Bengali, Marathi, Malayalam, Swahili, Afrikaans, Panjabi, Gujarati, Somali, Lithuanian, Norwegian, Estonian, Swedish, Telugu, Russian, Danish, Slovak, Japanese, Kannada, Polish, Vietnamese, Hebrew, Romanian, Nepali, Czech, Modern Greek, Albanian, Croatian, Slovenian, Bulgarian, Ukrainian, Welsh, Hungarian, and Latvian. </p> <p>The following table represents the data description for this dataset</p> <table> <tbody> <tr> <td> <p><strong>Attribute Name</strong></p> </td> <td> <p><strong>Attribute Description</strong></p> </td> </tr> <tr> <td> <p>Post ID</p> </td> <td> <p>Unique ID of each Instagram post</p> </td> </tr> <tr> <td> <p>Post Description</p> </td> <td> <p>Complete description of each post in the language in which it was originally published</p> </td> </tr> <tr> <td> <p>Date</p> </td> <td> <p>Date of publication in MM/DD/YYYY format</p> </td> </tr> <tr> <td> <p>Language</p> </td> <td> <p>Language of the post as detected using the Google Translate API</p> </td> </tr> <tr> <td> <p>Translated Post Description</p> </td> <td> <p>Translated version of the post description. All posts which were not in English were translated into English using the Google Translate API. No language translation was performed for English posts.</p> </td> </tr> <tr> <td> <p>Sentiment</p> </td> <td> <p>Results of sentiment analysis (using translated Post Description) where each post was classified into one of the sentiment classes: fear, surprise, joy, sadness, anger, disgust, and neutral</p> </td> </tr> <tr> <td> <p>Hate</p> </td> <td> <p>Results of hate speech detection (using translated Post Description) where each post was classified as hate or not hate</p> </td> </tr> <tr> <td> <p>Anxiety or Stress</p> </td> <td> <p>Results of anxiety or stress detection (using translated Post Description) where each post was classified as stress/anxiety detected or no stress/anxiety detected.</p> </td> </tr> </tbody> </table>
Weekly supervised Multilingual Data Set to train Named Entity Recognition for Symptom Extraction
<p>Data Sets were generated using the Weakly Supervised NER pipeline (https://github.com/HUMADEX/Weekly-Supervised-NER-pipline) to train the symptom extraction NER models. </p> <p><strong>Supported Languages and dataset locations for the specific language:</strong></p> <p> English (base language): https://huggingface.co/HUMADEX/english_medical_ner<br> German: https://huggingface.co/HUMADEX/german_medical_ner<br> Italian: https://huggingface.co/HUMADEX/italian_medical_ner<br> Spanish: https://huggingface.co/HUMADEX/spanish_medical_ner<br> Greek: https://huggingface.co/HUMADEX/german_medical_ner<br> Slovenian: https://huggingface.co/HUMADEX/slovenian_medical_ner<br> Polish: https://huggingface.co/HUMADEX/polish_medical_ner<br> Portuguese: https://huggingface.co/HUMADEX/portugese_medical_ner</p> <p> </p> <p><strong>Dataset Building </strong></p> <ul> <li>Data Integration and Preprocessing</li> <li>Data Cleaning</li> <li>Annotation with Stanza's i2b2 Clinical Model </li> <li>Translation into the targeted language</li> <li>Word Alignment </li> <li>Data Augmentation </li> </ul> <p><strong>Acknowledgement</strong><br>This dataset had been created as part of joint research of HUMADEX research group (https://www.linkedin.com/company/101563689/) and has received funding by the European Union Horizon Europe Research and Innovation Program project SMILE (grant number 101080923) and Marie Skłodowska-Curie Actions (MSCA) Doctoral Networks, project BosomShield ((rant number 101073222). Responsibility for the information and views expressed herein lies entirely with the authors.</p> <p><strong>Authors:</strong><br>dr. Izidor Mlakar, Rigona Sallauka, dr. Umut Arioz, dr. Matej Rojc</p> <p><strong>Please cite as:</strong></p> <p><span>Article title: Weakly-Supervised Multilingual Medical NER For Symptom Extraction For Low-Resource Languages</span><br><span>Doi: 10.20944/preprints202504.1356.v1</span><br><span>Website: </span><a title="https://www.preprints.org/manuscript/202504.1356/v1" href="https://www.preprints.org/manuscript/202504.1356/v1">https://www.preprints.org/manuscript/202504.1356/v1</a></p>
Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana (dataset)
<p>The dataset contains all the data required to reproduce the experiments done in the paper "Automatic translation and multilingual cultural heritage retrieval: a case study with transcriptions in Europeana", published in the 25th International Conference on Theory and Practice of Digital Libraries (<a href="http://www.tpdl.eu/tpdl2021/">TPDL'21</a>). In that work we run an experiment using the Europeana CH digital library as a use case, and we evaluated the effectiveness of a multilingual information retrieval strategy using machine translations to English as pivot language. We used the CEF translation service (eTranslation) for the translation of queries and content to English (<a href="https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation">https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation</a>).</p> <p>The dataset is also available at <a href="https://rnd-2.eanadev.org/share/crosslingual-search/">https://rnd-2.eanadev.org/share/crosslingual-search/</a>, and it is organized in four main folders:</p> <ul> <li><strong>queries</strong>: sample of 68 queries and their translations to English. The queries were issued in languages other than English from the Europeana Portal, using the Europeana’s 1914-1918 thematic collection, between January and August 2019.</li> <li><strong>transcriptions</strong>: sample of 18,257 handwriting transcriptions and its translations to English. The transcriptions are taken from the Europeana 1914-1918 thematic collection, and obtained from the Transcribathon crowdsourcing platform (https://europeana.transcribathon.eu/).</li> <li><strong>solr_configuration</strong>: Apache Solr search engine configuration used in the experiments (which replicates the one used in Europeana).</li> <li><strong>results</strong>: manual evaluation of the query translations, and automatic evaluation of the multilingual retrieval.</li> </ul> <p> </p>
Two Multilingual Students in Tandem Language Exchange
<p>Data from research on two multilingual students in a languahe tandem course: two interviews, two final reflections and two tandem learning diaries.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.