Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

27 results for “semeval”

Learn how ShareScore rates datasets ↗
zenodo52/100

SemEval-2022 Task 8: Multilingual news article similarity

<p>This dataset contains pairs of news articles drawn from the first half of 2020 and annotated for seven aspects of similarity:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong>&nbsp;Do the articles have similar tones?</li> </ul> <p>Further details are provided in</p> <blockquote> <p>Chen et al. (2022). SemEval-2022 Task 8: Multilingual news article similarity. In Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022).&nbsp;<a href="https://aclanthology.org/2022.semeval-1.155/">https://aclanthology.org/2022.semeval-1.155/</a></p> </blockquote> <p>The data in this repository includes pairs of URLs and annotations. The text of webpages is generally&nbsp;via the Internet Archive in this special collection: https://archive.org/details/2020-multilingual-news-article-similarity . A script to download and process the webpages is available at&nbsp;https://github.com/euagendas/semeval_8_2022_ia_downloader .&nbsp;</p>

opencc-by-4.0Apr 2022View details →
zenodo44/100

Swedish Test Data for SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>This data collection contains the Swedish test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection:</a></p> <p>- a Swedish text corpus pair (`corpus1/`, `corpus2/`)<br> - 31 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)<br> - the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</p> <p>We sample from the KubHist2 corpus, digitized by the National Library of Sweden, and available through the Spr&aring;kbanken corpus infrastructure Korp (<a href="https://www.researchgate.net/profile/Markus_Forsberg/publication/266352576_Korp_-_the_corpus_infrastructure_of_Sprakbanken/links/55bf1ee008aed621de121ba3/Korp-the-corpus-infrastructure-of-Sprakbanken.pdf">Borin et al., 2012</a>). The full corpus is available through a CC BY (attribution) license. Each word for which the lemmatizer in the Korp pipelien has found a lemma is replaced with the lemma. In cases where the lemmatizer cannot find a lemma, we leave the word as is (i.e., unlemmatized, no lower-casing). KubHist contains very frequent OCR errors, especially for the older data.More detail about the properties and quality of the Kubhist corpus can be found in (<a href="https://www.diva-portal.org/smash/get/diva2:1358014/FULLTEXT01.pdf#page=28">Adesam et al., 2019</a>).</p> <p>Lars Borin, Markus Forsberg, and Johan Roxendal. &quot;Korp-the corpus infrastructure of Spr&aring;kbanken.&quot; <em>LREC</em>. 2012.</p> <p>Adesam, Yvonne, Dana Dann&eacute;lls, and Nina Tahmasebi. &quot;Exploring the Quality of the Digital Historical Newspaper Archive KubHist.&quot; <em>DHN</em>. 2019.</p> <p>__Corpus 1__</p> <p>- based on: <a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1790-1830<br> - size: ~71 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains frequent OCR errors</p> <p>__Corpus 2__</p> <p>- based on:&nbsp;<a href="https://spraakbanken.gu.se/korp/?mode=kubhist">Kubhist2</a><br> - language: Swedish<br> - time covered: 1895-1903<br> - size: ~111 million tokens<br> - format: lemmatized, sentence length &gt; 9 (before removal of punctuation), no punctuation, sentences randomly shuffled<br> - encoding: UTF-8<br> - note: contains OCR errors</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (`corpus1/token/`, `corpus2/token/`). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>&nbsp;</p> <p>Reference:</p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi.<a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p>

opencc-by-2.0Feb 2020View details →
zenodo44/100

SemEval-2021 Task 12: Learning with Disagreements

<p>This repository contains the Post-Evaluation data for SemEval-2021 Task 12: Learning with Disagreement, a shared task on learning to classify with datasets containing disagreements.&nbsp;</p> <p>The aim of this shared task is to provide a unified testing framework for learning from disagreements using the best-known datasets containing information about disagreements for interpreting language and classifying images:</p> <p>&nbsp;&nbsp; &nbsp;1. LabelMe-IC: Image Classification using a subset of LabelMe images (Russell et al., 2008), is a widely used, community-created image classification dataset where images are assigned to one of 8 categories: highway, inside city, tall building, street, forest, coast, mountain, open country. Rodrigues and Pereira (2017) collected crowd labels for these images using Amazon Mechanical Turk (AMT).</p> <p>&nbsp;&nbsp; &nbsp;2. CIFAR10-IC: Image Classification using a subset of CIFAR-10 dataset, https://www.cs.toronto.edu/~kriz/cifar.html. The entire dataset consists of colour images in 10 categories (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck). Crowdsourced labels for this dataset were collected by Peterson et al (2019).</p> <p>&nbsp;&nbsp; &nbsp;3. PDIS: Information Status Classification using Phrase Detectives Information. Information Status Classification (IS) in Phrase Detectives (Poesio et al., 2019) dataset involves identifying the information status of a noun phrase: whether that noun phrase refers to new information or to old information.</p> <p>&nbsp;&nbsp; &nbsp;4. Gimpel-POS: Part-of-Speech tagging using the Gimpel dataset (Gimpel et al., 2011) for Twitter posts. Plank et al.(2014b) mapped the Gimpel tags to the universal tag set (Petrov et al., 2011), using these tags as gold, and collected crowdsourced labels.</p> <p>&nbsp;&nbsp; &nbsp;5. Humour: ranking one-line texts using pairwise funniness judgements (Simpson et al., 2019). Crowdworkers have annotated pairs of puns to indicate which is funniest. A gold standard ranking was produced using a large number of redundant annotations. The goal is to infer the gold standard ranking from a reduced number of crowdsourced judgements.</p> <p><br> The files contained in this data collection are as follows:<br> starting_kit.zip - Base models used provided for the shared task.&nbsp;<br> practice_phase_data.zip - The training and development data used during the Practice Phase of the competition.&nbsp;<br> test_phase_data.zip - The test data, used during the Evaluation Phase of the competition</p> <p>Details of format of each dataset for each task can be found on&nbsp;<a href="http://This data collection contains the post-evaluation data for SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection: This repository contains the Post-Evaluation data for SemEval-2021 Task 12: Learning with Disagreement, a shared task on learning to classify with datasets containing disagreements. The aim of this shared task is to provide a unified testing framework for learning from disagreements using the best-known datasets containing information about disagreements for interpreting language and classifying images: 1. LabelMe-IC: Image Classification using a subset of LabelMe images (Russell et al., 2008), is a widely used, community-created image classification dataset where images are assigned to one of 8 categories: highway, inside city, tall building, street, forest, coast, mountain, open country. Rodrigues and Pereira (2017) collected crowd labels for these images using Amazon Mechanical Turk (AMT). 2. CIFAR10-IC: Image Classification using a subset of CIFAR-10 dataset, https://www.cs.toronto.edu/~kriz/cifar.html. The entire dataset consists of colour images in 10 categories (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck). Crowdsourced labels for this dataset were collected by Peterson et al (2019). 3. PDIS: Information Status Classification using Phrase Detectives Information. Information Status Classification (IS) in Phrase Detectives (Poesio et al., 2019) dataset involves identifying the information status of a noun phrase: whether that noun phrase refers to new information or to old information. 4. Gimpel-POS: Part-of-Speech tagging using the Gimpel dataset (Gimpel et al., 2011) for Twitter posts. Plank et al.(2014b) mapped the Gimpel tags to the universal tag set (Petrov et al., 2011), using these tags as gold, and collected crowdsourced labels. 5. Humour: ranking one-line texts using pairwise funniness judgements (Simpson et al., 2019). Crowdworkers have annotated pairs of puns to indicate which is funniest. A gold standard ranking was produced using a large number of redundant annotations. The goal is to infer the gold standard ranking from a reduced number of crowdsourced judgements. The files contained in this data collection are as follows: starting_kit.zip - Base models used provided for the shared task. practice_phase_data.zip - The training and development data used during the Practice Phase of the competition. test_phase_data.zip - The test data, used during the Evaluation Phase of the competiton Details of format of each dataset for each task can be found here: https://competitions.codalab.org/competitions/25748#participate-get_data">Codalab</a>.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

Data for PAN at SemEval 2019 Task 4: Hyperpartisan News Detection

<p>Training, validation, and test data for the <a href="https://webis.de/events/semeval-19/">PAN @ SemEval 2019 Task 4: Hyperpartisan News Detection</a>.</p> <p>See the README for details.</p>

opencc-by-4.0Nov 2018View details →
zenodo40/100

SemEval-2024 Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes

<p><strong>Task description:</strong>&nbsp;SHROOM participants will need to detect grammatically sound output that contains incorrect semantic information (i.e. unsupported or inconsistent with the source input), with or without having access to the model that produced the output.</p> <p><strong>Overview of the task:</strong>&nbsp;The modern NLG landscape is plagued by two interlinked problems: On the one hand, our current neural models have a propensity to produce inaccurate but fluent outputs; on the other hand, our metrics are most apt at describing fluency, rather than correctness. This leads neural networks to &ldquo;hallucinate&rdquo;, i.e., produce fluent but incorrect outputs that we currently struggle to detect automatically. For many NLG applications, the correctness of an output is however mission critical. For instance, producing a plausible-sounding translation that is inconsistent with the source text puts in jeopardy the usefulness of a machine translation pipeline. With our shared task, we hope to foster the growing interest in this topic in the community.</p> <p>With SHROOM we adopt a post hoc setting, where models have already been trained and outputs already produced: participants will be asked to perform binary classification to identify cases of fluent overgeneration hallucinations in two different setups: model-aware and model-agnostic tracks. That is, participants must detect grammatically sound outputs which contain incorrect or unsupported semantic information, inconsistent with the source input, with or without having access to the model that produced the output. To that end, we will provide participants with a collection of checkpoints, inputs, references and outputs of systems covering three different NLG tasks: definition modeling (DM), machine translation (MT) and paraphrase generation (PG), trained with varying degrees of accuracy. The development set will provide binary annotations from at least five different annotators and a majority vote gold label.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

SemEval-2021 Task 1: Lexical Complexity Prediction

<p>We ran a shared task on Lexical Complexity Prediction via SemEval using this data. The data was released according to the following schedule:</p> <ul> <li> <p>Trial data available: July 31, 2020</p> </li> <li> <p>Training data available: September 4, 2020</p> </li> <li> <p>Test data available/Evaluation starts: January 11, 2021</p> </li> </ul> <p>See the task website for further information: <a href="https://sites.google.com/view/lcpsharedtask2021">https://sites.google.com/view/lcpsharedtask2021</a></p> <p>See the CodaLab site to view submissions: <a href="https://competitions.codalab.org/competitions/27420">https://competitions.codalab.org/competitions/27420</a></p> <p>If you find this dataset useful, please cite the task paper as:&nbsp; <em>Shardlow, M. Evans, R., Paetzold, G and Zampieri, M. SemEval-2021 Task 1: Lexical Complexity Prediction. In Proceedings of the 15th International Workshop on Semantic Evaluation. 2021.</em></p> <p>The script evaluate.py is the same evaluation script that was used as part of the shared task. You can run this by calling &quot;python evaluate.py &lt;results_path&gt;/res/ &lt;reference_path&gt;/ref/&quot; where res and ref are directories containing the system&#39;s results and the reference labels respectively.</p> <p>The trial, training and test data are available. The trial data comprises of 99 MWEs (29 bible, 33 biomed and 37 europarl) and 421 single word instances (143 bible, 135 biomed and 143 europarl).</p> <p>The training data comprises of 1,517 MWEs (505 bible, 514 biomed and 498 europarl) and 7,662 single word instances (2,574 bible, 2,576 biomed and 2,512 europarl).</p> <p>The test data comprises of 184 MWEs (66 bible, 53 biomed and 65 europarl) and 917 single word instances (283 bible, 289 biomed and 345 europarl). It is compressedd and encrypted using 7zip as above. The password will be released to those registered for the task. The order of the test data has been randomised to prevent systems from overfitting to the data ordering. All submissions should be made via CodaLab as above.</p> <p>The trial and training data is arranged by token and sorted by complexity (i.e., instances with the same token appear together in groups, and the groups are sorted by the complexity of the lowest scored item). Tokens that appear in one partition will not appear in another partition. We have deliberately included more than one instance of a token where possible to identify places where the context affects the complexity of a word. Consider, for example, the following two sentences from the trial data:</p> <ul> <li>We now have a proposal on the <em>table</em> which lays down strict emissions values for the next ten years and which simultaneously creates clarity and incentives for technological innovations.</li> <li>Mr President, in coordination with other groups, I would like to <em>table</em> an oral amendment concerning the draft bill in the Duma to ignore certain rulings of the European Court of Human Rights.</li> </ul> <p>When <em>table</em> is used as a noun in the first sentence (albeit in an abstract sense) it receives a score of 0.01 indicating it is very easily understood. However, when it is used as a verb in the second, it is scored at 0.23, indicating a more difficult word to comprehend.</p> <p>The data comes from three sources: biblical text, biomedical articles and proceedings of the European Parliament. These sources were selected as they contain a natural mixture of common language and difficult to understand expressions, whilst each containing vastly different domain-specific vocabulary.</p>

opencc-by-4.0Jul 2021View details →
zenodo40/100

SemEval-2021 Task 10: Source-Free Domain Adaptation for Semantic Processing

<p>Data sharing restrictions are common in NLP datasets. For example, Twitter policies do not allow sharing of tweet text, though tweet IDs may be shared. The situation is even more common in clinical NLP, where patient health information must be protected, and annotations over health text, when released at all, often require the signing of complex data use agreements. The SemEval-2021 Task 10 framework asks participants to develop semantic annotation systems in the face of data sharing constraints. A participant&#39;s goal is to develop an accurate system for a target domain when annotations exist for a related domain but cannot be distributed. Instead of annotated training data, participants are given a model trained on the annotations. Then, given unlabeled target domain data, they are asked to make predictions.</p> <p>Website: <a href="https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/">https://machine-learning-for-medical-language.github.io/source-free-domain-adaptation/</a></p> <p>CodaLab site: <a href="https://competitions.codalab.org/competitions/26152">https://competitions.codalab.org/competitions/26152</a></p> <p>Github repository: <a href="https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation">https://github.com/Machine-Learning-for-Medical-Language/source-free-domain-adaptation</a></p>

opencc-by-4.0Jul 2021View details →
zenodo36/100

SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p><strong>Authors</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi</p> <p><strong>Description</strong></p> <p>This data collection contains the <strong>post-evaluation</strong> data for <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:</p> <ul> <li>the starting kit to download data, and examples for competing in the CodaLab challenge including baselines</li> <li>the true binary change scores of the targets for Subtask 1, and their true graded change scores for Subtask 2 (<code>test_data_truth/</code>),</li> <li>the scoring program used to score submissions against the true test data in the evaluation and post-evaluation phase (<code>scoring_program/</code>),</li> <li>the results of the evaluation phase including <ul> <li>the final rankings of the participating teams by their best submission (<code>results/rankings_teams.csv</code>),</li> <li>the submitted files of each team (<code>results/submissions/</code>),</li> <li>an overview of the results for each submission ordered by team (<code>results/submissions_results.csv</code>),</li> <li>analysis plots (<code>plots/</code>) displaying the results: <ul> <li>under <code>per_target/</code> we provide the gold change scores and the normalized prediction error of target words plotted against their frequency and polysemy statistics,</li> <li>under <code>per_team/</code> we provide the model predictions from the best submission per team (per subtask) plotted against frequency/polysemy statistics and performance on gold data (gray lines give the correlation with the respective variable in the gold data); we also provide plots of visualizing the teams&#39; prediction similarities.</li> </ul> </li> </ul> </li> </ul> <p>Some remarks:</p> <ul> <li>the paper referenced below remains the only source for the rankings between teams,</li> <li>some teams were disqualified, and are thus removed from the analyses and the rankings present in the paper,</li> <li>some teams have changed names, resulting in a discrepancy between team names under <code>results/</code> and team names in the paper. The paper contains a key to match old names with new names.</li> </ul> <p><strong>Test Data </strong>for SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection can be found using the links below:</p> <ul> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-eng/">English</a></li> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-ger/">German</a></li> <li><a href="https://zenodo.org/record/3734089">Latin</a></li> <li><a href="https://zenodo.org/record/3730550">Swedish</a></li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi. 2020. <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. SemEval@COLING2020.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <pre><code>@inproceedings{schlechtweg2020semeval, title = "{S}em{E}val-2020 {T}ask 1: {U}nsupervised {L}exical {S}emantic {C}hange {D}etection", author = "Schlechtweg, Dominik and McGillivray, Barbara and Hengchen, Simon and Dubossarsky, Haim and Tahmasebi, Nina", booktitle = "To appear in Proceedings of the 14th International Workshop on Semantic Evaluation", year = "2020", address = "Barcelona, Spain", publisher = "Association for Computational Linguistics"}</code></pre> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo36/100

SemEval-2020 Task 5: Modelling Causal Reasoning in Language: Detecting Counterfactuals

<p><strong>SemEval-2020 Task 5</strong></p> <p>&nbsp;</p> <p><strong>Subtask-1:</strong> Recognizing Counterfactual Statements (RCS) -- Determine whether a given sentence is counterfactual or not.</p> <p><strong>Subtask-2: </strong>Detecting Antecedent and Consequent (DAC) -- Extract the antecedent and consequent part in a given counterfactual sentence.</p> <p>&nbsp;</p> <p>The released dataset consists of train/test data of both subtask-1 and subtask-2. In our competition, participants could only use the corresponding dataset in each subtask.</p> <p>&nbsp;</p> <p><strong>Task 5 Codalab Website:</strong> <a href="https://competitions.codalab.org/competitions/21691">https://competitions.codalab.org/competitions/21691</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)

<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab:&nbsp;<a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is&nbsp;<a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles

<p>This dataset contains the files and annotations for <a href="https://propaganda.qcri.org/semeval2020-task11/index.html">SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles</a>. The task was composed by two subtasks: span identification (SI) and technique classification (TC). This dataset includes the following:</p> <ul> <li>The text files for training, development, and testing sets for both the SI and the TC tasks.</li> <li>The gold-standard files for the training sets, for both the SI and the TC task</li> </ul> <p>Our propaganda identification initiative remains active. We keep a <a href="https://propaganda.qcri.org/ptc/leaderboard.php">live leader-board</a> reporting the performance of models submited up to date.</p> <p><strong>Reference</strong></p> <p>Giovanni Da San Martino, Alberto Barr&oacute;n-Cede&ntilde;o, Henning Wachsmuth, Rostislav Petrov, and Preslav Nakov. 2020. <a href="https://propaganda.qcri.org/">Task 11: Detection of Propaganda Techniques in News Articles</a>. In Proceedings of the 14th International Workshop on Semantic Evaluation (SemEval 2020). Barcelona, Spain (2020)</p> <p>&nbsp;</p> <pre><code>@InProceedings{SemEval20-11-DaSanMartino, author = "Da San Martino, Giovanni and Barr\'{o}n-Cede\~no, Alberto and Wachsmuth, Henning and Petrov, Rostislav and Nakov, Preslav", title = "{SemEval}-2020 Task 11: {D}etection of Propaganda Techniques in News Articles", pages = "", abstract = "We describe the outcome of the SemEval 2020 Task 11 on the detection of propaganda in news articles. We present two tasks. In the first task, systems are asked to identify specific text spans in a free text where propaganda is being applied. In the second task, systems are asked to identify the propaganda technique being applied in a text span. We describe the construction of the evaluation framework (dataset and evaluation metrics) as well as the approaches explored by the different participants. ", crossref = "SemEval20" }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

LatinISE test data for SemEval 2020 task 1 with additional token versions of the corpora

<p>This data collection contains the Latin test data for <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:&nbsp;</p> <ul> <li>a Latin text corpus pair (`corpus1/lemma`, `corpus2/lemma`)</li> <li>40 lemmas which have been annotated for their lexical semantic change between the two corpora (`targets.txt`)</li> <li>the annotated binary change scores of the targets for subtask 1, and their annotated graded change scores for subtask 2 (`truth/`)</li> </ul> <p>The corpus data have been automatically lemmatized and part-of-speech tagged, and have been partially corrected by hand. For homonyms, the lemmas are followed by the &#39;\#&#39; symbol and the number of the homonym according to the Lewis-Short dictionary of Latin when this number is greater than 1. For example, the lemma &#39;dico&#39; corresponds to the first homonym in the Lewis-Short dictionary and &#39;dico\#2&#39; corresponds to the second homonym, cf. Lewis-Short dictionary.</p> <p>__Corpus 1__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a>&nbsp;(McGillivray and Kilgarriff 2013), <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the second century before Christ (BC) to the end of the first century BC</li> <li>size: ~1.7 million tokens</li> <li>format: lemmatized, sentence length &gt;= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>__Corpus 2__</p> <ul> <li>based on: <a href="http://hdl.handle.net/11372/LRT-3170">LatinISE</a>&nbsp;(McGillivray and Kilgarriff 2013) , <a href="https://app.sketchengine.eu/#dashboard?corpname=preloaded/latinise_4">version on Sketch Engine</a></li> <li>language: Latin</li> <li>time covered: from the beginning of the first century after Christ (AD) to the end of the twenty-first century AD</li> <li>size: ~9.4 million tokens</li> <li>format: lemmatized, sentence length &gt;= 2, no punctuation, sentences randomly shuffled</li> <li>encoding: UTF-8</li> </ul> <p>Find more information on the data in the papers referenced below.</p> <p>Besides the official lemma version of the corpora for SemEval-2020 Task 1 we also provide the raw token version (<code>corpus1/token/</code>,&nbsp;<code>corpus2/token/</code>). It contains the raw sentences in the same order as in the lemma version. Find more information on the data and SemEval-2020 Task 1 in the paper referenced below.</p> <p>The creation of the data was supported by the CRETA center and the CLARIN-D grant funded by the German Ministry for Education and Research (BMBF).</p> <p><strong>References</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi <a href="https://competitions.codalab.org/competitions/20948">SemEval 2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. To appear in SemEval@COLING2020.</p> <p>McGillivray, B. and Kilgarriff, A. (2013). <a href="https://www.sketchengine.co.uk/wp-content/uploads/2015/05/Latin_historical_corpus_2013.pdf">Tools for historical corpus research, and a corpus of Latin</a>. In Paul Bennett, Martin Durrell, Silke Scheible, Richard J. Whitt (eds.), New Methods in Historical Corpus Linguistics, T&uuml;bingen: Narr.<br> &nbsp;</p>

opencc-by-4.0Aug 2020View details →
zenodo36/100

SemEval 2017 Task 3 Subtask E Train data (StackExchange)

<p>This is the train data that was used for Task 3, Subtask E of the SemEval-2017 Shared Task. More information on the task can be found here: http://alt.qcri.org/semeval2017/task3/</p>

opencc-by-4.0Sep 2016View details →
zenodo36/100

SemEval 2017 Task 3 Subtask E Test data (StackExchange)

<p>This is the test data that was used for Task 3, Subtask E of the SemEval-2017 Shared Task. More information on the task can be found here: http://alt.qcri.org/semeval2017/task3/</p>

opencc-by-4.0Jan 2017View details →
zenodo36/100

SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense

<h3>Data&nbsp; for the SemEval 2024 paper "SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense"</h3> <p>&nbsp;</p> <p>The data of the two subtasks is saved in the&nbsp;<strong>data</strong>&nbsp;folder,&nbsp;<em>BTDATA.zip</em>, which contains the data for the sentence puzzle and word puzzle.</p> <p>The data contained in <em>BTDATA.zip</em>&nbsp;are as follows:</p> <ul> <li>Semeval Competition <ul> <li>Training Data <ul> <li><code>SP_train.npy</code>&nbsp;(Semeval training data)</li> <li><code>WP_train.npy</code>&nbsp;(Semeval training data)</li> </ul> </li> <li>Test Data <ul> <li><code>SP_test.npy</code>&nbsp;(Semeval test data)</li> <li><code>WP_test.npy</code>&nbsp;(Semeval test data)</li> <li><code>SP_test_answer.npy</code>&nbsp;(Semeval test data answer)</li> <li><code>WP_test_answer.npy</code>&nbsp;(Semeval test data answer)</li> </ul> </li> </ul> </li> </ul> <div> <h3><strong>Relation to EMNLP 2023 Paper</strong></h3> <a href="https://github.com/1171-jpg/BrainTeaser#2-relation-to-semeval2024-task9"></a></div> <p>The relationship between EMNLP and SemEval involves using the same dataset but with different data splitting and utilization methods. In EMNLP, the entire dataset is employed for testing, while in SemEval, the dataset is divided into training and testing sets, with the training set comprising a significant majority.</p> <p>Our EMNLP paper results on GitHub are tested on the entire data in a&nbsp;<strong>zero-shot manner</strong>. In the SemEval2024-Task9, although the whole dataset is the same as our EMNLP paper, we allow people to&nbsp;<strong>train on 80% of the whole dataset</strong>, and we&nbsp;<strong>evaluate the system on the 20% left</strong>.</p> <ul> <li>EMNLP Zero-Shot Experiment <ul> <li><code>sentence_puzzle.npy</code>&nbsp;(on all sentence puzzle data)</li> <li><code>word_puzzle.npy</code>&nbsp;(on all word puzzle data)</li> </ul> </li> </ul> <p><strong>Note:</strong>&nbsp;To prevent automatic data crawlers,&nbsp;<em>BTDATA.zip</em>&nbsp;needs a password:&nbsp;<strong>brainteaser</strong></p>

opencc-by-4.0Apr 2024View details →
zenodo36/100

Luminoso Input Data for SemEval-2018 Task 10: "Capturing Discriminative Attributes"

<p>This is the data required to run Luminoso&#39;s entry to the SemEval-2018 task on Capturing Discriminative Attributes.</p> <p>This data includes:</p> <ul> <li>A recently-computed version of the <a href="https://github.com/commonsense/conceptnet-numberbatch">ConceptNet Numberbatch</a> word embeddings</li> <li>The&nbsp;output of an implementation of Semantic Matching Energy over ConceptNet</li> <li>A SQLite database containing the lead section of all articles on the <a href="http://en.wikipedia.org">English Wikipedia</a> on 2017-12-20</li> <li>The text file that that database is constructed from</li> <li>A SQLite database of words that co-occur in <a href="http://storage.googleapis.com/books/ngrams/books/datasetsv2.html">Google Books 2-grams</a></li> <li>The text file containing total counts of 2-grams in the Google Books data, which that database is constructed from</li> </ul> <p>For more information, see the paper &quot;Luminoso at SemEval-2018 Task 10: Distinguishing Attributes Using Text Corpora and Relational Knowledge&quot;, by Robyn&nbsp;Speer and Joanna Lowry-Duda, to appear in the proceedings of the SemEval workshop at&nbsp;NAACL 2018.</p>

opencc-by-sa-4.0Feb 2018View details →
zenodo36/100

Sampled sentence pairs from SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>Each dataset consists of samples, containing two sentences with positions of one of the given target words. In every sample, first sentence is taken from corpus1 and second from corpus2. Initial sentences were taken from&nbsp;https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd/.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo36/100

Data produced in the context of RCLN particpation to the Visual WSD task at SemEval 2023

<p>The data contain:</p> <p>- generated captions from train, trial and test images</p> <p>- generated images from the diffusion model</p> <p>refer to&nbsp;https://github.com/dbuscaldi/VisualWSD23 for code</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

SemEval-2020 Task 7: Assessing Humor in Edited News Headlines

<p>This is the task dataset for SemEval-2020 Task 7: Assessing Humor in Edited News Headlines.</p> <p>The task&rsquo;s dataset contains news headlines in which short edits were applied to make them funny, and the funniness of these edited headlines was rated using crowdsourcing. This task includes two subtasks, the first of which is to estimate the funniness of headlines on a humor scale in the interval 0-3. The second subtask is to predict, for a pair of edited versions of the same original headline, which is the funnier version.</p> <p>CodaLab page hosting the competition:<br> <a href="https://competitions.codalab.org/competitions/20970">https://competitions.codalab.org/competitions/20970</a></p> <p>Starter Github code (scripts for running baseline and evaluation):<br> <a href="https://github.com/n-hossain/semeval-2020-task-7-humicroedit">https://github.com/n-hossain/semeval-2020-task-7-humicroedit</a></p> <p>Task mailing list:<br> <a href="https://groups.google.com/forum/#!forum/semeval-2020-task-7-all">https://groups.google.com/forum/#!forum/semeval-2020-task-7-all</a><br> ----------------------------------------------------------------------</p> <p>ZIP contents:<br> -------------</p> <p>Folders:<br> &nbsp;&nbsp; &nbsp;- subtask-1: Dataset for the funniness regression subtask.<br> &nbsp;&nbsp; &nbsp;- subtask-2: Dataset for the &quot;Funnier of the Two&quot; classification subtask.</p> <p>Files:<br> &nbsp;&nbsp; &nbsp;- {train, dev, test}.csv: the task&#39;s dataset including labels<br> &nbsp;&nbsp; &nbsp;- train_funlines.csv: additional training data gathered from the FunLines competition (https://funlines.co)<br> &nbsp;&nbsp; &nbsp;- baseline.zip: contains csv file which is the output of the BASELINE system. This is a template of the output format that can be submitted to CodaLab for scoring.</p> <p><strong>Reference</strong></p> <p>Please cite the task paper when using this dataset:</p> <p>Nabil Hossain, John Krumm, Michael Gamon and Henry Kautz. 2020. Semeval-2020 Task 7: Assessing Humor in Edited News Headlines. In Proceedings of International Workshop on Semantic Evaluation (SemEval-2020).</p> <pre><code>BIBTEX:  @InProceedings{hossainSemEval2020Task7, author = {Hossain, Nabil and Krumm, John and Gamon, Michael and Kautz,Henry}, title = {SemEval-2020 {T}ask 7: {A}ssessing Humor in Edited News Headlines}, booktitle = {Proceedings of the 14th International Workshop on Semantic Evaluation ({S}em{E}val-2020)}, address = {Barcelona, Spain}, year = {2020}}</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets

<p>There are 2 sub-tasks: sentiment analysis for Spanglish (Spanish-English) and for Hinglish (Hindi-English).</p> <p>The sentiment classes are Positive, negative, neutral.&nbsp;</p> <p>Hinglish dataset has 20k instances.</p> <p>Spanglish dataset has ~19k instances.&nbsp;</p> <p>Website:&nbsp;<a href="https://ritual-uh.github.io/sentimix2020/">https://ritual-uh.github.io/sentimix2020/</a></p>

opencc-by-4.0Aug 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record