Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

9

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

9 results for “SemEval-2020”

Learn how ShareScore rates datasets ↗
zenodo36/100

SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p><strong>Authors</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky, and Nina Tahmasebi</p> <p><strong>Description</strong></p> <p>This data collection contains the <strong>post-evaluation</strong> data for <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>:</p> <ul> <li>the starting kit to download data, and examples for competing in the CodaLab challenge including baselines</li> <li>the true binary change scores of the targets for Subtask 1, and their true graded change scores for Subtask 2 (<code>test_data_truth/</code>),</li> <li>the scoring program used to score submissions against the true test data in the evaluation and post-evaluation phase (<code>scoring_program/</code>),</li> <li>the results of the evaluation phase including <ul> <li>the final rankings of the participating teams by their best submission (<code>results/rankings_teams.csv</code>),</li> <li>the submitted files of each team (<code>results/submissions/</code>),</li> <li>an overview of the results for each submission ordered by team (<code>results/submissions_results.csv</code>),</li> <li>analysis plots (<code>plots/</code>) displaying the results: <ul> <li>under <code>per_target/</code> we provide the gold change scores and the normalized prediction error of target words plotted against their frequency and polysemy statistics,</li> <li>under <code>per_team/</code> we provide the model predictions from the best submission per team (per subtask) plotted against frequency/polysemy statistics and performance on gold data (gray lines give the correlation with the respective variable in the gold data); we also provide plots of visualizing the teams&#39; prediction similarities.</li> </ul> </li> </ul> </li> </ul> <p>Some remarks:</p> <ul> <li>the paper referenced below remains the only source for the rankings between teams,</li> <li>some teams were disqualified, and are thus removed from the analyses and the rankings present in the paper,</li> <li>some teams have changed names, resulting in a discrepancy between team names under <code>results/</code> and team names in the paper. The paper contains a key to match old names with new names.</li> </ul> <p><strong>Test Data </strong>for SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection can be found using the links below:</p> <ul> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-eng/">English</a></li> <li><a href="https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd-ger/">German</a></li> <li><a href="https://zenodo.org/record/3734089">Latin</a></li> <li><a href="https://zenodo.org/record/3730550">Swedish</a></li> </ul> <p>Please find more information on the provided data in the paper referenced below.</p> <p><strong>Reference</strong></p> <p>Dominik Schlechtweg, Barbara McGillivray, Simon Hengchen, Haim Dubossarsky and Nina Tahmasebi. 2020. <a href="https://languagechange.org/semeval">SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</a>. SemEval@COLING2020.</p> <p>The resources are freely available for education, research and other non-commercial purposes.</p> <pre><code>@inproceedings{schlechtweg2020semeval, title = "{S}em{E}val-2020 {T}ask 1: {U}nsupervised {L}exical {S}emantic {C}hange {D}etection", author = "Schlechtweg, Dominik and McGillivray, Barbara and Hengchen, Simon and Dubossarsky, Haim and Tahmasebi, Nina", booktitle = "To appear in Proceedings of the 14th International Workshop on Semantic Evaluation", year = "2020", address = "Barcelona, Spain", publisher = "Association for Computational Linguistics"}</code></pre> <p>&nbsp;</p>

opencc-by-4.0May 2020View details →
zenodo36/100

SemEval-2020 Task 5: Modelling Causal Reasoning in Language: Detecting Counterfactuals

<p><strong>SemEval-2020 Task 5</strong></p> <p>&nbsp;</p> <p><strong>Subtask-1:</strong> Recognizing Counterfactual Statements (RCS) -- Determine whether a given sentence is counterfactual or not.</p> <p><strong>Subtask-2: </strong>Detecting Antecedent and Consequent (DAC) -- Extract the antecedent and consequent part in a given counterfactual sentence.</p> <p>&nbsp;</p> <p>The released dataset consists of train/test data of both subtask-1 and subtask-2. In our competition, participants could only use the corresponding dataset in each subtask.</p> <p>&nbsp;</p> <p><strong>Task 5 Codalab Website:</strong> <a href="https://competitions.codalab.org/competitions/21691">https://competitions.codalab.org/competitions/21691</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

SemEval-2020 Task 12: Multilingual Offensive Language Identification in Social Media (OffensEval 2020)

<p>The task involves three subtasks corresponding to the hierarchical taxonomy of the OLID schema (Zampieri et al., 2019) from OffensEval 2019. The task featured five languages and this upload is for the English language. In addition, English also featured Subtasks B and C. OffensEval 2020 was one of the most popular tasks at SemEval-2020 attracting a large number of participants across all subtasks and also across all languages. A total of 528 teams signed up to participate in the task, 145 teams submitted systems during the evaluation period, and 70 submitted system description papers.</p> <p>This upload includes a test set used in the paper describing the dataset used in the shared task as well as the official test set used in the shared task.</p> <p>The evaluation phase for English is available on Codalab:&nbsp;<a href="https://competitions.codalab.org/competitions/23285">https://competitions.codalab.org/competitions/23285</a></p> <p>The Website for the shared task is&nbsp;<a href="https://sites.google.com/site/offensevalsharedtask/home">https://sites.google.com/site/offensevalsharedtask/home</a></p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles

<p>This dataset contains the files and annotations for <a href="https://propaganda.qcri.org/semeval2020-task11/index.html">SemEval-2020 Task 11: Detection of Propaganda Techniques in News Articles</a>. The task was composed by two subtasks: span identification (SI) and technique classification (TC). This dataset includes the following:</p> <ul> <li>The text files for training, development, and testing sets for both the SI and the TC tasks.</li> <li>The gold-standard files for the training sets, for both the SI and the TC task</li> </ul> <p>Our propaganda identification initiative remains active. We keep a <a href="https://propaganda.qcri.org/ptc/leaderboard.php">live leader-board</a> reporting the performance of models submited up to date.</p> <p><strong>Reference</strong></p> <p>Giovanni Da San Martino, Alberto Barr&oacute;n-Cede&ntilde;o, Henning Wachsmuth, Rostislav Petrov, and Preslav Nakov. 2020. <a href="https://propaganda.qcri.org/">Task 11: Detection of Propaganda Techniques in News Articles</a>. In Proceedings of the 14th International Workshop on Semantic Evaluation (SemEval 2020). Barcelona, Spain (2020)</p> <p>&nbsp;</p> <pre><code>@InProceedings{SemEval20-11-DaSanMartino, author = "Da San Martino, Giovanni and Barr\'{o}n-Cede\~no, Alberto and Wachsmuth, Henning and Petrov, Rostislav and Nakov, Preslav", title = "{SemEval}-2020 Task 11: {D}etection of Propaganda Techniques in News Articles", pages = "", abstract = "We describe the outcome of the SemEval 2020 Task 11 on the detection of propaganda in news articles. We present two tasks. In the first task, systems are asked to identify specific text spans in a free text where propaganda is being applied. In the second task, systems are asked to identify the propaganda technique being applied in a text span. We describe the construction of the evaluation framework (dataset and evaluation metrics) as well as the approaches explored by the different participants. ", crossref = "SemEval20" }</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo36/100

Sampled sentence pairs from SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection

<p>Each dataset consists of samples, containing two sentences with positions of one of the given target words. In every sample, first sentence is taken from corpus1 and second from corpus2. Initial sentences were taken from&nbsp;https://www.ims.uni-stuttgart.de/en/research/resources/corpora/sem-eval-ulscd/.&nbsp;</p>

opencc-by-4.0Jun 2021View details →
zenodo32/100

SemEval-2020 Task 7: Assessing Humor in Edited News Headlines

<p>This is the task dataset for SemEval-2020 Task 7: Assessing Humor in Edited News Headlines.</p> <p>The task&rsquo;s dataset contains news headlines in which short edits were applied to make them funny, and the funniness of these edited headlines was rated using crowdsourcing. This task includes two subtasks, the first of which is to estimate the funniness of headlines on a humor scale in the interval 0-3. The second subtask is to predict, for a pair of edited versions of the same original headline, which is the funnier version.</p> <p>CodaLab page hosting the competition:<br> <a href="https://competitions.codalab.org/competitions/20970">https://competitions.codalab.org/competitions/20970</a></p> <p>Starter Github code (scripts for running baseline and evaluation):<br> <a href="https://github.com/n-hossain/semeval-2020-task-7-humicroedit">https://github.com/n-hossain/semeval-2020-task-7-humicroedit</a></p> <p>Task mailing list:<br> <a href="https://groups.google.com/forum/#!forum/semeval-2020-task-7-all">https://groups.google.com/forum/#!forum/semeval-2020-task-7-all</a><br> ----------------------------------------------------------------------</p> <p>ZIP contents:<br> -------------</p> <p>Folders:<br> &nbsp;&nbsp; &nbsp;- subtask-1: Dataset for the funniness regression subtask.<br> &nbsp;&nbsp; &nbsp;- subtask-2: Dataset for the &quot;Funnier of the Two&quot; classification subtask.</p> <p>Files:<br> &nbsp;&nbsp; &nbsp;- {train, dev, test}.csv: the task&#39;s dataset including labels<br> &nbsp;&nbsp; &nbsp;- train_funlines.csv: additional training data gathered from the FunLines competition (https://funlines.co)<br> &nbsp;&nbsp; &nbsp;- baseline.zip: contains csv file which is the output of the BASELINE system. This is a template of the output format that can be submitted to CodaLab for scoring.</p> <p><strong>Reference</strong></p> <p>Please cite the task paper when using this dataset:</p> <p>Nabil Hossain, John Krumm, Michael Gamon and Henry Kautz. 2020. Semeval-2020 Task 7: Assessing Humor in Edited News Headlines. In Proceedings of International Workshop on Semantic Evaluation (SemEval-2020).</p> <pre><code>BIBTEX:  @InProceedings{hossainSemEval2020Task7, author = {Hossain, Nabil and Krumm, John and Gamon, Michael and Kautz,Henry}, title = {SemEval-2020 {T}ask 7: {A}ssessing Humor in Edited News Headlines}, booktitle = {Proceedings of the 14th International Workshop on Semantic Evaluation ({S}em{E}val-2020)}, address = {Barcelona, Spain}, year = {2020}}</code></pre> <p>&nbsp;</p>

opencc-by-4.0Jul 2020View details →
zenodo32/100

SemEval-2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets

<p>There are 2 sub-tasks: sentiment analysis for Spanglish (Spanish-English) and for Hinglish (Hindi-English).</p> <p>The sentiment classes are Positive, negative, neutral.&nbsp;</p> <p>Hinglish dataset has 20k instances.</p> <p>Spanglish dataset has ~19k instances.&nbsp;</p> <p>Website:&nbsp;<a href="https://ritual-uh.github.io/sentimix2020/">https://ritual-uh.github.io/sentimix2020/</a></p>

opencc-by-4.0Aug 2020View details →
zenodo32/100

SemEval-2020 Task 3: Graded Word Similarity in Context

<p>For this tasks we ask participants to build systems that try to predict the effect that context has in human perception of similarity of words.</p> <p>We have seen very interesting work that uses local context to predict <strong>discrete</strong> changes in meaning: the different senses of a polysemous word. However context also has more subtle, <strong>continuous</strong> (<strong>graded</strong>) effects on meaning, even for words not necessarily considered polysemous.</p> <p>In order to be able to look at these effects we are building several datasets where we ask annotators to score how similar a pair of words are after they have read a short paragraph (which contains the two words). Each pair is scored within two of these paragraphs, allowing us to look at changes in similarity ratings due to context.</p> <p>CodaLab was used to run this task, you can see the dedicated website and the results of the participants at: https://competitions.codalab.org/competitions/20905</p>

opencc-by-4.0Jul 2020View details →
zenodo28/100

SemEval-2020 Task 8: Memotion Analysis- The Visuo-Lingual Metaphor!

<p>Memes typically induce humour and strive to be relatable. Many of them aim to express solidarity during certain life phases and thus, to connect with their audience. Some memes are directly humorous whereas others go for sarcastic dig at daily life events. Inspired by the various humorous effects of memes, we propose three tasks as follows:</p> <p>&bull;<strong>TaskA-Sentiment Classification:</strong> Given an Internet meme, the first task is to classify it as a positive or negative meme. We presume that a meme is not neutral.</p> <p>&bull;<strong>TaskB-Humor Classification: </strong>Given an Internet meme, the system has to identify the type of humour expressed. The categories are sarcastic, humorous, and offensive meme. If a meme does not fall under any of these categories, then it is marked as the other meme. A meme can have more than one category. For instance, Fig 3 is an offensive meme but sarcastic too.</p> <p><strong>&bull;TaskC-Scales of Semantic Classes:</strong> The third task is to quantify the extent to which a particular effect is being expressed. Details of such quantifications are reported in Table 1. Appropriateannotated data will be provided</p> <p>We have released 10K human-annotated Internet memes labelled with semantic dimensions namely sentiment, and type of humour that is, sarcastic, humorous, or offensive. The humour types are further quantified on a Like&nbsp;scale as in Table 1. The dataset will also contain the extracted captions/texts from the memes.</p>

opencc-by-4.0Jul 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record