Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
27
datasets available to search
ShareScore release 0.9.0
Dataset results
27 results for “semeval”
SemEval-2020 Task 3: Graded Word Similarity in Context
<p>For this tasks we ask participants to build systems that try to predict the effect that context has in human perception of similarity of words.</p> <p>We have seen very interesting work that uses local context to predict <strong>discrete</strong> changes in meaning: the different senses of a polysemous word. However context also has more subtle, <strong>continuous</strong> (<strong>graded</strong>) effects on meaning, even for words not necessarily considered polysemous.</p> <p>In order to be able to look at these effects we are building several datasets where we ask annotators to score how similar a pair of words are after they have read a short paragraph (which contains the two words). Each pair is scored within two of these paragraphs, allowing us to look at changes in similarity ratings due to context.</p> <p>CodaLab was used to run this task, you can see the dedicated website and the results of the participants at: https://competitions.codalab.org/competitions/20905</p>
SemEval-2013 Task 13: Word Sense Induction for Graded and Non-Graded Senses
<p>Data originally provided by the SemEval-2013 Task 13 organizers and which was available through the following link:</p> <p>https://www.cs.york.ac.uk/semeval-2013/task13/data/uploads/semeval-2013-task-13-test-data.zip</p>
SemEval-2010 Task 14: Word Sense Induction & Disambiguation
<p>Data originally provided by the SemEval-2010 Task 14 organizers and which was available through the following links:</p> <p>1. https://www.cs.york.ac.uk/semeval2010_WSI/files/evaluation.zip<br> 2. https://www.cs.york.ac.uk/semeval2010_WSI/files/training_data.tar.gz<br> 3. https://www.cs.york.ac.uk/semeval2010_WSI/files/test_data.tar.gz</p>
SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials
<p>This is the github repository hosting data and code for Task 2: Safe Biomedical Natural Language Inference for Clinical Trials at <a href="https://semeval.github.io/SemEval2024/" rel="nofollow">Semeval 2024</a>.</p> <p>For additional information about the task, please consult the official <a href="https://sites.google.com/view/nli4ct/home" rel="nofollow">website</a>.</p>
SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages
<p>This is the GitHub repository hosting data and code for SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages. For additional information about the task, please consult the official website.</p>
SemEval-2020 Task 8: Memotion Analysis- The Visuo-Lingual Metaphor!
<p>Memes typically induce humour and strive to be relatable. Many of them aim to express solidarity during certain life phases and thus, to connect with their audience. Some memes are directly humorous whereas others go for sarcastic dig at daily life events. Inspired by the various humorous effects of memes, we propose three tasks as follows:</p> <p>•<strong>TaskA-Sentiment Classification:</strong> Given an Internet meme, the first task is to classify it as a positive or negative meme. We presume that a meme is not neutral.</p> <p>•<strong>TaskB-Humor Classification: </strong>Given an Internet meme, the system has to identify the type of humour expressed. The categories are sarcastic, humorous, and offensive meme. If a meme does not fall under any of these categories, then it is marked as the other meme. A meme can have more than one category. For instance, Fig 3 is an offensive meme but sarcastic too.</p> <p><strong>•TaskC-Scales of Semantic Classes:</strong> The third task is to quantify the extent to which a particular effect is being expressed. Details of such quantifications are reported in Table 1. Appropriateannotated data will be provided</p> <p>We have released 10K human-annotated Internet memes labelled with semantic dimensions namely sentiment, and type of humour that is, sarcastic, humorous, or offensive. The humour types are further quantified on a Like scale as in Table 1. The dataset will also contain the extracted captions/texts from the memes.</p>
SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Black-Box Machine-Generated Text Detection
<p>Large language models (LLMs) are becoming mainstream and easily accessible, ushering in an explosion of machine-generated content over various channels, such as news, social media, question-answering forums, educational, and even academic contexts. Recent LLMs, such as ChatGPT and GPT-4, generate remarkably fluent responses to a wide variety of user queries. The articulate nature of such generated texts makes LLMs attractive for replacing human labor in many scenarios. However, this has also resulted in concerns regarding their potential misuse, such as spreading misinformation and causing disruptions in the education system. Since humans perform only slightly better than chance when classifying machine-generated vs. human-written text, there is a need to develop automatic systems to identify machine-generated text with the goal of mitigating its potential misuse.</p> <p>We offer three subtasks over two paradigms of text generation: (1) <strong>full text</strong> when a considered text is entirely written by a human or generated by a machine; and (2) <strong>mixed text</strong> when a machine-generated text is refined by a human or a human-written text paraphrased by a machine.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.