Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “shared task”

Learn how ShareScore rates datasets ↗
zenodo48/100

DIPROMATS 2024 - Shared Task 2: testing data for narrative identification

<p>Narratives are causally connected sequences of events that are selected and evaluated as meaningful for a particular audience. They make sense of the world by identifying the significance of people, places, objects, and events in time. In international relations, international actors create strategic narratives to &ldquo;construct a shared meaning of the past, present, and future of international politics to shape the behavior of domestic and international actors&rdquo;</p> <p>DIPROMATS 2024 Task 2 is a multiclass multilabel classification problem. Given a series of predefined narratives of each international actor, systems must determine which narrative the tweets belong to. Systems will receive the description of each narrative and a few examples of tweets in both languages (English and Spanish) that belong to each of them (few-shot learning). A tweet may be associated with one, several or none of the narratives.</p> <p>The few-shot training data can be found here: <a href="https://doi.org/10.5281/zenodo.10820961" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10820961</a></p> <p>These are the testing datasets for Englsih and Spanish. They are provided without the keys so the large language models can't be contaminated. If you are interested on testing your system, write anselmo@lsi.uned.es for details on submission and leaderboards.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Audio features for DSL-S 2023 shared task

<p>These files are features (i-vectors, x-vectors, MFCC features) extracted from the subset of <a href="https://commonvoice.mozilla.org/">Mozilla Common Voice</a> corpus version 12.0 used in the <a href="https://sites.google.com/view/vardial-2023">VarDial 2023</a> shared task on <a href="https://dsl-s.github.io/">Discriminating Between Similar Languages - Speech</a> (DSL-S 2023).</p> <p>This data set contains only the features for the training and development section (and the Common Voice meta data) for the nine languages included in the shared task (see shared task website for further information).</p> <p>&nbsp;</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

PharmaCoNER Silver Standard: Participant predictions in PharmaCoNER shared task

<p><strong>Introduction</strong></p> <p>Predictions in the test and background set of PharmaCoNER participants.</p> <p>&nbsp;</p> <p><strong>Zip structure</strong></p> <p>One directory per PharmaCoNER participant. Within each participant directory, there is one directory per run.</p> <p>&nbsp;</p> <p><strong>Corpus format description</strong></p> <p>For subtask 1 annotations are distributed in Brat format. See Brat webpage for more information <a href="https://brat.nlplab.org/standoff.html">https://brat.nlplab.org/standoff.html</a></p> <p>For subtask-2, codes are associated with each document are given in a TSV file with the following columns:&nbsp;</p> <pre>articleID code </pre> <p><br> <strong>Resources:</strong></p> <ul> <li><strong><a href="https://temu.bsc.es/pharmaconer/">Web</a></strong></li> <li><strong>Citation:&nbsp;</strong>A. G. Agirre, M. Marimon, A. Intxaurrondo, O. Rabal, M. Villegas, M. Krallinger, Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track, in: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks, 2019, pp. 1&ndash;10.</li> <li><a href="https://doi.org/10.5281/zenodo.4270157"><strong>Gold Standard corpus</strong></a></li> <li><a href="https://doi.org/10.5281/zenodo.3763276"><strong>Annotation guidelines</strong></a></li> <li><a href="https://github.com/TeMU-BSC/PharmaCoNER-Tagger"><strong>PharmaCoNER tagger</strong></a></li> </ul> <p>&nbsp;</p> <p>All credit&nbsp;to PharmaCoNER participants.&nbsp;</p> <p>For further information, please visit&nbsp;<a href="https://temu.bsc.es/pharmaconer/">https://temu.bsc.es/pharmaconer/</a>&nbsp;or email us at encargo-pln-life@bsc.es</p> <p>Copyright (c) 2018 Secretar&iacute;a de Estado para el Avance Digital (SEAD)</p>

opencc-by-4.0Nov 2020View details →
zenodo40/100

CLEF-HIPE-2020 Shared Task Named Entity Datasets

<p>CLEF-HIPE-2020 (Identifying Historical People, Places and other Entities) is a <strong>evaluation campaign on named entity processing</strong> on <strong>historical newspapers</strong> in <strong>French</strong>, <strong>German</strong> and <strong>English</strong>, which was organized in the context of the <a href="http://impresso-project.ch"><em>impresso</em></a> project and run as a <a href="https://clef2020.clef-initiative.eu/">CLEF 2020</a> Evaluation Lab. Data consists of manually annotated historical newspapers in French, German and English.<br> <br> For more information, please refer to:</p> <ul> <li>the CLEF-HIPE-2020 <a href="https://impresso.github.io/CLEF-HIPE-2020/">website</a>;</li> <li>the <a href="https://github.com/impresso/CLEF-HIPE-2020-eval">CLEF-HIPE-2020-eval repository</a>, for the necessary material to replicate the results of the shared task;</li> <li>the <a href="https://zenodo.org/record/3539085">CLEF-HIPE-2020 poster</a> presented at CLEF 2019 in Lugano, Switzerland;</li> <li>the CLEF-HIPE-2020 <a href="https://zenodo.org/record/3677171">participation guidelines</a> (v1.1);</li> <li>the <em>impresso</em> <a href="https://doi.org/10.5281/zenodo.3585749">Named Entity Annotation Guidelines</a> (v2.2.0);</li> <li>the <a href="https://infoscience.epfl.ch/record/281054">CLEF-HIPE-2020 Extended Overview</a> paper (bibtex below);</li> <li>the participant team <a href="http://ceur-ws.org/Vol-2696/">CEUR working note papers</a>;</li> <li>the workshop presentation <a href="https://www.youtube.com/playlist?list=PLB45F159nVx-3bee7G_1jdTfUAtsLD0FU">video records</a>;</li> </ul> <p>A second edition of HIPE is organised in 2022: <a href="https://hipe-eval.github.io/HIPE-2022/">https://hipe-eval.github.io/HIPE-2022/ </a></p> <p>Please cite this paper if you are using the datasets or find the shared task results relevant to your research:</p> <pre><code>@inproceedings{ehrmann_extended_2020, title = {Extended {Overview} of {CLEF HIPE} 2020: {Named Entity Processing} on {Historical Newspapers}}, booktitle = {{CLEF 2020 Working Notes}. {Working Notes} of {CLEF} 2020 - {Conference} and {Labs} of the {Evaluation Forum}}, author = {Ehrmann, Maud and Romanello, Matteo and Fl{\"u}ckiger, Alex and Clematide, Simon}, editor = {Cappellato, Linda and Eickhoff, Carsten and Ferro, Nicola and N{\'e}v{\'e}ol, Aur{\'e}lie}, year = {2020}, volume = {2696}, pages = {38}, publisher = {{CEUR-WS}}, address = {{Thessaloniki, Greece}}, doi = {10.5281/zenodo.4117566}, url = {https://infoscience.epfl.ch/record/281054}, }</code></pre> <p>&nbsp;</p>

opencc-by-nc-sa-4.0Mar 2020View details →
zenodo40/100

SIGTYP 2022 Shared Task: Prediction of Cognate Reflexes

<p>This is the data and code underlying the SIGTYP 2022 Shared Task on Cognate Reflex Prediction.</p>

opencc-by-4.0May 2022View details →
zenodo40/100

PAN Arabic Intrinsic Plagiarism Detection Shared Task Corpus

<p>Evaluation corpus for ARAbic INtrinsic plagiarism detection (InAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a>&nbsp;</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>InAra corpus comprises 2048 documents; 80% of them contain passages&nbsp;borrowed from other documents to simulate documents that contain&nbsp;plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 2 datasets:&nbsp;textual files and XML files.&nbsp;The textual files represent the suspicious documents i.e., the documents&nbsp;that contain artificial plagiarism; and the XML files are the plagiarism&nbsp;annotation i.e. they provide for each plagiarized passage its starting&nbsp;offset in the suspicious document and its length (offset and length are both expressed in characters). A suspicious document file and its plagiarism&nbsp;annotation file share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of InAra corpus is to evaluate automatic plagiarism&nbsp;detection methods, notably methods of the intrinsic approach. This&nbsp;approach consists in uncovering the plagiarized passages on the basis of&nbsp;the writing style inconsistency in a given suspicious document. As&nbsp;opposed to the external approach, the intrinsic approach does not&nbsp;necessitate any comparison of the suspicious document against the&nbsp;potential sources of plagiarism. Hence, InAra corpus is not appropriate for the evaluation of the external plagiarism detection because the source of plagiarism are not provided.</p> <p>It should be noted that some documents in InAra corpus contain religious&nbsp;quotations (e.g., Quran and Hadith). These quotations have a peculiar writing style and then a simple intrinsic plagiarism detection software can consider them as plagiarism. However, quotations are not plagiarism, and they are not&nbsp;annotated in the XML files in InAra. Hence, it is an important feature for the plagiarism detection systems evaluated on InAra to not consider religious quotations as plagiarism cases unless they appear as part of a larger&nbsp; plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose InAra corpus do not contain actual plagiarism&nbsp;cases. They are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the plagiarized passages lengths. This building method is the same&nbsp;used to construct PAN 2009-2011 corpora of plagiarism detection (see&nbsp;<a href="http://pan.webis.de">http://pan.webis.de</a> for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. SOURCES OF TEXTS&nbsp;</strong></p> <p>Texts used to build this corpus, either suspicious documents or the&nbsp;inserted passages, are taken mainly from the open library Arabic&nbsp;Wikisource (http://ar.wikisource.org), one of Wikimedia Foundation&nbsp;projects. A few numbers of documents were taken from other websites,&nbsp;namely:&nbsp;</p> <ul> <li>Create your own country blog: http://diycountry.blogspot.com&nbsp;</li> <li>Corpus of Classical Arabic (KSUCCA): http://ksucorpus.ksu.edu.sa&nbsp;</li> <li>Islamic book web site: http://www.islamicbook.ws&nbsp;</li> </ul> <p>&nbsp;</p> <p><strong>VII. COPYRIGHT AND AVAILABILITY&nbsp;</strong></p> <p>We were very careful to build the corpus with copyright-free texts only,&nbsp;to be able to make it publicly available without any sort of problems&nbsp;with texts owners.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using InAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on InAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>Additional information on the corpus building are in the papers:</p> <ul> <li>Bensalem, I., Rosso, P., Chikhi, S.: A New Corpus for the Evaluation&nbsp;of Arabic Intrinsic Plagiarism Detection. In: Forner, P., M&uuml;ller, H., Paredes, R., Rosso, P., and Stein, B. (eds.) CLEF 2013, LNCS, vol. 8138. pp. 53&ndash;58. Springer, Heidelberg (2013).</li> <li>Bensalem, I., Rosso, P., Chikhi, S.: Building Arabic Corpora from&nbsp;Wikisource. 10th ACS/IEEE International Conference on Computer Systems&nbsp;and Applications (AICCSA&rsquo;13),May 27-30 Fes/Ifran, Morocco (2013).IEEE.&nbsp;</li> </ul> <p>&nbsp;</p> <p>You may wish to&nbsp;compare the results of your experiments with the result of the following papers that used InAra corpus:</p> <ul> <li>Bensalem I, Rosso P, Chikhi S (2019) On the use of character n-grams&nbsp;as the only intrinsic evidence of plagiarism. Language Resources and&nbsp;Evaluation 53:363&ndash;396. doi: 10.1007/s10579-019-09444-w</li> <li>Mahgoub AY, Magooda A, Rashwan M, et al (2015) RDI System for&nbsp;Intrinsic Plagiarism Detection (RDI_RID), Working Notes for&nbsp;PAN-AraPlagDet at FIRE 2015. In: Majumder P, Mitra M, Agrawal M,&nbsp;&nbsp;Mehta P (eds) Post Proceedings of the Workshops at the 7th Forum for&nbsp;Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587. CEUR-WS.org, pp 129&ndash;130</li> </ul> <p>&nbsp;</p> <p><strong>IX. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection and&nbsp;you are retaining the corpus because it contains books you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;</p> <p>sources mentioned in Section VI where you can find the original content of&nbsp;the books you are looking for. We emphasize that we are not responsible&nbsp;for the results of any use of this corpus other than the evaluation of&nbsp;the intrinsic plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>X. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using InAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Jun 2015View details →
zenodo40/100

PAN Arabic External Plagiarism Detection Shared Task Corpus

<p>Evaluation Corpus for ARAbic EXternal plagiarism detection (ExAra Corpus)&nbsp;</p> <p>&nbsp;</p> <p>This corpus has been used in AraPlagDet 2015 shared task&nbsp;</p> <p>More details could be found in : <a href="http://araplagdet.misc-lab.org">https://araplagdet.misc-lab.org/</a>&nbsp;or <a href="http://pan.webis.de/fire15/pan15-web/index.html">https://pan.webis.de/fire15/pan15-web/index.html</a></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>I. SYNOPSIS&nbsp;</strong></p> <p>ExAra corpus comprises 2345 documents; almost half of them (suspecious doucuments) contain passages borrowed from the other half (source docucments) to simulate documents that contain plagiarized fragments. The corpus involves 2 parts: Training and test.</p> <p>&nbsp;</p> <p><strong>II. DESCRIPTION&nbsp;</strong></p> <p>Each part of the corpus (training and test) consists mainly of 3 datasets:&nbsp; 2 sets of textual files and 1 set of XML files. The 2 sets of the textual files are the suspicious documents (i.e. the documents that contain artificial plagiarism) and the source documents (i.e., the documents from which the suspicious passages have been plagiarised). The 3rd set of documents contains XML files, which are the plagiarism annotation, i.e., they provide for each plagiarized passage its starting offset and its length in both the suspicious and source documents (offset and length were both expressed in characters). A suspicious document file (.txt) and its plagiarism annotation file (.xml) share the same name.</p> <p>&nbsp;</p> <p><strong>III. PURPOSE&nbsp;</strong></p> <p>The purpose of ExAra corpus is to evaluate automatic plagiarism detection methods, notably methods of the External approach. This approach consists in uncovering the plagiarized passages on the basis of their similarity with passages in the source documents.</p> <p>It should be noted that some suspicious documents in ExAra corpus contain religious quotations (e.g., Quran and Hadith) and common phrases. Some of them appear also in some source documents, and hence a simple plagiarism detection software can consider them as plagiarism. However, quotations and common phrases are legitimate text reuse cases and are not annotated in the XML files in ExAra. Therefore, it is an important feature for the&nbsp;plagiarism detection systems evaluated on ExAra to not consider religious quotations and common phrases as plagiarism cases unless they appear as part of a larger plagiarism case.</p> <p>&nbsp;</p> <p><strong>IV. BUILDING METHODS&nbsp;</strong></p> <p>The documents that compose ExAra corpus do not contain actual plagiarism&nbsp; cases, they are rather artificial suspicious documents in which&nbsp;plagiarism was created automatically by a software that takes fragments&nbsp;of text from one or more sources documents and inserts them in another&nbsp;one according to a set of parameters, namely the percentage of plagiarism&nbsp;and the lengths of the plagiarized passages. Some of the plagiarised fragments are&nbsp;obfuscated manually or automatically before inserting them in the suspicious documents.</p> <p>This building method is the same used to construct PAN 2009-2011 corpora of plagiarism detection (see http://pan.webis.de for more information on PAN competition and its&nbsp;corpora).&nbsp;</p> <p>&nbsp;</p> <p><strong>V. LANGUAGE AND ENCODING&nbsp;</strong></p> <p>All the textual documents of this corpus are written in Arabic language&nbsp;and encoded in UTF-8 without BOM.</p> <p>&nbsp;</p> <p><strong>VI. HOW TO CITE THE CORPUS ?</strong></p> <p>If you publish a paper about your experimentations using ExAra corpus,&nbsp;please cite the following paper:</p> <ul> <li>Bensalem, I., Boukhalfa, I., Rosso, P., Abouenour, L., Darwish, K., &amp; Chikhi, S.:&nbsp;Overview of the AraPlagDet PAN@FIRE2015 Shared Task on Arabic Plagiarism Detection.&nbsp;In P. Majumder, M. Mitra, M. Agrawal, &amp; P. Mehta (Eds.), Post Proceedings of the Workshops at the 7th Forum for Information Retrieval Evaluation (FIRE 2015), Gandhinagar, India,&nbsp;December 4-6, CEUR proceedings vol. 1587 (pp. 111&ndash;122). CEUR-WS.org (2015).</li> </ul> <p>We encourage you to compare your method tested on ExAra with the methods of AraPlagDet&nbsp;competition described in the paper above.</p> <p>&nbsp;</p> <p><strong>VII. WARNING&nbsp;</strong></p> <p>It should be noted that the Arabic texts may contain quotations from the&nbsp;Quran and the Hadith; and due to the fact that text insertion is&nbsp;automatic and in random positions, it is possible that the plagiarized&nbsp;text is inserted unintentionally between Quranic verses or sentences of&nbsp;a Hadith cited in a document. Hence, the inserted passages may alter&nbsp;the meaning of the original text. For these reasons, this corpus must&nbsp;not be used outside the purpose for which it was built. Examples of the&nbsp;inappropriate use include using the corpus documents as a source of&nbsp;knowledge or distributing them without mentioning that they contain&nbsp;borrowed texts. If you are not interested in plagiarism detection, and you are retaining the corpus because it contains articles you want to read,&nbsp;then this corpus is not the right source. Please, you should refer to the&nbsp;sources mentioned in (Bensalem et al. 2015) (i.e.,the paper above) where&nbsp;you can find the original content of the articles you are looking for.</p> <p>We emphasize that we are not responsible for the results of any use of this corpus other than the evaluation of the external plagiarism detection methods.&nbsp;</p> <p>&nbsp;</p> <p><strong>VIII. CONTACT US</strong></p> <p>We will be happy to hear from you about your experience in using ExAra&nbsp;corpus. Please do not hesitate to contact us with the following email&nbsp;address: bens.imene@gmail.com</p> <p>&nbsp;</p> <p>Imene Bensalem&sup1;, Imene Boukhalfa&sup1;, Paolo Rosso&sup2;, Salim Chikhi&sup1;</p> <p>&sup1;MISC Lab. Constantine 2 university, Algeria</p> <p>&sup2;PRHLT, Universitat Polit&egrave;cnica de Val&egrave;ncia, Spain&nbsp;</p>

opencc-by-4.0Aug 2015View details →
zenodo40/100

SemEval-2024 Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes

<p><strong>Task description:</strong>&nbsp;SHROOM participants will need to detect grammatically sound output that contains incorrect semantic information (i.e. unsupported or inconsistent with the source input), with or without having access to the model that produced the output.</p> <p><strong>Overview of the task:</strong>&nbsp;The modern NLG landscape is plagued by two interlinked problems: On the one hand, our current neural models have a propensity to produce inaccurate but fluent outputs; on the other hand, our metrics are most apt at describing fluency, rather than correctness. This leads neural networks to &ldquo;hallucinate&rdquo;, i.e., produce fluent but incorrect outputs that we currently struggle to detect automatically. For many NLG applications, the correctness of an output is however mission critical. For instance, producing a plausible-sounding translation that is inconsistent with the source text puts in jeopardy the usefulness of a machine translation pipeline. With our shared task, we hope to foster the growing interest in this topic in the community.</p> <p>With SHROOM we adopt a post hoc setting, where models have already been trained and outputs already produced: participants will be asked to perform binary classification to identify cases of fluent overgeneration hallucinations in two different setups: model-aware and model-agnostic tracks. That is, participants must detect grammatically sound outputs which contain incorrect or unsupported semantic information, inconsistent with the source input, with or without having access to the model that produced the output. To that end, we will provide participants with a collection of checkpoints, inputs, references and outputs of systems covering three different NLG tasks: definition modeling (DM), machine translation (MT) and paraphrase generation (PG), trained with varying degrees of accuracy. The development set will provide binary annotations from at least five different annotators and a majority vote gold label.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Data set for Plos One Article "Force sharing and other collaborative strategies in a dyadic force perception task"

<p>Data set for Plos One Article :</p> <p>Tatti, F., Baud-Bovy G. (2018) &quot;Force sharing and other collaborative strategies in a dyadic force perception task&quot;. doi: 10.1371/journal.pone.0192754</p> <p>This study investigates how people might interact to extract information from the forces experienced while holding an object together.&nbsp;&nbsp; More specifically, the dyads (i.e. pairs formed two persons) participating to the study had to identify the direction of a small force applied to a jointly held object by a haptic device. This study included a condition where each participant responded independently and another one where the two participants had to agree upon a single negotiated response.</p> <p>The dataset (data.csv) contains the force produced by the haptic device and the average and standard deviation of the interaction force for all trials together with the responses of the participants. We also included the initial and final position of the haptic device and total distance traveled for each trial.</p> <p>The data are in comma separated text format and its description in a PDF document (readme.pdf).</p>

opencc-by-4.0Jan 2018View details →
dryad40/100

Interference in the shared-stroop task: a comparison of self- and other-monitoring

<p>Co-acting participants represent and integrate each other's actions, even when they are not required to monitor one another. However, monitoring the actions of a partner is an important component of successful interactions, and particularly of linguistic interactions. Moreover, monitoring others may rely on similar mechanisms to those that are involved in self-monitoring. In order to investigate the effect of monitoring on shared linguistic representations, we combined a monitoring task with the shared Stroop task. In the shared Stroop task, one participant named the colour of words in one colour (e.g., red) while ignoring stimuli in the other colour (e.g., green); the other participant either named the colour of words in the other colour or did not respond. Crucially, participants either had to provide feedback about the correctness of their partner's response (Experiment 3) or did not (Experiment 2). The results showed that interference was greater when both participants responded than when they did not, but only when partners provided feedback. We argue that feedback increased joint task interference because in order to monitor their partner, participants had to represent their target utterance, and this representation interfered with self-monitoring of their own utterance.</p>

opencc-zeroJul 2021View details →
dryad40/100

Interference in the shared-stroop task: a comparison of self- and other-monitoring

Open the record for dataset details and reuse information.

publicJul 2021View details →
OpenNeuro36/100

Shared and differential default-mode related patterns of activity in an autobiographical, a self-referential and an attentional task

Open the record for dataset details and reuse information.

openJan 2018View details →
zenodo36/100

HIPE-2022 Shared Task Named Entity Datasets

<p>HIPE-2022 datasets used for the <a href="https://hipe-eval.github.io/HIPE-2022/">HIPE 2022 shared task</a> on <strong>named entity recognition and classification (NERC) and entity linking (EL) in multilingual historical documents</strong>.&nbsp;</p> <p>HIPE-2022 datasets are based on six primary datasets assembled and prepared for the shared task. Primary datasets are composed of historical newspapers and classic commentaries covering ca. 200 years, feature several languages and different entity tag sets and annotation schemes. They originate from several European cultural heritage projects, from HIPE organizers&rsquo; previous research project, and from the previous HIPE-2020 campaign. Some are already published, others are released for the first time for HIPE-2022.</p> <p>The HIPE-2022 shared task assembles and prepares these primary datasets in HIPE-2022 release(s), which correspond to a single package composed of neatly structured and homogeneously formatted files.</p> <p>Primary datasets undergo the following preparation steps:</p> <ul> <li>conversion to the HIPE format (with correction of data inconsistencies and metadata consolidation);</li> <li>rearrangement or composition of train and dev splits.</li> </ul> <p>Please also refer to:</p> <ul> <li>HIPE-2022 shared task website: <a href="https://hipe-eval.github.io/HIPE-2022/">https://hipe-eval.github.io/HIPE-2022/</a></li> <li>HIPE-2022 data repository: <a href="https://github.com/hipe-eval/HIPE-2022-data">https://github.com/hipe-eval/HIPE-2022-data</a></li> </ul> <p>Here is an overview of the primary datasets:</p> <table> <tbody> <tr> <td> <p><strong>Dataset alias</strong></p> </td> <td> <p><strong>Readme</strong></p> </td> <td> <p><strong>Document type</strong></p> </td> <td> <p><strong>Languages</strong></p> </td> <td> <p><strong>Suitable for</strong></p> </td> <td> <p><strong>Project</strong></p> </td> </tr> <tr> <td> <p>hipe2020</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-hipe2020.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de, fr, en</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://impresso.github.io/CLEF-HIPE-2020">CLEF-HIPE-2020</a></p> </td> </tr> <tr> <td> <p>newseye</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-newseye.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de, fi, fr, sv</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://www.newseye.eu/">NewsEye</a></p> </td> </tr> <tr> <td> <p>sonar</p> </td> <td> <p>link</p> </td> <td> <p>historical newspapers</p> </td> <td> <p>de</p> </td> <td> <p>NERC-Coarse, EL</p> </td> <td> <p><a href="https://sonar.fh-potsdam.de/">SoNAR</a></p> </td> </tr> <tr> <td> <p>letemps</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-letemps.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>fr</p> </td> <td> <p>NERC-Coarse, NERC-Fine</p> </td> <td> <p>LeTemps</p> </td> </tr> <tr> <td> <p>topres19th</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-topres19th.md">link</a></p> </td> <td> <p>historical newspapers</p> </td> <td> <p>en</p> </td> <td> <p>NERC-Coarse, EL</p> </td> <td> <p><a href="https://livingwithmachines.ac.uk/">Living with Machines</a></p> </td> </tr> <tr> <td> <p>ajmc</p> </td> <td> <p><a href="https://github.com/hipe-eval/HIPE-2022-data/blob/pre-release/documentation/README-ajmc.md">link</a></p> </td> <td> <p>classical commentaries</p> </td> <td> <p>de, fr, en</p> </td> <td> <p>NERC-Coarse, NERC-Fine, EL</p> </td> <td> <p><a href="https://mromanello.github.io/ajax-multi-commentary/">AjMC</a></p> </td> </tr> </tbody> </table> <p>&nbsp;</p> <p>The HIPE-2022 team expresses her greatest appreciation to the partnering projects, namely <a href="https://mromanello.github.io/ajax-multi-commentary/">AJMC</a>, <em><a href="https://impresso-project.ch/">impresso</a>, </em><a href="https://impresso.github.io/CLEF-HIPE-2020/">HIPE-2020</a>, <a href="https://livingwithmachines.ac.uk/"><em>Living with Machines</em></a>, <a href="https://www.newseye.eu/"><em>NewsEye</em></a>, and <a href="https://sonar.fh-potsdam.de/"><em>SoNAR</em></a>, for contributing their NE-annotated datasets (and hiding a part thereof for the time of the evaluation campaign).</p>

opencc-by-nc-4.0Feb 2022View details →
zenodo36/100

CRAFT 2019 Shared Task data

<p>This data set consists of data used for the CRAFT Shared Task 2019.</p> <p>Version 3.1.3 of the CRAFT corpus was provided to participants as training data (CRAFT-3.1.3.tar.gz).</p> <p>During the evaluation phase, participants were provided&nbsp;the 30 plain text documents of the CRAFT evaluation set, along with ontologies used for concept annotation, other concept metadata files, and tokens required for the coreference resolution evaluation. (craft-st-2019-2019_test_data.tar.gz)</p> <p>Finally, the evaluation was completed using the gold standard annotation files available in&nbsp;evaluation-data.tar.gz.</p>

opencc-by-nc-sa-3.0Jul 2019View details →
ClinicalTrials.gov36/100

Building and Sustaining Interventions for Children: Task-sharing Mental Health Care in Low-resource Settings

ClinicalTrials.gov study NCT03243396. IPD Sharing: NO. Countries: 1. Publications: 76.

closedIPD-NOFeb 2026View details →
dryad36/100

Task sharing for point-of-care testing: Review of national health policies and implementation landscape in 19 African countries

Open the record for dataset details and reuse information.

publicDec 2025View details →
zenodo32/100

Shared mobility opporTunities And challenges foR European citieS (STARS) - Work Package 2 - Tasks 2.1 - 2.3

<p>These datasets contain data from the Work Package 2&nbsp;(WP2)&nbsp;in the project STARS. The aims of this WP were:&nbsp;</p> <p>1) To map existing car sharing services and quantitatively assess their relevance in their respective<br> mobility contexts.<br> 2) To define a classification scheme for such services to ease subsequent analyses.<br> 3) To highlight which are the changes in social practices and media usages that are related to the<br> diffusion of car sharing practices.<br> 4) To assess which are the near and medium term development of such services following the<br> trends and the planned implementation actions that are planned by different stakeholders.<br> 5) To point out the main national and European policy barriers and opportunities for the growth of<br> car-sharing.</p> <p>For more information on the project:&nbsp;<a href="http://stars-h2020.eu/">http://stars-h2020.eu/</a></p>

opencc-by-4.0Dec 2019View details →
zenodo32/100

Shared mobility opporTunities And challenges foR European citieS (STARS) - Work Package 5, Task 5.1

<p>These datasets contain information about car sharing members and non-members collected through a survey carried out as part of the Work Package&nbsp;5 (WP5) activities in the STARS project.</p> <p>The data were gathered in three European countries (Italy, Germany and Belgium) and contain&nbsp;information about travel habits, car ownership level and sociodemographic characteristics of the respondents.</p> <p>The data collected were used to:</p> <p>1) Define the maximum portion of travel demand that can be served by car sharing and how this will impact the demand for other travel means.</p> <p>2) Define a &ldquo;rupture scenario&rdquo;, where the benefits of car sharing are maximised.</p> <p>3) Quantify the gap between the business as usual scenario and the rupture scenario, and the related impacts.</p> <p>For more information about the project, please refer to&nbsp;<a href="http://stars-h2020.eu/">http://stars-h2020.eu/</a></p>

opencc-by-4.0Dec 2019View details →
zenodo32/100

Training data for the shared task Ideology and Power Identification in Parliamentary Debates (2024)

<p>This dataset contains a selection of speeches from <a href="https://www.clarin.eu/parlamint">ParlaMint</a> corpora (version 4.0) as the training set for &nbsp;the shared task on "<a href="https://touche.webis.de/clef24/touche24-web/ideology-and-power-identification-in-parliamentary-debates.html">Ideology and Power Identification in Parliamentary Debates</a>" in <a href="https://clef2024.imag.fr/">CLEF 2024</a>.</p> <p>All files are tab-separated text files with the following fields:</p> <ul> <li>"<em>id</em>" is a unique (arbitrary) ID for each text.</li> <li>"<em>speaker</em>" is a unique (arbitrary) ID for each speaker. There may be multiple speeches from the same speaker.</li> <li>"<em>sex</em>" is the (binary/biological) sex of the speaker. This information is collected from varying sources (typically data published by the respective parliament), and in some cases it may be unspecified or unknown.</li> <li>"<em>text</em>" is the transcribed text of the parliamentary speech. Real examples may include line breaks, and other special sequences escaped or quoted.</li> <li>"<em>text_en</em>" is an automatic English translation of the corresponding text. This field may be empty (obviously) &nbsp;for speeches in English, but the translations may be missing for a small number of non-English speeches as well.</li> <li>"<em>label</em>" is the binary/numeric label. For political orientation, 0 is left and 1 is right. For power identification 0 indicates coalition (or governing party) and 1 indicates opposition.</li> </ul> <p>File names indicate the task and the parliament. We provide data from&nbsp;the following national and regional parliaments.</p> <ul> <li>Austria (at)</li> <li>Bosnia and Herzegovina (ba)</li> <li>Belgium (be)</li> <li>Bulgaria (bg)</li> <li>Czechia (cz)</li> <li>Denmark (dk)</li> <li>Estonia (ee) [only political orientation]</li> <li>Spain (es)</li> <li>Catalonia (es-ct)</li> <li>Galicia (es-ga)</li> <li>Basque Country (es-pv) [only power]</li> <li>Finland (fi)</li> <li>France (fr)</li> <li>Great Britain (gb)</li> <li>Greece (gr)</li> <li>Croatia (hr)</li> <li>Hungary (hu)</li> <li>Iceland (is) [only political orientation]</li> <li>Italy (it)</li> <li>Latvia (lv)</li> <li>The Netherlands (nl)</li> <li>Norway (no) [only political orientation]</li> <li>Poland (pl)</li> <li>Portugal (pt)</li> <li>Serbia (rs)</li> <li>Sweden (se) [only political orientation]</li> <li>Slovenia (si)</li> <li>Turkey (tr)</li> <li>Ukraine (ua)</li> </ul> <p>The number of training instances and the class imbalance differs for each training set. We do not provide a fixed validation split. Please see the <a href="https://touche.webis.de/clef24/touche24-web/ideology-and-power-identification-in-parliamentary-debates.html">shared task website</a> for further description of the data set and the sampling process.</p>

opencc-by-4.0Jan 2024View details →
zenodo32/100

DIPROMATS 2024 - Shared Task 2: few-shot training data for narrative identification

<p>Narratives are causally connected sequences of events that are selected and evaluated as meaningful for a particular audience. They make sense of the world by identifying the significance of people, places, objects, and events in time. In international relations, international actors create strategic narratives to &ldquo;construct a shared meaning of the past, present, and future of international politics to shape the behavior of domestic and international actors&rdquo;</p> <p>DIPROMATS 2024 Task 2 is a multiclass multilabel classification problem. Given a series of predefined narratives of each international actor, systems must determine which narrative the tweets belong to. Systems will receive the description of each narrative and a few examples of tweets in both languages (English and Spanish) that belong to each of them (few-shot learning). A tweet may be associated with one, several or none of the narratives.</p> <p>These are the few-shot training datasets for Englsih and Spanish.</p> <p>These files don't contain the narratives description. You can find them in the testing dataset:</p> <p>Pe&ntilde;as, A., Fraile-Hern&aacute;ndez, J. M., Moral, P., Rodrigo, &Aacute;., Deriu, J., Sharma, R., Centeno, R., Rodr&iacute;guez-Garc&iacute;a, R., Giedemann, P., &amp; Reyes-Montesinos, J. (2024). DIPROMATS 2024 - Shared Task 2: testing data for narrative identification (1.0.0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.12663310" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.12663310</a></p>

opencc-by-4.0Mar 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record