Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
5
datasets available to search
ShareScore release 0.9.0
Dataset results
5 results for “authorship verification”
PAN20 Authorship Analysis: Authorship Verification
<p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts' writing styles.</p> <p>In the coming three years at PAN 2020 to PAN 2022, we develop a new experimental setup that addresses three key questions in authorship verification that have not been studied at scale to date:</p> <ul> <li> <p>Year 1 (PAN 2020): Closed-set verficiation.<br> Given a large training dataset comprising of known authors who have written about a given set of topics, the test dataset contains verification cases from a subset of the authors and topics found in the training data.</p> </li> <li> <p>Year 2 (PAN 2021): Open-set verification.<br> Given the training dataset of Year 1, the test dataset contains verification cases from previously unseen authors and topics.</p> </li> <li> <p>Year 3 (PAN 2022): <em>Suprise task</em>.<br> The task of the last year of this evaluation cycle (to be announced at a later time) will be designed with an eye on realism and practical application.</p> </li> </ul> <p>This evaluation cycle on authorship verification provides for a renewed challenge of increasing difficulty within a large-scale evaluation. We invite you to plan ahead and participate in all three of these tasks.</p> <p>More information at: <a href="https://pan.webis.de/clef20/pan20-web/author-identification.html">PAN @ CLEF 2020 - Authorship Verification</a></p> <p> </p> <p><strong>Citing the Dataset</strong></p> <p>If you use this dataset for your research, please be sure to cite the following paper:<br> <br> Sebastian Bischoff, Niklas Deckers, Marcel Schliebs, Ben Thies, Matthias Hagen, Efstathios Stamatatos, Benno Stein, and Martin Potthast. The Importance of Suppressing Domain Style in Authorship Analysis. CoRR, abs/2005.14714, May 2020.</p> <p>Bibtex:</p> <pre><code>@Article{stein:2020k, author = {Sebastian Bischoff and Niklas Deckers and Marcel Schliebs and Ben Thies and Matthias Hagen and Efstathios Stamatatos and Benno Stein and Martin Potthast}, journal = {CoRR}, month = may, title = {{The Importance of Suppressing Domain Style in Authorship Analysis}}, url = {https://arxiv.org/abs/2005.14714}, volume = {abs/2005.14714}, year = 2020 }</code></pre> <p> </p>
PAN22 Authorship Analysis: Authorship Verification
<p><strong>Download</strong></p> <p>Access to our corpus can be requested via the Aston Institute for Forensic Linguistics Databank: <a href="https://fold.aston.ac.uk/handle/123456789/17">https://fold.aston.ac.uk/handle/123456789/17</a></p> <p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts' writing styles. In previous editions of PAN, we explored the effectiveness of authorship verification technology in several languages and text genres. In the two most recent editions, cross-domain authorship verification using fanfiction texts was examined. Despite certain differences between fandoms, the task of cross-fandom authorship verification has proved to be relatively feasible. In the current edition, we focus on more challenging scenarios where each author verification case considers two texts that belong to different DTs (cross-DT authorship verification). This will allow us to study the ability of stylometric approaches to capture authorial characteristics that remain stable across DTs even when very different forms of expression are imposed by the DT norms.</p> <p>Based on a new corpus in English, we provide cross-DT authorship verification cases using the following DTs:</p> <ul> <li>Essays</li> <li>Emails</li> <li>Text messages</li> <li>Business memos</li> </ul> <p>The corpus comprises texts of around 100 individuals. All individuals have similar age (18-22) and are native English speakers. The topic of text samples is not restricted while the level of formality can vary within a certain DT (e.g., text messages may be addressed to family members or non-familial acquaintances).</p> <p>More information at: <a href="https://pan.webis.de/clef22/pan22-web/author-identification.html">Authorship Verification 2022</a></p>
VeriDark SilkRoad1 Authorship Verification Dataset
<pre>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</pre>
VeriDark DarkReddit+ Authorship Verification Dataset
<p>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</p>
VeriDark Agora Authorship Verification Dataset
<pre>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.