Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

5

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

5 results for “authorship verification”

Learn how ShareScore rates datasets ↗
zenodo32/100

PAN20 Authorship Analysis: Authorship Verification

<p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts&#39; writing styles.</p> <p>In the coming three years at PAN&nbsp;2020 to PAN&nbsp;2022, we develop a new experimental setup that addresses three key questions in authorship verification that have not been studied at scale to date:</p> <ul> <li> <p>Year 1 (PAN 2020): Closed-set verficiation.<br> Given a large training dataset comprising of known authors who have written about a given set of topics, the test dataset contains verification cases from a subset of the authors and topics found in the training data.</p> </li> <li> <p>Year 2 (PAN 2021): Open-set verification.<br> Given the training dataset of Year&nbsp;1, the test dataset contains verification cases from previously unseen authors and topics.</p> </li> <li> <p>Year 3 (PAN 2022):&nbsp;<em>Suprise task</em>.<br> The task of the last year of this evaluation cycle (to be announced at a later time) will be designed with an eye on realism and practical application.</p> </li> </ul> <p>This evaluation cycle on authorship verification provides for a renewed challenge of increasing difficulty within a large-scale evaluation. We invite you to plan ahead and participate in all three of these tasks.</p> <p>More information at:&nbsp;<a href="https://pan.webis.de/clef20/pan20-web/author-identification.html">PAN @ CLEF 2020 - Authorship Verification</a></p> <p>&nbsp;</p> <p><strong>Citing the Dataset</strong></p> <p>If you use this dataset for your research, please be sure to cite the following paper:<br> <br> Sebastian Bischoff, Niklas Deckers, Marcel Schliebs, Ben Thies, Matthias Hagen, Efstathios Stamatatos, Benno Stein, and Martin Potthast.&nbsp;The Importance of Suppressing Domain Style in Authorship Analysis.&nbsp;CoRR,&nbsp;abs/2005.14714,&nbsp;May&nbsp;2020.</p> <p>Bibtex:</p> <pre><code>@Article{stein:2020k, author = {Sebastian Bischoff and Niklas Deckers and Marcel Schliebs and Ben Thies and Matthias Hagen and Efstathios Stamatatos and Benno Stein and Martin Potthast}, journal = {CoRR}, month = may, title = {{The Importance of Suppressing Domain Style in Authorship Analysis}}, url = {https://arxiv.org/abs/2005.14714}, volume = {abs/2005.14714}, year = 2020 }</code></pre> <p>&nbsp;</p>

openDec 2019View details →
zenodo16/100

PAN22 Authorship Analysis: Authorship Verification

<p><strong>Download</strong></p> <p>Access to our corpus can be requested via the Aston Institute for Forensic Linguistics Databank:&nbsp;<a href="https://fold.aston.ac.uk/handle/123456789/17">https://fold.aston.ac.uk/handle/123456789/17</a></p> <p><strong>Task</strong></p> <p>Authorship verification is the task of deciding whether two texts have been written by the same author based on comparing the texts&#39; writing styles. In previous editions of PAN, we explored the effectiveness of authorship verification technology in several languages and text genres. In the two most recent editions, cross-domain authorship verification using fanfiction texts was examined. Despite certain differences between fandoms, the task of cross-fandom authorship verification has proved to be relatively feasible. In the current edition, we focus on more challenging scenarios where each author verification case considers two texts that belong to different DTs (cross-DT authorship verification). This will allow us to study the ability of stylometric approaches to capture authorial characteristics that remain stable across DTs even when very different forms of expression are imposed by the DT norms.</p> <p>Based on a new corpus in English, we provide cross-DT authorship verification cases using the following DTs:</p> <ul> <li>Essays</li> <li>Emails</li> <li>Text messages</li> <li>Business memos</li> </ul> <p>The corpus comprises texts of around 100 individuals. All individuals have similar age (18-22) and are native English speakers. The topic of text samples is not restricted while the level of formality can vary within a certain DT (e.g., text messages may be addressed to family members or non-familial acquaintances).</p> <p>More information at:&nbsp;<a href="https://pan.webis.de/clef22/pan22-web/author-identification.html">Authorship Verification 2022</a></p>

restrictedMar 2022View details →
zenodo16/100

VeriDark SilkRoad1 Authorship Verification Dataset

<pre>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</pre>

restrictedJul 2022View details →
zenodo16/100

VeriDark DarkReddit+ Authorship Verification Dataset

<p>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</p>

restrictedJul 2022View details →
zenodo16/100

VeriDark Agora Authorship Verification Dataset

<pre>VeriDark (Authorship Verification in the DarkNet) is a benchmark for evaluating authorship analysis methods in a cybersecurity context, by introducing datasets gathered from the DarkNet marketplace forums or from Darknet-related discussions on Reddit. This benchmark contains three datasets for authorship verification and one dataset for authorship identification.</pre>

restrictedJul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record