Skip to main content
zenodoopen

PAN23 Multi-Author Writing Style Analysis

<p>This is the dataset for the shared task on&nbsp;<a href="https://pan.webis.de/clef23/pan23-web/style-change-detection.html">Multi-Author Writing Style Analysis PAN@CLEF2023</a>. Please consult the task&#39;s page for further details on the format, the dataset&#39;s creation, and links to baselines and utility code.</p> <p><strong>Task:&nbsp;</strong>We ask participants to solve the following intrinsic style change detection task:&nbsp;<strong>for a given text, find all positions of writing style change on the paragraph-level</strong>&nbsp;(i.e., for each pair of consecutive paragraphs, assess whether there was a style change). The simultaneous change of authorship and topic will be carefully controlled and we will provide participants with datasets of three difficulty levels:</p> <ol> <li><strong>Easy:</strong>&nbsp;The paragraphs of a document cover a variety of topics, allowing approaches to make use of topic information to detect authorship changes.</li> <li><strong>Medium:</strong>&nbsp;The topical variety in a document is small (though still present) forcing the approaches to focus more on style to effectively solve the detection task.</li> <li><strong>Hard:</strong>&nbsp;All paragraphs in a document are on the same topic.</li> </ol> <p>All documents are provided in English and may contain an arbitrary number of style changes. However, style changes may only occur between paragraphs (i.e., a single paragraph is always authored by a single author and contains no style changes).</p> <p><strong>Data:&nbsp;</strong>To develop and then test your algorithms, three datasets including ground truth information are provided (<em>dataset1</em>&nbsp;for the easy task,&nbsp;<em>dataset2</em>&nbsp;for the medium task, and&nbsp;<em>dataset3</em>&nbsp;for the hard task).</p> <p>Each dataset is split into three parts:</p> <ol> <li><em>training set:</em>&nbsp;Contains 70% of the whole dataset and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em>&nbsp;Contains 15% of the whole dataset and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em>&nbsp;Contains 15% of the whole dataset, no ground truth data is given. This set is used for evaluation.</li> </ol> <p>You are free to use additional external data for training your models. However, we ask you to make the additional data utilized freely available under a suitable license.</p> <p><strong>Versioning:</strong>&nbsp;</p> <ul> <li>1.0: initial upload</li> </ul>

ShareScore

32/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
8
Reuse readiness
8
Engagement
4

Topics