PAN23 Multi-Author Writing Style Analysis
<p>This is the dataset for the shared task on <a href="https://pan.webis.de/clef23/pan23-web/style-change-detection.html">Multi-Author Writing Style Analysis PAN@CLEF2023</a>. Please consult the task's page for further details on the format, the dataset's creation, and links to baselines and utility code.</p> <p><strong>Task: </strong>We ask participants to solve the following intrinsic style change detection task: <strong>for a given text, find all positions of writing style change on the paragraph-level</strong> (i.e., for each pair of consecutive paragraphs, assess whether there was a style change). The simultaneous change of authorship and topic will be carefully controlled and we will provide participants with datasets of three difficulty levels:</p> <ol> <li><strong>Easy:</strong> The paragraphs of a document cover a variety of topics, allowing approaches to make use of topic information to detect authorship changes.</li> <li><strong>Medium:</strong> The topical variety in a document is small (though still present) forcing the approaches to focus more on style to effectively solve the detection task.</li> <li><strong>Hard:</strong> All paragraphs in a document are on the same topic.</li> </ol> <p>All documents are provided in English and may contain an arbitrary number of style changes. However, style changes may only occur between paragraphs (i.e., a single paragraph is always authored by a single author and contains no style changes).</p> <p><strong>Data: </strong>To develop and then test your algorithms, three datasets including ground truth information are provided (<em>dataset1</em> for the easy task, <em>dataset2</em> for the medium task, and <em>dataset3</em> for the hard task).</p> <p>Each dataset is split into three parts:</p> <ol> <li><em>training set:</em> Contains 70% of the whole dataset and includes ground truth data. Use this set to develop and train your models.</li> <li><em>validation set:</em> Contains 15% of the whole dataset and includes ground truth data. Use this set to evaluate and optimize your models.</li> <li><em>test set:</em> Contains 15% of the whole dataset, no ground truth data is given. This set is used for evaluation.</li> </ol> <p>You are free to use additional external data for training your models. However, we ask you to make the additional data utilized freely available under a suitable license.</p> <p><strong>Versioning:</strong> </p> <ul> <li>1.0: initial upload</li> </ul>
ShareScore
32/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 8
- Reuse readiness
- 8
- Engagement
- 4