PAN18 Author Identification: Attribution
<p>We provide a corpus which comprises a set of cross-domain authorship attribution problems in each of the following 5 languages: English, French, Italian, Polish, and Spanish. Note that we specifically avoid to use the term 'training corpus' because <strong>the sets of candidate authors of the development and the evaluation corpora are not overlapping</strong>. Therefore, your approach should not be designed to particularly handle the candidate authors of the development corpus.</p> <p>Each problem consists of a set of known fanfics by each candidate author and a set of unknown fanfics located in separate folders. The file <code>problem-info.json</code> that can be found in the main folder of each problem, shows the name of folder of unknown documents and the list of names of candidate author folders.</p> <p>The true author of each unknown document can be seen in the file <code>ground-truth.json</code>, also found in the main folder of each problem.</p> <p>In addition, to handle a collection of such problems, the file <code>collection-info.json</code>includes all relevant information. In more detail, for each problem it lists its main folder, the language (either <code>"en"</code>, <code>"fr"</code>, <code>"it"</code>, <code>"pl"</code>, or <code>"sp"</code>) and encoding (always <code>UTF-8</code>) of its documents.</p> <p>More information: <a href="https://pan.webis.de/clef18/pan18-web/authorship-attribution.html">Link</a></p>
ShareScore
24/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 8
- Access
- 8
- Reuse readiness
- 0
- Engagement
- 0