Skip to main content
zenodoopen

PAN18 Author Identification: Attribution

<p>We provide&nbsp;a corpus which comprises a set of cross-domain authorship attribution problems in each of the following 5 languages: English, French, Italian, Polish, and Spanish. Note that we specifically avoid to use the term &#39;training corpus&#39; because <strong>the sets of candidate authors of the development and the evaluation corpora are not overlapping</strong>. Therefore, your approach should not be designed to particularly handle the candidate authors of the development corpus.</p> <p>Each problem consists of a set of known fanfics by each candidate author and a set of unknown fanfics located in separate folders. The file <code>problem-info.json</code> that can be found in the main folder of each problem, shows the name of folder of unknown documents and the list of names of candidate author folders.</p> <p>The true author of each unknown document can be seen in the file <code>ground-truth.json</code>, also found in the main folder of each problem.</p> <p>In addition, to handle a collection of such problems, the file <code>collection-info.json</code>includes all relevant information. In more detail, for each problem it lists its main folder, the language (either <code>&quot;en&quot;</code>, <code>&quot;fr&quot;</code>, <code>&quot;it&quot;</code>, <code>&quot;pl&quot;</code>, or <code>&quot;sp&quot;</code>) and encoding (always <code>UTF-8</code>) of its documents.</p> <p>More information:&nbsp;<a href="https://pan.webis.de/clef18/pan18-web/authorship-attribution.html">Link</a></p>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
8
Reuse readiness
0
Engagement
0

Topics