Skip to main content
zenodoopen

PAN15 Author Identification: Verification

<p>We provide you with a training corpus that comprises a set of author verification problems in several languages/genres. Each problem consists of some (up to five) known documents by a single person and exactly one questioned document. All documents within a single problem instance will be in the same language. However, their genre and/or topic may differ significantly. The document lengths vary from a few hundred to a few thousand words.</p> <p>The documents of each problem are located in a separate folder, the name of which (problem ID) encodes the language of the documents. The following list shows the available sub-corpora, including their language, type (cross-genre or cross-topic), code, and examples of problem IDs:</p> <p>Language; Type; Code; Problem IDs<br> Dutch; Cross-genre; DU; DU001, DU002, DU003, etc.<br> English; Cross-topic; EN; EN001, EN002, EN003, etc.<br> Greek; Cross-topic; GR; GR001, GR002, GR003, etc.<br> Spanish; Cross-genre; SP; SP001, SP002, SP003, etc.</p> <p>The ground truth data of the training corpus found in the file <code>truth.txt</code> include one line per problem with problem ID and the correct binary answer (Y means the known and the questioned documents are by the same author and N means the opposite). For example:</p> <pre>EN001 N EN002 Y EN003 N ...</pre>

ShareScore

24/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
8
Reuse readiness
0
Engagement
0

Topics