Skip to main content
zenodoopen

Genre-KI-04

<p>The web genre corpus 2004 (Genre-KI-04) is designed for the evaluation of techniques for genre classification. It consists of 1239 web documents classified into 8 genres and basic meta data for each of the files.</p> <p>The corpus consists of the HTML documents grouped into directories according to their respective genre. The first lines of each document contain the meta information for each document in a HTML comment. This information includes the URL the document was downloaded from as well as the document title and the parsed text.</p> <p>A definition of the genres can be found in the paper (http://dx.doi.org/10.1007/978-3-540-30221-6_20) or in the corpus. The distribution of the documents among the genres is summarized below.</p> <ul> <li>Articles: 127 (10.2%)</li> <li>Discussion: 127 (10.3%)</li> <li>Download: 152 (12.3%)</li> <li>Help: 140 (11.3%)</li> <li>Link lists: 208 (16.8%)</li> <li>Portrait (non private): 179 (14.4%)</li> <li>Portrait (private): 131 (10.6%)</li> <li>Shop: 175 (14.1%)</li> <li><strong>Sum: </strong>1239 (100%)</li> </ul>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
16
Reuse readiness
8
Engagement
0

Topics