Skip to main content
zenodoopen

Webis-Simple-Sentences-17 Corpus

<p>A corpus of 471,085,690 English sentences extracted from the ClueWeb12 Web Crawl. The sentences were sampled from a larger corpus to achieve a level of sentence complexity similar to the one of sentences that humans make up as a memory aid for remembering passwords. Sentence complexity was determined by syllables per word.</p> <p>The corpus is split in training and test set as it is used in the associated publication.&nbsp; The test set is extracted from part 00 of the ClueWeb12, while the training set is extracted from the other parts.</p> <p>More information on the corpus can be found on the corpus web page at our university (listed under documented by).</p>

ShareScore

44/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
8
Access
20
Reuse readiness
8
Engagement
0

Topics