Enhanced Latin Lemma Dataset
<h2><strong>Overview</strong></h2> <p>The Latin Lexicon Dataset contains information about Latin words collected through webscraping from Wiktionary. The dataset includes various linguistic features such as part of speech, lemma, aspect, tense, verb form, voice, mood, number, person, case, and gender. Additionally, it provides source URLs and links to the Wiktionary pages for further reference. The dataset aims to contribute to linguistic research and analysis of Latin language elements.</p> <h3><strong>Versions of the Dataset</strong></h3> <p>This dataset is available in three versions, each offering varying levels of refinement:</p> <ul> <li><strong>wiki_latin_data_v1.csv(v1):</strong> The initial raw version, containing all webscraped data without extensive cleaning or filtering.</li> <li><strong>wiki_latin_data_v2.csv(v2):</strong> A more processed version, where some inconsistencies and duplicates were removed, and linguistic features were better aligned.</li> <li><strong><strong>wiki_latin_data_v3.csv (v3)</strong>:</strong> The most refined version, offering a clean, well-organized dataset with comprehensive linguistic features and translation equivalents with minimal errors. This version is recommended for most use cases.</li> </ul> <h3><strong>Data Source:</strong></h3> <ul> <li>Webscraped from Wiktionary</li> </ul> <h3><strong>Produced by:</strong></h3> <ul> <li>Python-based web scraping algorithms</li> </ul>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 12
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 0
- Engagement
- 4