Skip to main content
zenodoopen

German Innsbruck Corpus (GermInnC) 1800-1950

<p><strong>A digital corpus on variation in German (1800-1950)</strong></p> <p>The <em>German Innsbruck Corpus</em> <em>(GermInnC) 1800-1950</em> is a digitised corpus built after the fashion of the <em>German Manchester Corpus (GerManC) 1650-1800</em> (cf. Scheible et al. 2011; Durrell et al. 2012). Hence, the corpus design of the GermInnC is balanced according to period, region and genre.</p> <p>The GermInnC consists of ca. 840,000 tokens, ca. 120,000 per genre (seven in total: Drama, Humanities, Legal texts, Narrative prose, Newspapers, Scientific texts, Sermons). It is subdivided into three periods, 1800-1850, 1851-1900 und 1901-1950, as well as five regions, North German, West Central German, East Central German, West Upper German (including Switzerland), East Upper German (including Austria).</p> <p>The corpus can be retrieved in a raw version, a lemmatised, fully-annotated version, or an &ldquo;all data&rdquo; file (including metadata annotation of file names and periods) for further import and processing. The <em>Stuttgart Tag Set </em>(STTS) and the POS-Tagger <em>TreeTagger </em>was used for linguistic annotation.</p> <p>Two documentation files (word and excel, both included in the download package), provide a more detailed description of the corpus and the digitisation.</p> <p>The corpus may be of interest to all scholars working on the history of the German language, standardisation of German, variation and change, historical sociolinguistics, and Germanic linguistics.</p> <p>The corpus was generously funded by the early career funding of the University of Innsbruck (October 2018 through September 2019).</p>

ShareScore

36/100

Overall dataset sharing score

Score breakdown

These five areas show where the dataset supports — or may limit — practical reuse.

Stewardship
8
Harmonization
4
Access
12
Reuse readiness
8
Engagement
4

Topics