SEEFLEX - The Corpus of Secondary English as a Foreign Language (EFL) Exams
<blockquote> <p>This version of the corpus may have been superseded by the <a href="https://github.com/tobib92/SEEFLEX/" target="_blank" rel="noopener">GitHub repository</a></p> </blockquote> <p> </p> <p><strong>Abstract:</strong></p> <p>This report presents the <em>Corpus of Secondary School English as a Foreign Language (EFL) Exams (SEEFLEX)</em>. In Germany, upper secondary school EFL exams feature recurring tasks targeting diverse text types. The <em>SEEFLEX</em> was developed to investigate how students complete these tasks linguistically and whether they meet the curricular requirements. The corpus contains data from 575 transcribed authentic curriculum-based examinations (n<sub>texts</sub> = 1979, total = ~625.000 words). The metadata include standardized receptive vocabulary assessments, a cognition scale, the participants’ reading habits, social background, and their language experience and proficiency. Extensive xml mark-up was added to investigate the influence of inter alia source material, structural text features, and selected language mistakes. An online repository provides full-text access as well as ample additional resources, including an interactive Shiny application to investigate register variation in the corpus.</p>
ShareScore
36/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 12
- Reuse readiness
- 8
- Engagement
- 4