Speech Corpus of Interpreted Premier Press Conferences (SCIPPC)
<p>SCIPPC v1.0 is a parallel corpus of consecutive interpreting between Mandarin Chinese and English and vice versa in two Chinese premiers’ press conferences in March 2003–2007 and 2013–2017. The conferences were held after sessions of the National People’s Congress and the Chinese People’s Political Consultative Conference. They were moderated by spokespersons of the Congress and the Chinese Ministry of Foreign Affairs and attended by journalists, who asked the premiers questions.</p> <p>SCIPPC v1.0 includes source speeches by approximately 170 speakers and interpretations by six different staff interpreters of the Chinese Ministry of Foreign Affairs, who worked into their B language. It contains 192,209 tokens (source: 108,296, target: 83,913; Chinese: 112,528, English: 79,681) and 19 h 43 min 5 s of video recordings. It is fully transcribed and aligned at the recording–transcript and source–target transcript levels.</p>
ShareScore
44/100
Overall dataset sharing score
Score breakdown
These five areas show where the dataset supports — or may limit — practical reuse.
- Stewardship
- 8
- Harmonization
- 4
- Access
- 16
- Reuse readiness
- 8
- Engagement
- 8