Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
9
datasets available to search
ShareScore release 0.7.1
Dataset results
9 results for “interpreting corpus”
Speech Corpus of Interpreted Premier Press Conferences (SCIPPC)
<p>SCIPPC v1.0 is a parallel corpus of consecutive interpreting between Mandarin Chinese and English and vice versa in two Chinese premiers’ press conferences in March 2003–2007 and 2013–2017. The conferences were held after sessions of the National People’s Congress and the Chinese People’s Political Consultative Conference. They were moderated by spokespersons of the Congress and the Chinese Ministry of Foreign Affairs and attended by journalists, who asked the premiers questions.</p> <p>SCIPPC v1.0 includes source speeches by approximately 170 speakers and interpretations by six different staff interpreters of the Chinese Ministry of Foreign Affairs, who worked into their B language. It contains 192,209 tokens (source: 108,296, target: 83,913; Chinese: 112,528, English: 79,681) and 19 h 43 min 5 s of video recordings. It is fully transcribed and aligned at the recording–transcript and source–target transcript levels.</p>
Webis Query Interpretation Corpus 2022 (Webis-QInC-22)
<p><strong>Webis-QInC-22</strong><em> </em></p> <p>The Webis Query Interpretation Corpus 2022 (Webis-QInC-22) contains manually selected explicit entities, implicit entities and entity based interpretations for 3,026 web queries. These web queries were either obtained from existing entity linking datasets or ambiguity queries from various sources.</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 4. The frequency of different source segments of "but" in the Cantonese sub-corpus
<p>Likewise, Figure 4 lists all the Cantonese source segments of but and their frequency. The results show that most of the use of but was contributed by daan (hai) — its closest equivalence in Cantonese. However, there were 52 instances of the use of but which corresponded to no source segments at all in the Cantonese sub-corpus, indicating, again, the possibility of explicitation. The rest of the 16 instances of but were attributed by the use of the Cantonese markers ji (而 (frequency=8; close in meaning to however and but, indicating concession or contrast), followed by bat gwo (frequency=7) and koek (卻 (frequency=1; close in meaning to however and but, indicating concession or contrast).</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 3. The frequency of different source segments of "however" in the Cantonese sub-corpus
<p>The use of however and but in the English sub-corpus was also investigated through looking into their source segments in the Cantonese sub-corpus. Figure 3 lists all the Cantonese source segments of however and their frequency. The figures show that the use of however mostly resulted from the employment of daan (hai) in the Cantonese source text (frequency=52). There were 24 instances of however which corresponded to no source segment at all, indicating a possible trend of explicitation — a feature of interpreted or translated language. Only 21 cases of however were caused by the use of bat gwo — its closest equivalence in Cantonese. The rest 4 cases resulted from the use of the Cantonese marker ho si (可是) (frequency=3; close in meaning to both however and but in English, indicating concession or contrast) and zeon gun (儘管 (frequency=1; close in meaning to although or in spite of in English, indicating concession or topic change).</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 2. The frequency of different renditions of "daan (hai)" in the English sub-corpus
<p>The renditions of daan (hai) show a similar pattern (Figure 2). Among the 287 cases of daan (hai), the majority were interpreted into but (frequency=74) — its closest equivalence in English that signals denial and contrast. A total of 60 daan (hai) received no interpretation at all, which, again, indicates a possible mitigation strategy employed by the interpreter(s). Similar to the case of bat gwo, however — the other frequently used English marker apart from but — was the next most often employed rendition of daan (hai) in the English sub-corpus. In addition to but and however, daan (hai) was also rendered into the following English markers: although, despite, while, nevertheless, yet, having said that / that said, nonetheless, though, notwithstanding, on the other, regardless, even so, and after, most of which signal concession and topic change, with an even higher degree of subtlety as compared to however.</p>
A Corpus-driven Study of Contrastive Markers in Cantonese‒English Political Interpreting-Figure 1. The frequency of different renditions of "bat gwo" in the English sub-corpus
<p>The renditions of bat gwo and daan (hai) were closely examined by looking into how they were interpreted. Figure 1 lists all the renditions of bat gwo and their frequency. The figures show that bat gwo was most often interpreted into however (frequency=22) — its closest equivalence in English. There were, however, 7 cases that bat gwo was interpreted into but — its stronger and less subtle correspondence that indicates denial and contrast. Apart from rendering into these two most common English contrastive markers, there were also 5 cases that bat gwo was not interpreted at all, suggesting a possible mitigation strategy employed by the interpreter(s). Likewise, the rest of bat gwo were rendered into other markers including nevertheless, nonetheless, while, yet, having said that / that said, all of which indicate concession and topic change, yet with an even higher degree of subtlety as compared to however.</p>
European Parliament Interpreting Corpus (EPIC)
<p>EPIC v2.0 is a parallel and trilingual (English, Italian, and Spanish) corpus of European Parliament (EP) speeches and their simultaneous interpretations. The data were collected from EP sessions in Feb–Apr, and July 2004, including the speeches of 175 speakers and an unknown number of interpreters. The current version of the EPIC (v2.0) contains 692,585 tokens (source: 247,385, target: 445,200) and 83 h 36 min 14 s of audiovisual recordings (source videos: 27 h 32 min 20 s, target audio: 56 h 3 min 54 s). It is fully transcribed, annotated, and aligned at the recording–transcript level.</p>
BSL-Hansard: A parallel, multimodal corpus of English and interpreted British Sign Language data from parliamentary proceedings
<p>BSL-Hansard is a novel open source and multimodal resource composed by combining Sign Language video data in BSL and English text from the official transcription of British parliamentary sessions. This paper describes the method followed to compile BSL-Hansard including time alignment of text using the MAUS (Schiel, 2015) segmentation system, gives some statistics about this dataset, and suggests experiments. These primarily include end-to-end Sign Language-to-text translation, but is also relevant for broader machine translation, and speech and language processing tasks.</p> <p>This dataset will be useful for translation between BSL and English, or for studies in BSL or English down to the phonetic level.</p>
Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff
<p>Dataset for publication "Knowledge-driven compound interpretation. A corpus study on German complex nouns headed by -stoff"</p> <p>To appear in: SKASE Journal of Theoretical Linguistics, ISSN: <a href="https://portal.issn.org/resource/ISSN/1336-782X">1336-782X</a></p> <p>Authors: Olav Mueller-Reichau (University of Leipzig, Germany) and Matthias Irmer (OntoChem GmbH, Halle (Saale), Germany)</p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.