Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

27

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

27 results for “German language”

Learn how ShareScore rates datasets ↗
zenodo48/100

Polifonia Corpus - Books Module Metadata - German Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)

<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie &quot;Forrest Gump&quot; (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation)&nbsp;and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

A German Language Labeled Dataset of Tweets

<p>Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan.&nbsp;</p><p>The dataset consists of 18 samples of tweets with the keyword "Juden" or "Israel." The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over the indicated time period. Each sample was annotated by two expert annotators using an Annotation Portal that visualizes the live tweets in context. We provide the annotation results based on the agreement of two annotators, after discussing discrepancies (Jikeli et al. 2022: 3-6).&nbsp;</p><p>&nbsp;Overall, 335 tweets (4%) were labelled as antisemitic following the IHRA Working Definition of Antisemitism. 1345 tweets (17 %) come from 2019, 1364 tweets (17 %) from 2020, 2639 tweets (33 %) from 2021 and 2700 tweets (34 %) from 2022.&nbsp;</p><p>About half of the tweets, a total of 4,493 tweets (56 %) come from queries with the keyword "Juden," which is representative of a continuous time period from January 2019 to December 2022: 864 tweets (19 %) come from 2019, 891 tweets (20 %) from 2020, 1364 tweets (30 %) from 2021 and 1374 (31 %). 148 out of the 4493 tweets, so 3% from the query with "Juden" are antisemitic.&nbsp;</p><p>The other part of the tweets, a total of 3,555 (44 %)&nbsp; results of queries with the keyword "Israel". 481 tweets (14 %) of the keywords containing Israel stem from 2019, 473 (13 %) come from 2020, 1275 tweets (36 %) from 2021 and 1326 tweets (37 %) are from 2022. Out of all tweets from the "Israel" query, 187 (5 %)&nbsp; are antisemitic.&nbsp;</p><p>The csv file contains diacritics and special characters of the German language (e.g., "ä", "ü", "ö", "ß"), which should be taken into account when opening it with anything other than a text editor.&nbsp;</p><p><strong>Acknowledgements</strong>&nbsp;</p><p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;&nbsp;</p><p>We are grateful for the support of Indiana University's Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University's Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, Emma Shriberg and Victor Tschiskale.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Metadata and annotation data for the XSample corpus on German academic language

<p>The XSample corpus war created in the project <em>XSample</em> (https://www.izus.uni-stuttgart.de/fokus/fdm-projekte/xsample/) at Universit&auml;t Stuttgart in 2021 by Melanie Andresen and Axel Pichler. It contains 135 German academic journal articles, 45 each from the disciplines linguistics, literary studies and philosophy. The texts themselves cannot be made public for copyright reasons. However, metadata and some annotation data are published here.</p> <p><strong>xsample-metadata.csv</strong><br> This file contains metadata on the texts in the corpus, like journal, title, authors, text length, and the URL to the original paper. It also contains two analytical metrics, &#39;past-ratio&#39; and &#39;temp-expr-ratio&#39;, that are based on the annotations in the other two files. The variable &#39;past-ratio&#39; expresses the proportion of verbs in past tense relative to all finite verbs in the text. The variable &#39;temp-expr-ratio&#39; gives the number of temporal expressions per 1000 token.</p> <p><strong>xsample-heidel.csv</strong><br> This file contains all temporal expressions found and classified by the annotation tool <em>HeidelTime</em> (https://github.com/HeidelTime/heideltime, V. 2.2.1, Str&ouml;tgen &amp; Gertz 2013 ). The variable &#39;position&#39; expresses the position of the first character of the temporal expression in the text in characters.</p> <p><strong>xsample-sticker2.csv</strong><br> This file contains all finite verbs found and classified by the annotation tool <em>sticker2</em> (https://github.com/stickeritis/sticker2). The variable &#39;position&#39; expresses the position of the first character of the finite verb in the text in characters.</p> <p><strong>References</strong><br> Str&ouml;tgen, Jannik &amp; Michael Gertz. 2013. Multilingual and cross-domain temporal tagging. <em>Language Resources and Evaluation</em>. Springer 47(2). 269&ndash;298. <a href="https://doi.org/10.1007/s10579-012-9179-y">https://doi.org/10.1007/s10579-012-9179-y</a>.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Polifonia Corpus - Periodicals Module Metadata - German Language

<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p>

opencc-by-4.0Jun 2022View details →
zenodo40/100

A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (annotation)

<p>This dataset contains the annotation of speech spoken in the research cut (Hanke et al. 2014; Hanke et al., 2016) of the movie &quot;Forrest Gump&quot; (Zemeckis, 1994) and its audio-description that was broadcast as an additional audio track (Koop et al., 2009) for visually impaired listeners on Swiss public television. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>

opencc-by-4.0Dec 2020View details →
zenodo40/100

Polifonia Corpus - Encyclopedic Module Metadata - German Language

<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at <a href="http://Full description at https://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

German-Swiss survey on the use of (digital) language resources

<p>This survey on the use of (digital) language resources in German-speaking Switzerland complements the Europe-wide survey conducted by the European Network of e-Lexicography (ENeL) and was carried out in close cooperation with the Leibniz Institute for the German Language (IDS) in Mannheim. The aim was to obtain information on the use of language resources from the user&#39;s perspective.</p>

opencc-by-4.0Dec 2020View details →
zenodo36/100

Lombard speech database for German language

<p>This is a publication of Lombard speech database for German language. Additionally, a GitHub project has been created where information about database updates and related data will be stored:</p> <p>https://github.com/Telecommunication-Telemedia-Assessment/Lombard-Speech-database.git</p> <p>For citations please use:</p> <p>Sołoducha et al., "Lombard speech database for German Language", Proc. of German Annual Conference on Acoustics (DAGA), 2016</p>

opencc-by-4.0Mar 2016View details →
zenodo36/100

Overview: GERMAN LANGUAGE OER FOR SOCIAL SCIENCE METHODS EDUCATION

<p>This document contains the corpus of identified OER as well as OER-like on social science research methods in the German language and the used codes to classify the identified resources. This document is provided for secondary use as well as expansion and revision.</p> <p>The collection and classification of the corresponding OER and OER-like was last updated in August 2021 and has to be discussed accordingly. OER and OER-like that were created later or were offline during data collection are not included.</p> <p>The codings provided are only covering content-based aspects and not feature e.g. like license framework or opportunity to comment and discuss.</p>

opencc-by-4.0May 2022View details →
zenodo36/100

Polifonia Corpus - Encyclopedic Module Data - German Language

<p>We release the data of the Encyclopedic Module of the Polifonia Textual Corpus (Wikipedia pages), collected by selecting from <a href="http://lcl.uniroma1.it/babeldomains/">BabelNet domains</a> all the <a href="https://www.wikipedia.org">Wikipedia</a> musical pages.</p> <p>Full description at <a href="/github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo36/100

Supplementary materials for "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study"

<div> <pre>This is the supplementary data for the article "Phonetic differences between affirmative and feedback head nods in German Sign Language (DGS): A pose estimation study" by Anastasia Bauer, Anna Kuder, Marc Schulder and Job Schepens.<br><br>The supplementary data consists of three components, stored in separate directories:<br>- <code>annotations/</code>: The manual annotations of head nod categories, produced by Anna Kuder and Anastasia Bauer.<br>- <code>pose_analysis/</code>: Code and input/output files for the pose-based automatic analysis of phonetic attributes head nods, produced by Marc Schulder.<br>- <code>statistical_analysis/</code>: Code for the statistical analysis of the other two components and for the creation of related figures, produced by Job Schepens.<br><br>For further details, see the README files of the respective directories.</pre> </div>

opencc-by-4.0Apr 2024View details →
zenodo32/100

Connecting the Dots Variables of Literary History and Emotions in German-language Poetry - Code and Data

<p><strong>Usage</strong></p><p>Experiments can be replicated by running run.py. Please take a look at run.py, since there are some choices to be made within the code. Since pymc4 and gpu usage is under heavy development at the time, we advise to stick to the excact package versions from requirements.txt.</p><p><strong>Content</strong></p><ul><li>data<ul><li>input.tsv - tab-separated table containing corpus metadata (including emotions)&nbsp;</li></ul></li><li>factors<ul><li>model.py - declaration the HGLM in pymc3</li><li>preprocessing.py - code to transform the data in input.tsv to fit the HGLM</li></ul></li><li>plots</li><li>requirements.txt - required python version and packages</li><li>run.py - wrapper to start the experiments (priors and targeted emotions need to be decalered in the code)</li></ul>

opencc-by-4.0Nov 2023View details →
zenodo32/100

Modeling Languages for Digital Twins - A Survey Among the German Automotive Industry

<p>This repository contains the replication package for the paper _Modeling Languages for Digital Twins: A Survey Among the German Automotive Industry_ by J&eacute;r&ocirc;me Pfeiffer, Dominik Fuch&szlig;, Thomas K&uuml;hn, Robin Liebhart, Dirk Neumann, Christer Neim&ouml;ck, Christian Seiler, Anne Koziolek, and Andreas Wortmann.&nbsp;<br>The paper has been submitted to the practice track of &nbsp;[MODELS 2024](https://conf.researchr.org/track/models-2024/models-2024-technical-track#Practice-Track).</p> <h3>Data</h3> <p>This replication package contains all information from the survey:<br>- `results.csv`: A csv version of all data exported from LimeSurvey (German). Personal information from the participants has been removed. This file can be imported to reproduce the extraction results described in our paper.<br>- `survey_german.pdf`: The pdf version of the original survey in German. &nbsp;<br>- `survey_german.md`: A markdown version of the original survey in German.&nbsp;<br>- `survey_english.md`: A markdown version of the survey translated into English.&nbsp;</p> <h3>Selection of participants and distribution</h3> <p>With both versions, the survey can be executed again with a different target audience in English or German. In our case we wanted to reach as much participants from diverse work areas as possible, where we invited the participants by email via an internal mailing list of 189 members of the SofDCar project. &nbsp;To improve the response rate, we implemented two deadline extensions from the initial one-month-long time frame with 2 weeks of additional response time. Together with the deadline extension, we sent a mail to inform and remind the members of the consortium of the survey.</p> <h3>Data extraction</h3> <p>In total, we had 96 participants, of which 43 completed the questionnaire. For incomplete survey responses, we took only the available answers and did not include the missing answers in our data analysis. For data analysis we utilized the commercial Tool IBM SPSS and custom python scripts.</p> <h2>Research Questions&nbsp;</h2> <p>- RQ1: How is the DT understood in the automotive industry?<br>&nbsp; &nbsp; - RQ1.1: For which phases of automotive development are DTs<br>important?<br>&nbsp; &nbsp; - RQ1.2: What are desired properties of DTs?<br>&nbsp; &nbsp; - RQ1.3: What are desired purposes of using DTs?<br>&nbsp; &nbsp; - RQ1.4: How do these purposes change in relation to different phases of automotive development?<br>- RQ2: Which modeling languages and modeling tools are currently employed in the automotive industry?<br>&nbsp; &nbsp; - RQ2.1: Which kinds of models are important during automotive development?<br>&nbsp; &nbsp; - RQ2.2: How important are which models in the phases of automotive development?<br>&nbsp; &nbsp; - RQ2.3: Which tools are used to create and maintain these models?</p>

opencc-by-4.0Jun 2024View details →
ClinicalTrials.gov32/100

Pilot Study to Test the Feasibility and the Efficacy of the German Language Adapted PRO-SELF© Plus Pain Control Program

ClinicalTrials.gov study NCT00920504. IPD Sharing: Not stated. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Design and Validation of a German Language Questionnaire for Measuring Alarm Fatigue in Intensive Care Units

ClinicalTrials.gov study NCT04994600. IPD Sharing: UNDECIDED. Countries: 1. Publications: 0.

restrictedIPD-UNDECIDEDFeb 2026View details →
ClinicalTrials.gov32/100

Parent-adolescent Communication: Validation of a German Language Scale and Its Longitudinal Association With Adolescent Mental Health

ClinicalTrials.gov study NCT05332236. IPD Sharing: UNDECIDED. Countries: 1. Publications: 2.

restrictedIPD-UNDECIDEDFeb 2026View details →
zenodo28/100

THE NECESSITY OF PSYCHOLINGUISTICS IN THE DEVELOPMENT OF LANGUAGE SKILLS OF THE PUPILS IN GERMAN LANGUAGE CLASSES

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

THE IMPORTANCE OF PSYCHOLINGUISTICS IN THE DEVELOPMENT OF LANGUAGE SKILLS OF STUDENTS IN GERMAN LANGUAGE LESSONS

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →
zenodo28/100

IMPORTANCE OF METHODS AND TECHNIQUES IN TEACHING GERMAN LANGUAGE

Open the record for dataset details and reuse information.

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record