Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
724
datasets available to search
ShareScore release 0.9.0
Dataset results
724 results for “german”
Molecular Markers for Predicting Treatment Outcome in Patients with Rectal Cancer: A Comprehensive Analysis from the German Rectal Cancer Trials
GEO Series GSE40492. Homo sapiens. 245 samples. Type: Expression profiling by array.
Fig. 5 in The German-Russian deep-sea expedition KuramBio (Kurile Kamchatka biodiversity studies) on board of the RV Sonne in 2012 following the footsteps of the legendary expeditions with RV Vityaz
Fig. 5. Stations sampled during the KuramBio expedition with RV Sonne in 2012.
Tweets in German language
<p>With this dataset, we aim to understand mechanisms of COVID-19 epidemic-related social behavior in German speaking countries deploying methods of computational social science and digital epidemiology. </p><p>- extracted 1632030 Tweets (text and date only) with #Coronavirus in the language German using Twitter API (prospective joined with academic Twitter retrospective function) between 15.01.2020-31.03.2023</p><p> </p><p>German Research Foundation for COVINT project (458528774)</p>
GeNeG: German News Knowledge Graph
<p>GeNeG is a knowledge graph constructed from news articles on the topic of refugees and migration, collected from German online media outlets. GeNeG contains rich textual and metadata information, as well as named entities extracted from the articles' content and metadata and linked to Wikidata. The graph is expanded with up to three-hop neighbors from Wikidata of the initial set of linked entities.</p> <p>GeNeG comes in three flavors:</p> <ul> <li>Base GeNeG: contains textual information, metadata, and linked entities extracted from the articles.</li> <li>Entities GeNeG: derived from the Base GeNeG by removing all literal nodes, it contains only resources and it is enriched with three-hop Wikidata neighbors of the entities extracted from the articles.</li> <li>Complete GeNeG: the combination of the Base and Entities GeNeG, it contains both literals and resources.</li> </ul> <p>Information about uploaded files:</p> <p>(all files are b-zipped and in the N-Triples format.)</p> <table align="left"> <thead> <tr> <th scope="col"><strong>File</strong></th> <th scope="col"><strong>Description</strong></th> </tr> </thead> <tbody> <tr> <td>geneg_<em>type</em>-metadata.nt.bz2</td> <td>Metadata about the dataset, described using void vocabulary.</td> </tr> <tr> <td>geneg_<em>type</em>-instances_types.nt.bz2</td> <td>Class definitions of articles and events.</td> </tr> <tr> <td>geneg_<em>type</em>-instances_labels.nt.bz2</td> <td>Labels of instances.</td> </tr> <tr> <td>geneg_<em>type</em>-instances_metadata_literals.nt.bz2</td> <td>Relations between news article resurces and metadata literals (e.g. URL, publishing date, modification date, polarity score, stance).</td> </tr> <tr> <td>geneg_<em>type</em>-instances_metadata_resources.nt.bz2</td> <td>Relations between news article resources and metadata entities (i.e. publishers, authors, keywords).</td> </tr> <tr> <td>geneg_<em>type</em>-instances_content_relations.nt.bz2</td> <td>Relations between news article resources and content components (e.g. titles, abstracts, article bodies).</td> </tr> <tr> <td>geneg_<em>type</em>-instances_event_mapping.nt.bz2</td> <td>Mapping of news article resources to events.</td> </tr> <tr> <td>geneg_<em>type</em>-event_relations.nt.bz2</td> <td>Relations between news events and entities mentioned (i.e. actors, places, mentions).</td> </tr> <tr> <td>geneg_<em>type</em>-wiki_relations.nt.bz2</td> <td>Relations between news event Wikidata entities and their <em>k</em>-hop entities neighbors from Wikidata.</td> </tr> </tbody> </table> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p> </p> <p><strong>Changelog</strong></p> <p>v1.0.1</p> <ul> <li>Stance annotations have been added to the Base and Complete GeNeG.</li> </ul>
Supplementary material for a mixed-methods study on research impact in a study on lived experiences of German occupation children
<p>Study material (questionnaire) as well as material regarding the qualitative analysis of open-ended questions, s.a. category system, analysis table, additional statistic analyses.</p> <p> </p> <p>Supplement 1: Questionnaire used in the study</p> <p>Supplement 2: Category system of qualitative analysis with sample quotes</p> <p>Supplement 3: Analysis table of qualitative analysis with all corresponding quotes</p> <p>Supplement 4: Statistical comparison of responders (participating in the impact study) to non-responders of the initial study on lived experiences</p>
SPSS Responses of German Energy Commons to Questionnaire
Open the record for dataset details and reuse information.
Collection of Individual Educational Plans (German)
<p>After obtaining permission from the relevant authorities, Individual Educational Plans (IEPs) were collected in one federal state in Germany at the beginning of 2019. Participation was entirely voluntary. The sample contained 112 fully anonymised IEPs from seven schools.</p> <p>The documents contained in this dataset have all been written in German.</p> <p>From primary schools came 54.5 % of the IEPs and 45.5 % from secondary schools. Forty-one IEPs (36.6 %) did not have a clear indication of grade level. Encompassed were at least documents from grade levels one to eight.</p> <p>60.7 % of the IEPs were written in a table format, 39.3 % in a list format. In most cases (70.5 %), it was unclear whether the IEPs were produced for pupils with identified special educational needs. Of the documents, 24.1 % were related to ‘special educational needs in learning’ (‘sonderpädagogischer Unterstützungsbedarf im Bereich Lernen’). Only 5.4 % of the IEPs were related to other support areas mentioned explicitly.</p> <p><strong>Documents cataloguing</strong></p> <p>Funds supported the study from the ‘Teacher Training and Education Research Centre’ (CeLeB) of Hildesheim University to employ a student research assistant, facilitating data collection and preparation. Nevertheless, the collection is a compilation of documents that was collected with comparatively few resources.</p> <p>No manual for the codes is therefore available.</p> <p>During the data collection, various documents were provided by the schools as IEPs. However, 8 of the 120 documents assembled in the present file were merely cover sheets to groups of IEPs (documents 18, 26, 30, 40, 42 and 43; document No 47 contains only observations and is not a complete IEP).</p> <p>Codes were used on-site in the schools during collection to catalogue the documents.</p> <p>The code consists of four parts: AA-BB-CC-DDD</p> <p>Part AA: indicates the number of the school in the sample.</p> <p>Part BB: indicates the type of school (01=primary; 02=secondary)</p> <p>Part CC: indicates the type of document (01 = IEP; 02 = school internal curriculum: whereby only IEPs are included in this dataset)</p> <p>Part DDD: consecutive page numbers per school included in the sample</p>
German Local Protest News (GLPN) dataset
<p>This dataset contains excerpts from newspaper articles of four German local newspapers labelled for relevancy in protest event analysis.</p> <p>It can be used to train machine learning models to detect news articles containing mentions of protest event for political analysis.</p> <p>For using a model trained on this data, it is recommended to preprocess new data in similar ways like this dataset.</p> <p>To retrieve the excerpts, we the following steps have been taken:</p> <ol> <li>split articles into sentences</li> <li>tag sentences that match the following regular expression: protest_regex = re.compile(r'protest|versamm|demonstr|kundgebung|kampagne|soziale bewegung|hausbesetz|streik|unterschriftensammlung|hasskriminalität|unruhen|aufruhr|aufstand|boykott|riot|aktivis|widerstand|mobilisierung|petition|bürgerinitiative|bürgerbegehren|aufmarsch', re.UNICODE | re.IGNORECASE)</li> <li>tag sentences predecessing or succeeding tagged sentences</li> <li>concatenate all tagged sentences to the excerpt.</li> </ol> <p>See the following code on github for an example:</p> <ul> <li><a href="https://github.com/Leibniz-HBI/protest-event-analysis/blob/main/utils.py">https://github.com/Leibniz-HBI/protest-event-analysis/blob/main/utils.py</a> contains the function "reformat_df" that preprocesses a column named "text" of a given dataframe</li> <li><a href="https://github.com/Leibniz-HBI/protest-event-analysis/blob/main/task-A_prediction.py">https://github.com/Leibniz-HBI/protest-event-analysis/blob/main/task-A_prediction.py</a> contains an example on how to apply a model on new data</li> </ul> <p>Experiments on this dataset are described in the following paper:</p> <p>> Wiedemann, G., Dollbaum, J. M., Haunss, S., Daphi, P., Meier, L. D. (2022): A Generalized Approach to Protest Event Detection in German Local News, In: Proceedings of the 13th International Conference on Language Resources and Evaluation (LREC 2022). Marseille, France. European Language Resources Association (ELRA).</p> <p>In case of questions on the dataset, please contact Gregor Wiedemann at the Leibniz-Institute for Media Research (HBI): g.wiedemann@leibniz-hbi.de</p> <p> </p>
Monthly Samples of German Tweets (2019 - 2022)
<p><strong>Due to size limitations, this dataset is no longer updated. Further data as of January 2013 can be found in this dataset: <a href="https://doi.org/10.5281/zenodo.7670097">https://doi.org/10.5281/zenodo.7670097</a></strong></p> <p>This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API using the following filters:</p> <ul> <li>terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em></li> <li>language: <em>'de'</em></li> </ul> <p>This filter combination should record a 1% sample of (almost) all German tweets (in German it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>).</p> <p>This dataset might be useful for the following use cases:</p> <ul> <li>Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets)</li> <li>Social Network Analysis (Twitter network)</li> <li>Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...)</li> <li>Sharing political (or other domain-specific) content</li> <li>Bot detection</li> <li>and more ...</li> </ul> <p>This dataset will be updated monthly. Each sample (starting in April 2019) will follow the following naming pattern:</p> <ul> <li>german-tweet-sample-<em><YEAR></em>-<em><MONTH></em>.zip (size: ~ 1GB)</li> </ul> <p>It will contain several bunches of recorded JSON gzipped files.</p>
SMILE Swiss German Sign Language Dataset
<p><strong>Description</strong></p> <p>The SMILE Swiss German Sign Language Dataset consists of videos, joint coordinates and annotations of 100 isolated signs of a Swiss German Sign Language (Deutschschweizerische Gebärdensprache, DSGS) vocabulary production test. All items were produced multiple times by 16 adult L1 signers and 22 adult L2 learners of DSGS. Associated linguistic transcriptions and annotations are available for second path data of 10 adult L1 signers and 18 adult L2 learners of DSGS.</p> <p>The dataset has been created in the context of developing an assessment system for lexical signs of DSGS in the SNSF project SMILE.</p> <p>More precisely, for each participant, the following files are available:</p> <ul> <li>Kinect color video (.mp4); 1920x1080 Pixels @ 30 FPS;</li> <li>Kinect Pose Information (.csv); 25 Joints; 3D Joint Coordinates and Angles;</li> <li>OpenPose output (.json); 2D Joint Coordinates and Confidences;</li> <li>iLex annotation files (.xml); linguistic annotations.</li> </ul> <p> </p> <p><strong>Reference</strong></p> <p>If you use this database, please cite the following publication:</p> <p><em>Sarah Ebling, Necati Cihan Camgöz, Penny Boyes Braem, Katja Tissi, Sandra Sidler-Miserez, Stephanie Stoll, Simon Hadfield, Tobias Haug, Richard Bowden, Sandrine Tornay, Marzieh Razavi, and Mathew Magimai-Doss. SMILE Swiss German Sign Language Dataset. In Proceedings of the 11th Language Resources and Evaluation Conference (LREC 2018), pages 4221–4229, 2018.</em></p>
TeCoPhy: A Text Corpus of German Physics Texts
<p>TeCoPhy is a Text Corpus of German Physics Texts. Most of the texts are taken from textbooks at school and university level, but other sources were included as well. The corpus is a collection of sentences. From each book, at most 14% of the text is included in the corpus. The corpus consists of 236,278 sentences with 5,32 Million tokens collected from 223 different sources.</p> <p>The distribution consists of two files: an XML file with metadata on the sources and a file with the sentences taken from those sources.</p>
German Character Recognition Dataset
<p>The dataset contains 282,472 grayscale images, each measuring 40 x 40 pixels, depicting a diverse range of 82 distinct German characters, digits and mathematical symbols.</p> <p>In contrast to the MNIST dataset, where image alignment varies, all the images in this dataset are perfectly aligned. They are centered within a 40 x 40 bounding box, ensuring they touch either the left and right sides or the top and bottom borders. This alignment significantly simplifies the training task, leading to excellent performance metrics.</p> <p>The training and testing data is stored in two separate CSV files. In each file, the first column represents the Unicode character, while the subsequent 1600 values correspond to the grayscale values of the flattened image. If you find any aspect unclear, please refer to our attached code, which offers a comprehensive logic for training a CNN in PyTorch. You can easily select the specific classes on which you intend to train. Notably, when exclusively training on the digits from 0 to 9, we achieved an impressive accuracy and Matthews Correlation Coefficient (MCC) of roughly 99% on the test data.</p>
The TongueSwitcher Corpus of German-English Code-Switching
<p>This is the TongueSwitcher Corpus of German-English code-switching tweets. Included are the train and dev sets with automatic word (and subword for mixed words) language identification, alongside the human-annotated corpus with test and interlingual homograph sets.</p> <p>BibTeX entry and citation info</p> <p>@inproceedings{sterner2023tongueswitcher,<br> author = {Igor Sterner and Simone Teufel},<br> title = {TongueSwitcher: Fine-Grained Identification of German-English Code-Switching},<br> booktitle = {Sixth Workshop on Computational Approaches to Linguistic Code-Switching},<br> publisher = {Empirical Methods in Natural Language Processing},<br> year = {2023},<br>}</p>
Stollen Bread From German for Christmas
Traditional stollen bread from German for Christmas. Source: Objaverse 1.0 / Sketchfab
German-Swiss survey on the use of (digital) language resources
<p>In this document the raw data of an online survey (n=255) on the use of (digital) language resources in German-speaking Switzerland is made available to the public.The corresponding codes are described in more detail in the codebook also published on Zenodo (cf. 10.5281/zenodo.4314348). In addition, the website of the project "Digital Language Resources - Empirical Analyses and Perspectives" (cf. https://www.ds.uzh.ch/de/projekte/digitale-sprachressourcen) offers a detailed presentation of the results elaborated on the basis of this data. The study complements the Europe-wide survey conducted by the European Network of e-Lexicography (ENeL) and was carried out in close cooperation with the Leibniz-Institute for the German Language (IDS) in Mannheim. </p> <p> </p> <p>T</p>
Translation, Cross-Cultural Adaptation and Validation of the Lymphedema Quality of Life Questionnaire (LYMQOL) in German Patients with Lymphedema of the Upper or Lower Limbs
<p>Raw data collected during interviews 1 and 2.</p>
Hate speech and personal attack dataset in German social media
<p>This dataset contains 43735 German tweet ids and corresponding annotations for Hate Speech label and 43734 German tweet ids and corresponding annotations for Personal attack label. The creation of this dataset was part of the project DACHS “A Data-driven Approach to Countering Hate Speech” funded by the Rights, Equality and Citizenship Programme of the European Union.</p>
Monthly Samples of German Tweets (2023)
<blockquote> <p><strong>Important:</strong> Due to data format changes in the Twitter Streaming API, recording will be discontinued in March 2023 until further notice, as the data provided can no longer be provided in the same format.</p> </blockquote> <p>This dataset contains German tweets and Twitter accounts recorded from the public Twitter Streaming API using the following filters:</p> <ul> <li>terms: <em>'a'</em>, <em>'e'</em>, <em>'i'</em>, <em>'o'</em>, <em>'u'</em>, and <em>'n'</em></li> <li>language: <em>'de'</em></li> </ul> <p>This filter combination should record a 1% sample of (almost) all German tweets (in German, it is very unlikely that terms do not contain vowels or the frequently used character <em>'n'</em>).</p> <p>This dataset might be helpful for the following use cases:</p> <ul> <li>Natural language processing (focussing on Twitter specifics in German, there exist only little German datasets)</li> <li>Social Network Analysis (Twitter network)</li> <li>Identifying behavioural patterns (retweeting, quoting, replying, hate speech, ...)</li> <li>Sharing political (or other domain-specific) content</li> <li>Bot detection</li> <li>and more ...</li> </ul> <p>This dataset will be updated monthly. Each sample (starting in January 2023) will follow the following naming pattern:</p> <ul> <li>german-tweet-sample-<em><YEAR></em>-<em><MONTH></em>.zip (size: ~ 1GB)</li> </ul> <p>It will contain several bunches of recorded JSON gzipped files.</p>
Grocery prices in German e-commerce
<p>The dataset comprises price quotes of large online full-assortment German grocery retailers automatically collected from open sources during 09/2019-09/2020.</p> <p>The data is provided "as is" and will not be maintained or updated; by downloading the dataset, you agree to use it only for non-commercial, scholarly purposes. </p> <p> </p>
German Childhood Cancer Registry
Currently 81,323 cases are registered, and annually another 2,300 are reported. Almost 48,400 of these are in active longterm surveillance. Below we report the average annual cases among children aged 0 to 17 years between 2015 and 2024.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.