Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
724
datasets available to search
ShareScore release 0.9.0
Dataset results
724 results for “german”
In 2017, Plantix, a free smartphone app that helps identify plant damage, was introduced to the Indian state of Andhra Pradesh, with an extension partner. Plantix was created by Progressive Environmental and Agricultural Technologies (PEAT), a German startup. Two PEAT cofounders, Charlotte Schuman (second from the right) and Alex Kennepohl (center, with eyeglasses), confer about the smartphone app with students from Angrau University. Farmers and gardeners can transmit their plant images to Plantix, which uses deep learning and computer vision to help identify diseases and pests. The smartphone app offers symptom descriptions, treatment recommendations, and potential preventive actions. Photographs: Courtesy of PEAT GmbH. in Deep learning brings speed, accuracy to the life sciences.
In 2017, Plantix, a free smartphone app that helps identify plant damage, was introduced to the Indian state of Andhra Pradesh, with an extension partner. Plantix was created by Progressive Environmental and Agricultural Technologies (PEAT), a German startup. Two PEAT cofounders, Charlotte Schuman (second from the right) and Alex Kennepohl (center, with eyeglasses), confer about the smartphone app with students from Angrau University. Farmers and gardeners can transmit their plant images to Plantix, which uses deep learning and computer vision to help identify diseases and pests. The smartphone app offers symptom descriptions, treatment recommendations, and potential preventive actions. Photographs: Courtesy of PEAT GmbH.
MFD-2-TM-de: German Moral Foundation Dictionary 2 by Translation Matrix
<p>This dataset was created for the research paper: "Reiter-Haas, M., Kopeinik, S., & Lex, E. (2021, May). Studying moral-based differences in the framing of political tweets. In <em>Proceedings of the International AAAI Conference on Web and Social Media</em> (Vol. 15, pp. 1085-1089).". Available at <a href="https://ojs.aaai.org/index.php/ICWSM/article/view/18135/17938">https://ojs.aaai.org/index.php/ICWSM/article/view/18135/17938</a></p> <p>The translation were created using a translation matrix of positive and negative valence words. The dictionary was used to measure effects in the embedding space for which it is suitable. However, it is to note that the dictionary does not work in a classical sense, as it contains duplicates (e.g., words of the opposite poles). While this was also the case in the original dictionary (e.g., for "wound"), the effect has become much more pronounced in the translated version.</p> <p>Format:</p> <ul> <li>Lines starting with % denote a new section that the following words belong to (e.g., care.vice for harm).</li> <li>Otherwise, one word per line.</li> </ul> <p>For implementation details, refer to the <a href="https://github.com/socialcomplab/icwsm21-framing">GitHub Repo</a></p> <p> </p>
Alphabetical comprehensive false friends lists and top 100 frequency lists in German-English & Chinese-Japanese
<p>The FF pairs in German and English were sourced from "False Friends: A Short Dictionary: Reclam premium Sprachtraining" by Burkhard Dretzke and Margaret Nester. I systematically compiled a list of all 794 pairs of FFs from the German-English dictionary, initially arranging them alphabetically in spreadsheets. Subsequently, I queried the frequency of each entry in both German and English corpora, sorting the word lists in each language in descending order based on frequency per million. I aimed to compile a list of the most relevant FFs by selecting those pairs present in the top 100 entries of both lists, resulting in 34 pairs that appeared prominently in both languages. Additionally, I collected data on FFs where one language exhibited high frequency although the other showed low or minimal occurrence. Excluding the aforementioned 34 pairs, I further filtered the German and English frequency lists to identify 33 words each, along with their corresponding FFs in the other language, yielding a total of 66 additional FF pairs. Thus, I compiled a comprehensive list of the top 100 FFs in German and English based on combined frequency of occurrence. Meanwhile, the Chinese and Japanese FFs were gathered from "2136 Japanese Kanji Character Dictionary" and published by Liaoning People's Publishing House. The methodology for collecting data on Chinese-Japanese FF pairs follows a similar procedure but with greater difficulty and complexity. Given that the concept and phenomenon of FFs originate from European linguistics, there is a relative scarcity of corresponding academic materials in Asian languages such as Chinese and Japanese, for instance, a readily available FFs' dictionary. Consequently, I used a Japanese kanji dictionary as my primary data source. This dictionary allows for the retrieval of Japanese kanji words using 2136 commonly used Chinese characters, selected by the Japanese Ministry of Culture for their high frequency of usage in social life and officially announced in November 2010. With nearly 15,000 entries, each word in the dictionary is accompanied by detailed Chinese interpretations and examples, facilitating the differentiation between true cognates and FFs with the same form. Following a meticulous examination of almost 15,000 entries, I identified 700 pairs of Chinese-Japanese FFs. Then I queried the frequency of each entry in both Chinese and Japanese corpora, sorting the word lists in each language in descending order based on frequencies. Subsequently, I selected 48 pairs that appeared in the top 100 frequency list in both Chinese and Japanese, and 26 words for Chinese and 26 words for Japanese on their own, akin to the methodology employed for German-English data collection, ultimately compiling a comprehensive list of the top 100 FFs in Chinese and Japanese based on combined frequency of occurrence.</p>
Supplementary material for the article: 'Kein' subjects are hard: Exploring German-speaking children's behavior with negative indefinites
<div> <p>This data set includes the supplementary material for the article "Kein subjects are hard: Exploring German-speaking children’s behavior with negative indefinites". All content is documented in the README.md file.</p> <p>Abstract: In this paper, we present converging results from three studies investigating children’s production or comprehension of the negative indefinite <em>kein</em> in German. An elicited production study found that 3- to 6-year-old children and adults exhibit different patterns with respect to the production of <em>kein</em>: children, but not adults, exhibit an asymmetry with respect to the position where they produce negative indefinites, in that they use negative indefinites more frequently in object than in subject position. A corpus study investigating spontaneous speech replicated this asymmetry for children, but this time found it also present for adults. Finally, the asymmetry is corroborated by a comprehension study indicating a processing cost for negative indefinite subjects, relative to negative indefinite objects. We argue that these patterns are most straightforwardly captured by an explanation that assumes the decomposition approach to the German negative indefinite <em>kein</em>: rather than a single semantic unit (i.e., negative quantifier), <em>kein</em> is decomposed into a negative part and an indefinite part.</p> </div>
Modelling input data for the case study of the paper "Uncertainty-Based Market-Clearing Models: A Comparative Analysis of the Dutch, French, and German Markets".
<p>This data package includes the modelling input data to replicate the results of the case study included in the paper "Uncertainty-Based Market-Clearing Models: A Comparative<br>Analysis of the Dutch, French, and German Markets". </p> <p>The case study models the Dutch, French and German day-ahead electricity markets, in which the existing capacities of electricity generation and upward- and downward reserve capacities are considered, in addition to 105 wind output realization scenarios for each simulation day. A detailed description of the case study is provided in the readme file.</p> <p>This supplementary data package includes the following files:</p> <p>- Meta Data – Netherlands.xlsx: Dataset containing the meta data for the Dutch case study</p> <p>- Meta Data – France.xlsx: Dataset containing the meta data for the French case study</p> <p>- Meta Data – Germany.xlsx: Dataset containing the meta data for the German case study</p> <p>- Readme.txt: Includes a detailed description of the data packages</p>
Modeling Languages for Digital Twins - A Survey Among the German Automotive Industry
<p>This repository contains the replication package for the paper _Modeling Languages for Digital Twins: A Survey Among the German Automotive Industry_ by Jérôme Pfeiffer, Dominik Fuchß, Thomas Kühn, Robin Liebhart, Dirk Neumann, Christer Neimöck, Christian Seiler, Anne Koziolek, and Andreas Wortmann. <br>The paper has been submitted to the practice track of [MODELS 2024](https://conf.researchr.org/track/models-2024/models-2024-technical-track#Practice-Track).</p> <h3>Data</h3> <p>This replication package contains all information from the survey:<br>- `results.csv`: A csv version of all data exported from LimeSurvey (German). Personal information from the participants has been removed. This file can be imported to reproduce the extraction results described in our paper.<br>- `survey_german.pdf`: The pdf version of the original survey in German. <br>- `survey_german.md`: A markdown version of the original survey in German. <br>- `survey_english.md`: A markdown version of the survey translated into English. </p> <h3>Selection of participants and distribution</h3> <p>With both versions, the survey can be executed again with a different target audience in English or German. In our case we wanted to reach as much participants from diverse work areas as possible, where we invited the participants by email via an internal mailing list of 189 members of the SofDCar project. To improve the response rate, we implemented two deadline extensions from the initial one-month-long time frame with 2 weeks of additional response time. Together with the deadline extension, we sent a mail to inform and remind the members of the consortium of the survey.</p> <h3>Data extraction</h3> <p>In total, we had 96 participants, of which 43 completed the questionnaire. For incomplete survey responses, we took only the available answers and did not include the missing answers in our data analysis. For data analysis we utilized the commercial Tool IBM SPSS and custom python scripts.</p> <h2>Research Questions </h2> <p>- RQ1: How is the DT understood in the automotive industry?<br> - RQ1.1: For which phases of automotive development are DTs<br>important?<br> - RQ1.2: What are desired properties of DTs?<br> - RQ1.3: What are desired purposes of using DTs?<br> - RQ1.4: How do these purposes change in relation to different phases of automotive development?<br>- RQ2: Which modeling languages and modeling tools are currently employed in the automotive industry?<br> - RQ2.1: Which kinds of models are important during automotive development?<br> - RQ2.2: How important are which models in the phases of automotive development?<br> - RQ2.3: Which tools are used to create and maintain these models?</p>
Job satisfaction and performance orientation in German emergency medical service – a nationwide survey
<p>In a nationwide German cross-sectional questionnaire survey, we collected data on job satisfaction an perfomance orientation of paramedics in Germany.</p> <p>This file contains the underlying data set as a result of the online questionnaire.</p>
Data on concept 'Team of host country' in German, English and Russian
<p>This dataset contains the corpus of German, English and Russian utterances in which the concept <em>team of host country</em> is expressed. The corpus was collected during the 2019 IIHF Ice Hockey World Championship from the articles devoted to this tournament which were published on the webpages of German, English and Russian mass media.</p>
Fig. 2 Morphological characteristics. a Double broken infraorbital ridge IOR. b in Molecular identification and morphological characteristics of native and invasive Asian brush-clawed crabs (Crustacea: Brachyura) from Japanese and German coasts: Hemigrapsus penicillatus (De Haan, 1835) versus Hemigrapsus takanoi Asakura & Watanabe 2005
Fig. 2 Morphological characteristics. a Double broken infraorbital ridge IOR. b Symmetrically arranged spots on the dorsal face. c Male pleon fold down, first pleopods facing distal/ lateral (setae of the head around a scoop-like nose appears as dark area at the end of the pleopod). d Soft hair near the movable finger of a female chela (HtY2)
HOCON34k: A Corpus of Hate speech in Online Comments from German Newspapers
<p>We have compiled a dataset containing 34,223 comments in German, authored by users from online-platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by a group of 29 volunteers, using a binary annotation approach. The inter-rater reliability for hate speech is 0.4428 across all annotators and increases to 0.6078 when considering an optimized subset of 12 annotators, as measured by Fleiss’ Kappa. Additionally, we present a baseline text classification using BERT, achieving an MCC-score up to 0.32 and an F2-score up to 0.64 in our initial experiment on this new corpus. The data set, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.</p>
The (digital) Wellbeing of German Smartphone Users_ Muench & Carolus
Open the record for dataset details and reuse information.
Generation-based and consumption-based grid emission intensity time series for German federal states 01/2020 - 07/2024
<p>The dataset contains time series estimating the consumption mix, generation-based and consumption-based grid emission intensity of German federal states for the period January 2020 to end of June 2024 as shown on the <a href="https://co2map.de">CO2Map</a> website. For an explanation of the underlying methodology see the <a href="https://co2map.de/methodology.html">CO2Map Methology</a> section, a detailed scientific article currently is in preparation.</p> <p>All time series have an hourly resolution (time stamps are given in UTC) and contain values for the 13 federal territorial states in Germany, with the city states merged into them:<br><br>BW: Baden-Wuerttemberg<br>BY: Bavaria<br>BB: Brandenburg and Berlin<br>HE: Hesse<br>MV: Mecklenburg-Western Pomerania<br>NI: Lower Saxony and Bremen<br>NW: North Rhine-Westphalia<br>RP: Rhineland-Palatinate<br>SL: Saarland<br>SN: Saxony<br>ST: Saxony-Anhalt<br>SH: Schleswig-Holstein and Hamburg<br>TH: Thuringia</p> <p><strong>"Emission intensity"</strong><br>This folder contains the following files for 2020 to 2023:<br>"region_consumption_based_emission_intensity.csv": Consumption-based emission intensity in kgCO2/kWh<br>"region_generation_based_emission_intensity.csv": Generation-based emission intensity in kgCO2/kWh</p> <p><strong>"Import tables"</strong><br>The folder contains the raw data representing the hourly consumption and storage mix including imports for federal states in MWh. This data allows to calculate generation-based and consumption-based emission intensity time series using arbitrary technology-specific emission factors for the technology types Biomass, Gas, Hard coal, Hydro, Lignite, Nuclear, Other fossil, Other renewable, Solar, Storage, Wind Offshore, Wind Onshore. Each file contains time series for one region as given above and for one year. The columns represent the hourly amount of load/storage/load+storage originating from a certain region and technology type, with the regions including local generation (i.e. the same federal state), imports from other federal states in Germany, and imports from other European countries (the origin of the imports is determined using a flow-tracing algorithm, see the <a href="https://co2map.de/methodology.html">CO2Map Methology</a> section).</p> <p>For each year and each federal state there is one file for load (representing electricity consumption), one file for storage (representing storage charging) and one file for total (sum of load and storage). The file "2023_import_table_load_DE_BW.csv", for instance, contains the mix of hourly load in Baden-Wuerttemberg. The first entry 01.01.2023 00:00 (UTC) then contains the mix for the electricity consumption in the first hour (UTC) of the year 2023 in BW. The sum over all columns for this row yields 4667.80 MWh, which is the approximated total load in this hour in BW (based on a regionalization method). The largest entry in the row is from "DE_NI,HB / Wind Onshore" (614.57 MWh), which is the amount of onshore wind power generated in Lower Saxony and Bremen and consumed in Baden-Wuerttemberg (based on the flow-tracing method). The second largest entry is from "DE_BW / Wind Onshore", which is the amount of onshore wind power generated and consumed locally in Baden-Wuerttemberg. The largest import to BW from other countries in this hour has been 115.18 MWh of nuclear power from Belgium. The file "2023_import_table_load_DE_BW.csv" contains an aggregated version of the same file, with the originating regions corresponding to the same federal state, imports from other federal states in Germany, and imports from other countries (each separated per technology).</p> <p><strong>"Export tables"<br></strong>The folder "Export tables" is similar to the one with "Import tables", but gives data about the location where generation from German regions is used. It contains the raw data representing the hourly generation mix associated with the location where it is used (locally and in other regions). The columns represent the hourly amount of load/storage/load+storage originating from the given region. </p> <p>For each year and each federal state there is one file for load (representing electricity consumption), one file for storage (representing storage charging) and one file for total (sum of load and storage). The file "2023_export_table_total_DE_SH,HH.csv", for instance, contains the mix of hourly generation in Schleswig-Holstein and Hamburg which is associated with electricity consumption and storage charging both locally and abroad. The first entry 01.01.2023 00:00 (UTC) then contains the amount of generation from SH,HH used in each region in the first hour (UTC) of the year 2023. The sum over all columns for this row yields 5730.40 MWh, which is the approximated total generation in this hour in SH,HH (based on a regionalization method). Since SH,HH is a net exporter in this hour, the sum over all columns associated with SH,HH represents the local load+storage charging (1842.40 MWh). The sum over all columns with "Wind Onshore" yields the generation from onshore wind in SH,HH in this hour (4318.05 MWh) (again based on a regionalization method). The different entries for "Wind Onshore" then tell where this onshore wind is used (load and storage charging), with the main contributions locally in SH,HH (1388.33 MWh), and exports to Norway (1582.03 MWh) and UK (682.67 MWh). The destination of these exports is estimated based on the flow tracing method in the same way as the origin of imports are determined.</p> <p>For the data sources and methods used to calculate the time series contained in this dataset see the CO2Map <a href="https://co2map.de/methodology.html">Methodology </a>and <a href="https://co2map.de/about.html">About </a>sections. </p> <p>Please communicate any questions, corrections or comments to mirko.schaefer [at] inatech.uni-freiburg.de</p> <p>All authors acknowledge funding from Elektrizitätswerke Schönau (EWS) through the Sonnencent program, ID 00009280. Tim Fürmann acknowledges funding from DFG (SPP 1984), project ID 450860949. Ramiz Qussous and Robin L. Grether would like to thank the German Federal Government, the German State Governments, and the Joint Science Conference (GWK) for their funding and support as part of the NFDI4Energy consortium, funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – 501865131.</p>
Olfactory Learning Supports an Adaptive Sugar-Aversion Gustatory Phenotype in the German Cockroach
<p>An association of food sources with odors prominently guides foraging behavior in animals. To understand the interaction of olfactory memory and food preferences, we used glucose-averse (GA) German cockroaches. Multiple populations of cockroaches evolved a gustatory polymor-phism where glucose is perceived as a deterrent and enables GA cockroaches to avoid eating glucose-containing toxic baits. Comparative behavioral analysis using an operant conditioning paradigm revealed that learning and memory guide foraging decisions. Cockroaches learned to associate specific food odors with fructose (phagostimulant, reward) within only 1 hr of condi-tioning session, and with caffeine (deterrent, punishment) after only three 1 hr conditioning ses-sions. Glucose acted as reward in wild type (WT) cockroaches, but GA cockroaches learned to avoid an innately attractive odor that was associated with glucose. Olfactory memory was retained for at least 3 days after three 1 hr conditioning sessions. Our results reveal that specific tastants can serve as potent reward or punishment in olfactory associative learning, which reinforces gustatory food preferences. Olfactory learning therefore reinforces behavioral resistance of GA cockroaches to sugar-containing toxic baits. Cockroaches may also generalize their olfactory learning to baits that contain the same or similar attractive odors even if they do not contain glucose.</p>
German nobles in Paris, 1854
<p>The 1851 census showed the presence of 12,245 Germans in Paris, the vast majority of whom were craftsmen. The "Adressbuch 1854" project, carried out at the German Historical Institute between 2002 and 2006, made use of a work identified in the holdings of the library of the city of Paris. Some 4,700 addresses were listed in a directory that can still be consulted in the archives. The project consisted of integrating the entries in the book into a database that could be consulted online.<br> The dataset focuses on the noble persons who were listed in the book. Their occupation, location and rank. For some, the legion of honor is indicated.</p>
PoeTree. Poetry Corpora in Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Russian, Slovenian, and Spanish
<p>PoeTree is a dataset comprising nearly 335,000 poems / 90,000,000 tokens in 11 languages (Czech, English, French, German, Hungarian, Italian, Norwegian, Portuguese, Spanish, Slovenian, and Russian). Each corpus has been deduplicated, enriched with Universal Dependencies, provided with additional metadata and converted into a unified JSON structure (schema available at <a href="https://versologie.cz/poetree/json-schema">https://versologie.cz/poetree/json-schema</a>).</p> <ul> <li>cs (~80k poems) <ul> <li>derived from <a href="https://github.com/versotym/corpusCzechVerse" target="_blank" rel="noopener">Corpus of Czech Verse</a></li> </ul> </li> <li>de (~74k poems) <ul> <li>derived from <a href="https://metricalizer.de/" target="_blank" rel="noopener">Metricalizer</a> and <a href="https://github.com/tnhaider/DLK" target="_blank" rel="noopener">Deutsches Lyrik Korpus</a></li> </ul> </li> <li>en (~40k poems) <ul> <li>based on texts from <a href="https://gutenberg.org/" target="_blank" rel="noopener">Project Gutenberg</a></li> </ul> </li> <li>es (~9k poems) <ul> <li>derived from <a href="https://github.com/bncolorado/CorpusSonetosSigloDeOro" target="_blank" rel="noopener">Corpus of Spanish Golden-Age Sonnets</a> and <a href="https://github.com/pruizf/disco" target="_blank" rel="noopener">Diachronic Spanish Sonnet Corpus</a></li> </ul> </li> <li>fr (~18k poems) <ul> <li>derived from <a href="https://crisco4.unicaen.fr/verlaine" target="_blank" rel="noopener">Malherbə</a></li> </ul> </li> <li>hu (~13k poems) <ul> <li>derived from <a href="https://github.com/ELTE-DH/poetry-corpus/" target="_blank" rel="noopener">ELTE Poetry Corpus</a></li> </ul> </li> <li>it (~40k poems) <ul> <li>derived from <a href="http://www.bibliotecaitaliana.it/" target="_blank" rel="noopener">Biblioteca Italiana</a></li> </ul> </li> <li>no (~3k poems) <ul> <li>derived from <a href="https://github.com/norn-uio/norn-poems" target="_blank" rel="noopener">NORN Poems</a></li> </ul> </li> <li>pt (~5k poems) <ul> <li>derived from <a href="https://github.com/adiel-mittmann/poemas" target="_blank" rel="noopener">Poemas</a></li> </ul> </li> <li>ru (~45k poems) <ul> <li>derived from <a href="https://ruscorpora.ru/en/" target="_blank" rel="noopener">Corpus of Russian Poetry</a></li> </ul> </li> <li>sl (~5k poem)<br> <ul> <li>based on texts from <a href="https://en.wikisource.org/" target="_blank" rel="noopener">wikisource</a></li> </ul> </li> </ul> <p><em>new in v. 1.0.0:</em></p> <ul> <li><em>PoeTree.no added</em></li> <li><em>PoeTree.(cs,de,en,fr,hu,it,ru,sl) enriched with geolocation mentions</em></li> <li><em>Updated and corrected metadata in PoeTree.(de,en,es,ru)</em></li> <li><em>Multiple text corrections in PoeTree.ru</em></li> </ul>
GiCCS: A German in-Context Conversational Similarity Benchmark
<p>We introduce GiCCS, a first conversational STS evaluation benchmark for German. We collected the similarity annotations for GiCCS using best-worst scaling and presenting the target items in context, in order to obtain highly-reliable context-dependent similarity scores. In our paper, we present benchmarking experiments for evaluating LMs on capturing the similarity of utterances. Results<br> suggest that pretraining LMs on conversational data and providing conversational context can be useful for capturing similarity of utterances in dialogues. GiCCS will be publicly available to encourage benchmarking of conversational LMs.</p>
MHG4SNA: Middle high german texts annotated for social network analysis
<p><strong>Description</strong></p> <p>This corpus contains multiple middle high german texts with annotations for social network analysis. It contains annotations of: named entities and entity mentions (including partial coreference resolution), direct speech, narrator's comments. See below for further description.</p> <p>The annotated texts are part of my dissertation on social network analysis of arthurian romances. The research was developed in the context of the DH center <a href="https://www.creta.uni-stuttgart.de/">CRETA</a> at the University of Stuttgart.</p> <p> </p> <p><strong>Texts</strong></p> <ul> <li>Wolfram von Eschenbach: 'Parzival', in: Wolfram von Eschenbach: Werke, ed. by Karl Lachmann, 5th edition, Berlin 1891, pp. 11–388.</li> <li>Hartmann von Aue: 'Erec', ed. by Albert Leitzmann continued by Ludwig Wolff, 7th edition by Kurt Gärtner, Tübingen 2006 (Altdeutsche Textbibliothek 39).</li> <li>Hartmann von Aue: 'Iwein', ed. by G. F. Benecke and K. Lachmann, revised by Ludwig Wolff, 7th edition, part 1: Text, Berlin 1968.</li> <li>Wolfram von Eschenbach: 'Willehalm', in: Wolfram von Eschenbach: Werke, ed. by Karl Lachmann, 5th edition, Berlin 1891, pp. 421–640.</li> <li>'Das Rolandslied des Pfaffen Konrad', ed. by Carl Wesle, 3rd edition by Peter Wapnewski, Tübingen 1985 (Altdeutsche Textbibliothek 69).</li> </ul> <p>All texts are part of the MHDBDB (<a href="http://mhdbdb.sbg.ac.at/">Mittelhochdeutsche Begriffsdatenbank</a>).</p> <p> </p> <p><strong>Annotations</strong></p> <p>The texts contain annotations of different categories, as described in the following sections.</p> <p> 1.<em> Named Entities and Entity Mentions</em></p> <p>I annotated all namend entities and entity mentions that belong to the categories PER and LOC. PER stands for 'person' and refers to real persons as well as fictional characters. LOC stands for 'location' and includes real and fictional places.</p> <p>I annotated named entities (e.g. 'Parzival' as PER or 'Nantes' as LOC) as well as entity mentions referring to an instance of PER or LOC (e.g. 'the knight' for Parzival, or 'the city' for Nantes). I did not annotate pronouns. Entity references can contain multiple words, e.g. 'the lovely queen Ginover', and they can be nested, e.g. '[the son of [the king Gahmuret]]'.</p> <p>The annotations follow the guidelines created for multiple categories and disciplines in the context of CRETA. They are published <a href="https://www.creta.uni-stuttgart.de/cute/datenmaterial/annotationsrichtlinien-1-1/index.html">here</a>. </p> <p> 2.<em> Entity Grounding</em></p> <p>All annotated entity references are mapped to the entity instance that they refer to. E.g. the refences 'Parzival', 'Herzeloyde's son', 'the young man', 'the red knight' etc. all refer to the character instance 'Parzival'. The entity grounding takes into consideration the context of the entity mentions since one and the same expression can refer to different instances (in one context 'the king' refers to Arthur, in another context to Gahmuret)</p> <p> 3.<em> Direct Speech (DS)</em></p> <p>Passages of direct speech have been annotated by detecting quotation marks. They are tagged as 'DS'. There are a few cases of embedded direct speech (passages of direct speech containing another passage of direct speech); these cases are annotated as well.</p> <p> 4. <em>Narrator's comments (EK)</em></p> <p>As additional category I annotated passages that contain statements of the narrator, narrator's comments, extensive descriptions or digressions (e.g. an excursus to a specific topic). These passages are not part of the fictional world or lead to a pause in the timeline of events. The are annotated as 'EK' ('EK': passages that aren't part of the diegesis, 'EK2': passages that lead to a pause, e.g. comments or descriptions).</p> <p> 5. <em>Segmentation</em></p> <p>The texts are subdivided in passages of 30 verses. Since some text's editions ('Parzival', 'Willehalm') contain a formal segmentation in passages of 30 verses each, the same kind of segmentation has been transfered to the other texts. This means 'segment 1' contains the first 30 verses, 'segment 2' contains verses 31-60 and so on.</p> <p>According to the editions by Lachmann, 'Parzival' and 'Willehalm' are also subdivided in chapter-like books (Parzival: book 1 to 16, Willehalm: book 1 to 9). The other texts are similarly subdivided in chapter-like sections following common content-based divisions.</p> <p> </p> <p><strong>Social Network Analysis</strong></p> <p>The data can be used to explore and analyse the social network of the texts. SNA can be performed via gephi [4] using the gefx files.</p> <p>The social network is based on co-occurrences using a) the annotated and grounded entities, and b) the text segmentation in segments of 30 verses each. A relation between two or more entities is extracted whenever they co-occur in a segment.</p> <p> </p> <p><strong>Data downloads</strong></p> <p>The annotated texts can be downloaded in multiple formats: conll, csv, and gexf.</p> <p> 1. <em>Conll </em></p> <p>The files contain seven columns:</p> <ul> <li>(1) token,</li> <li>(2) POS-tag, tagged using a <a href="https://www.ims.uni-stuttgart.de/forschung/ressourcen/werkzeuge/pos-tag-mhg/">middle high german pos tagger</a>,</li> <li>(3) number of segment,</li> <li>(4) Entity reference annotation indicating the intance that the entity reference refers to. '-' if there is no entity reference,</li> <li>(5) EK: '1' in case there is an annotation of 'EK', '0' if not,</li> <li>(6) EK2: '1' in case there is an annotation of 'EK2', '0' if not,</li> <li>(7) DS: '1' if the token is tagged as direct speech, '0' if not.</li> </ul> <p> 2. <em>Csv</em></p> <p>The csv files contain all annotations of the category PER including entity grounding.</p> <p>The files contain the following columns:</p> <ul> <li>begin and end (start and end of the entity reference expression, character offset),</li> <li>doc_id (document id),</li> <li>buch (book number),</li> <li>quote (entity reference expression),</li> <li>coref (the entity instance that the expression refers to),</li> <li>overlap (indicates if there is an overlap, relevant for embedded entities),</li> <li>ek and ek2 (narrator's comment),</li> <li>ds (direct speech),</li> <li>space (annotations of the space where the story takes action, can be ignored here),</li> <li>segnr (number of segment),</li> <li>em (embedded),</li> <li>klasse (entity class),</li> <li>xrange (technical, relevant for annotation view).</li> </ul> <p> 3. <em>Gexf</em></p> <p>These files can be used to import the data to gephi. It is based on the annotation and grounding of entities (categorie PER). A relation between entities is based on co-occurrence (whenever two or more entities co-occur in a segment, they have a relation; with more relations, the intensitiy of their relation grows). The text segmentation is described above.</p> <p>Embedded entities are excluded. Entities mentioned in direct speech (DS) or in comments (EK) can optionally be selected or deselected. These optional filters are indicated in the name of the files.</p> <p>To visualize the graph dynamically, one can use the text segmentation as timeline.</p> <p> </p> <p>release v1.0.0: data publication in the context of my dissertation. </p>
ChatGPT - Questions and Answers to the Political Coordinates Test in English, French, Italian and German
<p>This dataset contains the questions and answers resulting from the administration of the policy coordination test in English, French, Italian and German to ChatGPT. The version used was GPT-4.</p>
RescueSpeech: A German Corpus for Speech Recognition in Search and Rescue Domain
<p>Dear User,</p> <p>We are thrilled to introduce our latest release - the <strong>RescueSpeech</strong> audio dataset, comprising authentic German speech recordings obtained from simulated search and rescue (SAR) exercises. The dataset contains manually annotated recordings from native German speakers, which were initially captured at 44.1 kHz and later down-sampled to 16 kHz to obtain a set of mono-speaker-single channel audio recordings. In order to protect the identity of the speakers, their names have been anonymized.</p> <p>The RescueSpeech dataset is divided into two sets, each designed for different tasks: Automatic Speech Recognition (ASR) and Speech Enhancement.</p> <p>1. For the ASR task, the dataset spans a duration of 1 hour and 36 minutes. It comprises a collection of clean-noisy pairs, where the noisy utterances are created by introducing contaminations from five different noise types sourced from the AudioSet dataset. These noise types include emergency vehicle siren, breathing, engine, chopper, and static radio noise. To match the 2412 clean utterances in the dataset, we have synthesized an equal number of corresponding noisy utterances. Additionally, we have provided the noise waveform files used to create the noisy utterances, ensuring transparency and reproducibility in the research community.</p> <p>2. The Speech Enhancement task dataset is larger in size compared to the ASR dataset. The primary objective of this dataset is to facilitate the fine-tuning of speech enhancement models, particularly for the five SAR noise types mentioned earlier: emergency vehicle siren, breathing, engine, chopper, and static radio noise. Given the limited duration of clean audio available (1 hour and 36 minutes), we have synthesized multiple noisy utterances with varying noise types and signal-to-noise ratio (SNR) levels, all derived from a single clean utterance. This augmentation approach allows us to generate a more extensive dataset for speech enhancement purposes while preserving the original speaker distribution.</p> <p>By providing these diverse datasets, we aim to support advancements in ASR and Speech Enhancement research, enabling the development and evaluation of robust systems that can handle real-world scenarios encountered during search and rescue operations.<br> </p>
Raw data related to "Tröndle et al (2023): Public preferences for phasing-out fossil fuels in the German building and transport sectors"
<p>Raw survey data related to "Tröndle et al (2023): Public preferences for phasing-out fossil fuels in the German building and transport sectors".</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.