Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
70
datasets available to search
ShareScore release 0.9.0
Dataset results
70 results for “language analysis”
A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)
<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie "Forrest Gump" (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation) and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>
Phlorest phylogeny derived from Birchall et al. 2016 'A combined comparative and phylogenetic analysis of the Chapacuran language family'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall, Joshua, Michael Dunn, and Simon J. Greenhill. 2016. A combined comparative and phylogenetic analysis of the Chapacuran language family. International Journal of American Linguistics 82 (3): 255–84. doi: 10.1086/687383</p> </blockquote>
Phlorest phylogeny derived from Kitchen et al. 2009 'Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Kitchen A, Ehret C, Assefa S & Mulligan CJ. 2009. Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Proceedings of the Royal Society B: Biological Sciences, 270(1668), 2703-2710.</p> </blockquote>
Phlorest phylogeny derived from Lee & Hasegawa 2011 'Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages'
<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee S, Hasegawa T (2011) Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proceedings of the Royal Society B: Biological Sciences, 278(1725):3662–9.</p> </blockquote>
Data associated with the article 'Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis'
<p>The efficacy of oral language comprehension interventions varies, but the reasons for this variation have received little attention. A meta-analysis was conducted to examine intervention factors associated with the efficacy (as expressed with effect sizes) of oral language comprehension interventions in children under the age of 18 with or at risk for (Developmental) Language Disorder, (D)LD.</p> <p>The meta-analysis article together with this additional material comprise the content needed for a thorough understanding and replication of the results.</p> <p>This dataset is based on two systematic scoping reviews on oral language comprehension interventions (Tarvainen et al., 2020, 2021). Further information from the sourced articles was extracted for this study titled ‘Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis’. </p> <p>In the future, we hope that this data is used with a growing body of oral language comprehension interventions to conduct further and more detailed examinations of intervention factors associated with efficacy.</p> <p>References:</p> <p>Tarvainen, S., Launonen, K., & Stolt, S. (2021). Oral language comprehension interventions in school-age children and adolescents with developmental language disorder: A systematic scoping review. <em>Autism & Developmental Language Impairments</em>, <em>6</em>, 1–24. https://doi.org/10.1177/23969415211010423</p> <p>Tarvainen, S., Stolt, S., & Launonen, K. (2020). Oral language comprehension interventions in 1–8-year-old children with language disorders or difficulties: A systematic scoping review. <em>Autism & Developmental Language Impairments</em>, <em>5</em>, 1–24. https://doi.org/10.1177/2396941520946</p> <p> </p>
CLDF dataset derived from Birchall et al.'s "A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family" from 2016
<p>Cite the source of the dataset as:</p> <blockquote> <p>Birchall J, Dunn M, & Greenhill SJ. 2016. A Combined Comparative and Phylogenetic Analysis of the Chapacuran Language Family. International Journal of American Linguistics 82(3). 255–284.</p> </blockquote>
CLDF dataset derived from Lee and Hasegawa's "Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages" from 2011
<p>Cite the source of the dataset as:</p> <blockquote> <p>Lee, Sean and Hasegawa, Toshikazu (2011). Bayesian phylogenetic analysis supports an agricultural origin of Japonic languages. Proceedings of the Royal Society B: Biological Sciences, 278(1725), 3662–3669. doi:10.1098/rspb.2011.0518.</p> </blockquote>
Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis
<p>This repository contain datasets and results for the paper:</p> <p><strong>Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis</strong></p> <p> </p> <p><strong>Github repository for the code: </strong></p> <p><a href="https://github.com/siebeniris/QuantifyingLanguageConfusion/tree/main">Quantifying Language Confusion GitHub repo</a></p> <p> </p> <p><strong>DATA</strong> include the following datasets:</p> <p>i) raw language graphs and</p> <p>ii) the calculated language similarities from the language graphs,</p> <p>iii) <strong>MTEI</strong>: the files from the <a href="https://github.com/siebeniris/vec2text_exp/tree/aaai">experimental results of multilingual inversion attacks</a>, and calculated language confusion entropy from the data;</p> <p>iv) <strong>LCB</strong>: the files from the <a href="https://github.com/for-ai/language-confusion?tab=Apache-2.0-1-ov-file#readme">language confusion benchmark</a> and calculated language confusion entropy from the data </p> <p> </p> <p><strong>Results</strong> include aggregated results for further analysis:</p> <p>i) <strong>inversion_language_confusion</strong>: results from MTEI</p> <p>ii) <strong>prompting_language_confusion</strong>: results from LCB</p> <p> </p> <p> </p>
Supplementary Materials to "Subgrouping in a `dialect continuum': A Bayesian phylogenetic analysis of the Mixtecan language family"
<p>SM0: metadata on the languages of the sample</p> <p>SM 1: custom word list</p> <p>SM2: prose explanation of cognate coding and IPA conversion</p> <p>SM3: annotated cognate sets</p> <p>SM4: nexus files of the broad and fine grained cognate coding</p> <p>SM5: NeighborNet visualization with coloring by Josserand (1983)'s groupings and by groupings from our analysis</p> <p>SM6: BEAST2 xml files</p> <p>SM7: MCC trees from BEAST2 analysis</p> <p>SM8: DensiTree visualization and visualization of full MCC tree of best performing model</p>
Domain-Specific Language domain analysis and evaluation: a systematic literature review
<p>In order to successfully implement Domain-Specific Languages (DSLs), it is needed to systematically define and to support its development process; namely its Evaluation and the Domain Analysis phase. For that purpose, the studies were systematically selected from the most relevant venues that focus on the implementation of DSLs, in order to get insight if and how these development phases were performed. The special focus was given to the human-machine DSLs (excluding the machine-machine languages), the involvement of its end-users in the development process and the evaluation of DSLs usability, i.e. quality in use of DSLs. </p> <p>Preliminary results give us a notion that there is increased the state of practice of performing the evaluation of the DSLs, mostly including usability concerns, at least after its implementation. Generally, the quality of the reviewed studies was high. On another hand, rarely the assessments are done during domain analysis, which in general is not reporting inclusion of end-users or consideration of different use-cases, although the majority of studies refer to target non-programmers and to contribute easy in use. </p> <p>We did collect the valuable body of primary studies that are giving us answers to our research questions, however, to raise the credibility of the conclusions we should extend the analysis to other venues. Also, as we get insights into the categorization of practices we could specify more concretely answers that would give us means to perform more detailed meta-analysis. </p>
CLDF dataset derived from Kitchen et al.'s "Bayesian phylogenetic analysis of Semitic languages" from 2009
<p>Cite the source of the dataset as:</p> <blockquote> <p>Bayesian phylogenetic analysis of Semitic languages identifies an Early Bronze Age origin of Semitic in the Near East. Andrew Kitchen, Christopher Ehret, Shiferaw Assefa, Connie J. Mulligan. Proc. R. Soc. B 2009 -; DOI: 10.1098/rspb.2009.0408. Published 29 April 2009</p> </blockquote>
Language Function Analysis 2011 Corpus (LFA-11)
<p>The Language Function Analysis 2011 Corpus (LFA-11) is a German text corpus of promotional text, reviews and blog posts on music and smartphones. The texts were manually classified with respect to their topic relevance, language function, and sentiment polarity.</p> <p>The purpose of the corpus is to provide textual data for the development and evaluation of approaches to language function analysis and sentiment analysis. Therefore, each text is classified by language function (personal, commercial, or informational) as well as by sentiment (positive, negative, neutral).</p> <p>The corpus consists of two separated collections, which contain the texts about <em>music</em> and <em>smartphones</em> respectively. The music collection consists of 2,713 promotional texts and reviews from both users and professionals. The smartphone collection contains 2,093 blog posts on smartphones from the <a href="http://spinn3r.com/">Spinn3r corpus</a>.</p>
Sentiment Analysis and Cross-lingual Word Embeddings for Endangered Languages
<p>A sentiment analyzer and cross-lingual word embeddings for endangered languages (e.g., Erzya, Moksha, Skolt Sami, Komi-Zyrian).</p>
OntoLAMA: LAnguage Model Analysis for Ontology Subsumption Inference
<h3><strong>About</strong></h3> <p>OntoLAMA is a set of language model (LM) probing datasets for ontology subsumption inference. The work follows the "LMs-as-KBs" literature but focuses on conceptualised knowledge extracted from formalised KBs such as the OWL ontologies. Specifically, the subsumption inference (SI) task is introduced and formulated in the Natural Language Inference (NLI) style, where the sub-concept and the super-concept involved in a subsumption axiom are verbalised and fitted into a template to form the premise and hypothesis, respectively. The sampled axioms are verified through ontology reasoning. The SI task is further divided into Atomic SI and Complex SI where the former involves only atomic named concepts and the latter involves both atomic and complex concepts. Real-world ontologies of different scales and domains are used for constructing OntoLAMA and in total there are four Atomic SI datasets and two Complex SI datasets.</p> <p> </p> <table> <tbody> <tr> <th>Dataset Source</th> <th>#Concepts</th> <th>#EquivAxioms</th> <th>#Datasets(Train/Dev/Test)</th> </tr> </tbody> <tbody> <tr> <td>Schema.org</td> <td>894</td> <td>N/A</td> <td> <p>Atomic SI: 808/404/2, 830</p> </td> </tr> <tr> <td>DOID</td> <td>11,157</td> <td>N/A</td> <td> <p>Atomic SI: 90,500/11,312/11,314</p> </td> </tr> <tr> <td>FoodOn</td> <td>30,995</td> <td>2,383</td> <td> <p>Atomic SI: 768,486/96,060/96,062</p> <p>Complex SI: 3,754/1,850/13,080</p> </td> </tr> <tr> <td>GO</td> <td>43,303</td> <td>11,456</td> <td> <p>Atomic SI: 772,870/96,608/96,610</p> <p>Complex SI: 72,318/9,040/9,040</p> </td> </tr> <tr> <td>MNLI</td> <td>N/A</td> <td>N/A</td> <td> <p>biMNLI: 235,622/26,180/12,906</p> </td> </tr> </tbody> </table> <p> </p> <h3><strong>Citation</strong></h3> <p>The relevant paper has been accepted at Findings of ACL 2023: <a href="https://aclanthology.org/2023.findings-acl.213/">https://aclanthology.org/2023.findings-acl.213/</a>.</p> <pre>```<br>@inproceedings{he2023language, title={Language Model Analysis for Ontology Subsumption Inference}, author={He, Yuan and Chen, Jiaoyan and Jimenez-Ruiz, Ernesto and Dong, Hang and Horrocks, Ian}, booktitle={Findings of the Association for Computational Linguistics: ACL 2023}, pages={3439--3453}, year={2023} }<br>```</pre> <h3><strong>Links</strong></h3> <ul> <li>See instructions at: <a href="https://krr-oxford.github.io/DeepOnto/ontolama/">https://krr-oxford.github.io/DeepOnto/ontolama/</a></li> <li>We have made available a convenient access of these datasets through Huggingface: <a href="https://huggingface.co/datasets/krr-oxford/OntoLAMA">https://huggingface.co/datasets/krr-oxford/OntoLAMA</a></li> <li>The arxiv version is available at: <a href="https://arxiv.org/abs/2302.06761">https://arxiv.org/abs/2302.06761</a></li> </ul> <h3><strong>Contact</strong></h3> <p>Yuan He (<code>yuan.he(at)cs.ox.ac.uk</code>)</p>
Token-based data sets for the analysis of the academic language of literary studies and linguistics
<p>These are the token-based data sets used for my PhD thesis ("Potentiale syntaktischer Annotationen für die datengeleitete Sprachbeschreibung am Beispiel der Wissenschaftssprachen der Germanistik", publication in progress).</p> <p>For python scripts and further data see https://github.com/melandresen/dissertation.</p> <p>Due to copyright law, the annotated texts of the corpus could only be published without the token layer. The files provided here include the token-based frequency data that have been derived from the origial texts and can be used as input to the analysis scripts in the GitHub-Repository.</p>
The TEI, Its Foundations and Impact, and How It Fits Into Today's Needs and Practices for Language Analysis
<p>6<sup>th</sup> Lecture</p>
Appendix_Results_of_quantitative_qualitative_analysis (42-language sample)
<p>A dataset showing results of quantitative and qualitative analysis on the basis of a 42-language sample.</p>
Appendix_results_qual_analysis_summarized (42-language sample)
<p>A .pdf file that plots verb scores (1 to 3) in main and adverbial clauses in the sample languages covered by the CIEP (42-language sample).</p>
Replication Package for "Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis"
<h1>Replication Package for the Paper: “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis”</h1> <p>This replication package includes the raw data, questionnaire answers, and a Python notebook needed for reproducing the results detailed in the paper titled “Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis.”</p> <h2><a></a>Repository Structure</h2> <ol> <li><strong>Scenarios:</strong> Contains an Excel file encompassing all 141 scenarios collected (in Italian).</li> <li><strong>Training and Validation Messages:</strong> Includes the jsonl files necessary for fine-tuning the model.</li> <li><strong>Testing Messages and Ground Truth:</strong> Contains the messages utilized for testing the models.</li> <li><strong>Results:</strong> Contains Excel files with the responses from the 2 human experts and the 5 model as well as the review of the 3 human reviewer.</li> <li><strong>Tables:</strong> Contains the full Wilcoxon Test Results for H01 and H02 as well as the raw RQs results.</li> </ol> <h2><a></a>Replication Process</h2> <p>To replicate the results of our study, open the provided Python Notebook in Google Colab and follow the instructions to seamlessly reproduce the results.</p> <h1><a></a>Instructions for Use</h1> <p>To utilize this replicability package, refer to the steps outlined in the notebook file.</p> <h1><a></a>Remarks</h1> <p>If you encounter any issues or have any questions, please reach out to the authors of the paper. We will be glad to assist you!</p>
Data for manuscript: "Longitudinal Analysis of Sentiment and Emotion in News Media Headlines Using Automated Labelling with Transformer Language Models"
<p>This data set contains automated sentiment and emotionality annotations of 23 million headlines from 47 popular news media outlets popular in the United States. </p> <p>The set of 47 news media outlets analysed (listed in Figure 1 of the main manuscript) was derived from the AllSides organization <a href="https://www.allsides.com/blog/updated-allsides-media-bias-chart-version-11">2019 Media Bias Chart v1.1</a>. The human ratings of outlets’ ideological leanings were also taken from this chart and are listed in Figure 2 of the main manuscript. </p> <p>News articles headlines from the set of outlets analyzed in the manuscript are available in the outlets’ online domains and/or public cache repositories such as The Internet Wayback Machine, Google cache and Common Crawl. Articles headlines were located in articles’ HTML raw data using outlet-specific XPath expressions. </p> <p>The temporal coverage of headlines across news outlets is not uniform. For some media organizations, news articles availability in online domains or Internet cache repositories becomes sparse for earlier years. Furthermore, some news outlets popular in 2019, such as <em>The Huffington Post</em> or <em>Breitbart</em>, did not exist in the early 2000’s. Hence, our data set is sparser in headlines sample size and representativeness for earlier years in the 2000-2019 timeline. Nevertheless, 18 outlets in our data set have chronologically continuous partial or full headline data availability fulfilling our inclusive criteria (see manuscript Methods) since the year 2000. Figure S 1 in the SI reports the number of headlines per outlet and per year in our analysis.</p> <p>In a small percentage of articles, outlet specific XPath expressions might fail to properly capture the content of the headline due to the heterogeneity of HTML elements and CSS styling combinations with which articles text content is arranged in outlets online domains. After manual testing, we determined that the percentage of headlines following in this category is very small. Additionally, our method might miss detecting some articles in the online domains of news outlets. To conclude, in a data analysis of over 23 million headlines, we cannot manually check the correctness of every single data instance and hundred percent accuracy at capturing headlines’ content is elusive due to the small number of difficult to detect boundary cases such as incorrect HTML markup syntax in online domains. Overall however, we are confident that our headlines set is representative of headlines in print news media content for the studied time period and outlets analyzed.</p> <p>The list of compressed files in this data set is listed next:</p> <p>-analysisScripts.rar contains the analysis scripts used in the main manuscript as well as aggregated data of sentiment and emotionality automated annotations of the headlines and human annotations of a subset of headlines sentiment and emotionality used as ground truth. </p> <p>-models.rar contains the Transformer sentiment and emotion annotation models used in the analysis. Namely: </p> <p>Siebert/sentiment-roberta-large-english from https://huggingface.co/siebert/sentiment-roberta-large-english. This model is a fine-tuned checkpoint of <a href="https://huggingface.co/roberta-large">RoBERTa-large</a> (<a href="https://arxiv.org/pdf/1907.11692.pdf">Liu et al. 2019</a>). It enables reliable binary sentiment analysis for various types of English-language text. For each instance, it predicts either positive (1) or negative (0) sentiment. The model was fine-tuned and evaluated on 15 data sets from diverse text sources to enhance generalization across different types of texts (reviews, tweets, etc.). See more information from the original authors at https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>DistilbertSST2.rar is the default sentiment classification model of the HuggingFace Transformer library https://huggingface.co/ This model is only used to replicate the results of the sentiment analysis with sentiment-roberta-large-english </p> <p>DistilRoberta j-hartmann/emotion-english-distilroberta-base from https://huggingface.co/j-hartmann/emotion-english-distilroberta-base. The model is a fine-tuned checkpoint of <a href="https://huggingface.co/distilroberta-base">DistilRoBERTa-base</a>. The model allows annotation of English text with Ekman's 6 basic emotions, plus a neutral class. The model was trained on 6 diverse datasets. Please refer to the original author at https://huggingface.co/j-hartmann/emotion-english-distilroberta-base for an overview of the data sets used for fine tuning. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromSentimentRobertaLargeModel.rar URLs of headlines analyzed and the sentiment annotations of the siebert/sentiment-roberta-large-english Transformer model. https://huggingface.co/siebert/sentiment-roberta-large-english</p> <p>-headlinesDataWithSentimentLabelsAnnotationsFromDistilbertSST2.rar URLs of headlines analyzed and the sentiment annotations of the default HuggingFace sentiment analysis model fine-tuned on the SST-2 dataset. https://huggingface.co/</p> <p>-headlinesDataWithEmotionLabelsAnnotationsFromDistilRoberta.rar URLs of headlines analyzed and the emotion categories annotations of the j-hartmann/emotion-english-distilroberta-base Transformer model. https://huggingface.co/j-hartmann/emotion-english-distilroberta-base</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.