Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
70
datasets available to search
ShareScore release 0.9.0
Dataset results
70 results for “language analysis”
Mocap video examples for the analysis of Sign Language movements
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p> <p> </p> <p> </p>
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p>
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p>
Mocap video examples for the analysis of Sign Language motion
<p>These mocap videos support my PhD thesis "Extracting human characteristics from motion: the case of identity in Sign Language" carried out from October 2018 to October 2021. The original mocap data is taken from the <a href="https://www.ortolang.fr/market/corpora/mocap1/">MOCAP1</a> corpus of French Sign Language. The videos have been generated using Python code available as part of the <a href="https://github.com/felixbgd/PLmocap">PLmocap</a> library.</p> <p> </p>
Dataset for "Non-canonical possessive constructions in Negidal and other Tungusic languages: a new analysis of the so-called 'alienable possession' suffix"
<p>This is the dataset used in the paper: Aralova, N. & Pakendorf, B. (2023). Non-canonical possessive constructions in Negidal and other Tungusic languages: a new analysis of the so-called “alienable possession” suffix. Special Issue “Re-assessing the explanatory potential of alienability contrasts”, guest-edited by Françoise Rose & An Van linden. <em>Linguistics</em>. <a href="https://doi.org/10.1515/ling-2022-0030">https://doi.org/10.1515/ling-2022-0030</a></p> <p>For more details, see the ReadMe file.</p>
Data from: The appropriateness of language found in research consent form templates: a computational linguistic analysis
Background: To facilitate informed consent, consent forms should use language below the grade eight level. Research Ethics Boards (REBs) provide consent form templates to facilitate this goal. Templates with inappropriate language could promote consent forms that participants find difficult to understand. However, a linguistic analysis of templates is lacking. Methods: We reviewed the websites of 124 REBs for their templates. These included English language medical school REBs in Australia/New Zealand (n=23), Canada (n=14), South Africa (n=8), the United Kingdom (n=34), and a geographically-stratified sample from the United States (n=45). Template language was analyzed using Coh-Metrix linguistic software (v.3.0, Memphis, USA). We evaluated the proportion of REBs with five key linguistic outcomes at or below grade eight. Additionally, we compared quantitative readability to the REBs' own readability standards. To determine if the template's country of origin or the presence of a local REB readability standard influenced the linguistic variables, we used a MANOVA model. Results: Of the REBs who provided templates, 0/94 (0%, 95% CI=0-3.9%) provided templates with all linguistic variables at or below the grade eight level. Relaxing the standard to a grade 12 level did not increase this proportion. Further, only 2/22 (9.1%, 95% CI= 2.5-27.8) REBs met their own readability standard. The country of origin (DF= 20, 177.5, F=1.97, p=0.01), but not the presence of an REB-specific standard (DF=5, 84, F=0.73, p=0.60), influenced the linguistic variables. Conclusions: Inappropriate language in templates is an international problem. Templates use words that are long, abstract, and unfamiliar. This could undermine the validity of participant informed consent. REBs should set a policy of screening templates with linguistic software.
Appendix_results_qual_analysis_summarized (39-language sample)
<p>A dataset showing a summary of the results of qualitative analysis.</p>
Graph Neural Network vs. Large Language Model: A Comparative Analysis for Bug Report Priority and Severity Prediction
Open the record for dataset details and reuse information.
Replication data for the paper "Leveraging Large Language Models for Comprehensive Psychological Analysis: Insights from Four Theoretical Frameworks"
<p>This is a replication data for the paper titled "Leveraging Large Language Models for Comprehensive Psychological Analysis: Insights from Four Theoretical Frameworks" submitted for a blind review.</p> <p>Abstract</p> <p>The rapid advancement of generative Artificial Intelligence (AI) has significantly transformed various research domains. This paper introduces a novel, fully automated methodology for applying Large Language Models (LLMs) to psychological text analysis. The approach includes prompt design for zero-shot and few-shot learning, model internal consistency analysis, autonomous machine evaluation, and additional human validation. Applied to four psychological theories—Self-Determination Theory, the Big Five Personality Traits, Psychological Well-being, and Cognitive Behavioral Therapy—this methodology is tested on a dataset of 25,780 emails written by a senior executive (called Person X) over 16 years. The analysis involves extracting psychological characteristics from the emails and regressing these characteristics against personal, professional, and environmental factors. The results demonstrate that the methodology provides unique insights into the examined psychological theories, offering a detailed understanding of how various factors influence psychological states and traits over time. This research highlights the potential of LLMs in capturing and analyzing complex psychological patterns in large text corpora, contributing a robust framework for future studies and practical applications in psychological assessment and intervention. The findings underscore the transformative impact of generative AI in psychological research, opening new avenues for understanding human behavior through advanced language models.</p> <p>The zipped file contains five csv files:</p> <ol> <li>Email_classification-csv: LLM (GPT-3.5 Turbo) classification of 25,780 emails for four psychological theories: SDT, Big Five, PWB and CBT.</li> <li>SDT_regression_data.csv</li> <li>Big_Five_regression_data.csv</li> <li>PWB_regression_data.csv</li> <li>CBT_regression_data.csv</li> </ol> <p>For 2-5 files the dependent variable is monthy percentage share of emails the were assigned a given value for categories of one of the four psychological theories analyzed. </p> <p>Linear regression model has been applied, where dependent variable is the percentage of emails in a specified category that assigned a specific value in this category. For example in Big Five Traits Model, for the Openness category, for each month we calculated percentage of emails that exhibit <em>High</em> or <em>Low</em> openness, or <em>None</em> if the content of the email does not provide enough information to assess whether the specific need is relevant. Two dependent variables were created: <em>Openness-high</em> and <em>Openness-low</em> and regressed on all independent variables. Regressions were not run for the <em>None</em> values.</p> <p>Descriptions of independent variables:</p> <p>- <em>income_index</em>: Person X salary income and consulting fees in a given month, normalized to [0,1].</p> <p>- <em>card_spending</em>: Person X credit card expenditures in a given month, normalized to [0,1].</p> <p>- <em>abroad_far</em>: dummy variable set to 1 for months when Person X worked in Central Asia</p> <p>- <em>abroad_near</em>: dummy variable set to 1 when Person X worked in other EU country</p> <p>- <em>death_1_war</em>: variable set to 1 in a month when Person X’ farther in law passed away. In the same month Russia invaded Ukraine. The variable was set to .75 in the following month, and to .5 in the month after that.</p> <p>- <em>death_2</em>: variable set to 1 in a month when Person X’ mother passed away. The variable was set to .75 in the following month, and to .5 in the month after that.</p> <p>- <em>court_case</em>: dummy variable set to 1 for months with the emotionally engaging inheritance court case involving other family members.</p> <p>- <em>BIG4_partner</em>: dummy variable set to 1 for months when Person X worked as a partner in BIG4 accounting firm, which resulted in adopting a professional activity sharply different from the usual Person X habits.</p> <p>- <em>AI_company</em>: dummy variable set to 1 for months when Person X worked as C-level executive at a company specializing in artificial intelligence.</p> <p>- <em>elections</em>: dummy variable set to 1 for months when Person X unsuccessfully run in parliamentary elections</p> <p>- <em>covid_lockdown</em>: dummy variable set to 1 for month where Polish government imposed tough measures during two covid lockdowns.</p> <p>- <em>no_receive</em>: number of different email recipients each month, normalized to [0,1].</p> <p>- <em>avg_length</em>: average number of words in emails sent each month, normalized to [0,1].</p> <p> While the email data was collected for January 2008 – March 2014 period, financial data was available from October 2009. There were some months where no emails with more than 10 words were sent, yielding 166 monthly observations used for regressions, before removing outliers.</p> <p> Independent variables were tested for multicollinearity, outlier months were removed, regressions were estimated with robust standard errors, and a range of standard tests were conducted for normality and autocorrelation of residuals, confirming good statistical properties of estimated models.</p> <p>Due to privacy concerns, the email texts cannot be publicly shared. However, the classifications of psychological categories derived from the email texts, along with all other relevant data, are made publicly available in this open access repository, with the consent of email author.</p> <p> </p>
Analysis of Language and Auditory Abilities in Cochlear Implanted Children
ClinicalTrials.gov study NCT00947778. IPD Sharing: Not stated. Countries: 1. Publications: 0.
Data from: The appropriateness of language found in research consent form templates: a computational linguistic analysis
Open the record for dataset details and reuse information.
THE LINGUACULTURAL CONCEPT OF DOUBT: A COMPARATIVE ANALYSIS OF ENGLISH AND UZBEK LANGUAGES
Open the record for dataset details and reuse information.
COMPARATIVE ANALYSIS OF THE CONCEPT "LOVE" IN THE PHRASEOLOGY OF ENGLISH AND UZBEKI LANGUAGES
Open the record for dataset details and reuse information.
SEMANTIC AND STRUCTURAL ANALYSIS OF THE TOURISTIC TERMS IN THE ENGLISH AND UZBEK LANGUAGES.
Open the record for dataset details and reuse information.
COMPARATIVE ANALYSIS OF SURGICAL TERMS IN UZBEK AND ENGLISH LANGUAGE
Open the record for dataset details and reuse information.
Supplementary material 1 from: Vyshedskiy A, Mahapatra S, Dunn R (2017) Linguistically deprived children: meta-analysis of published research underlines the importance of early syntactic language use for normal brain development. Research Ideas and Outcomes 3: e20696. https://doi.org/10.3897/rio.3.e20696
Linguistic isolates performance in verbal and nonverbal tests
Figure 5 from: Vyshedskiy A, Mahapatra S, Dunn R (2017) Linguistically deprived children: meta-analysis of published research underlines the importance of early syntactic language use for normal brain development. Research Ideas and Outcomes 3: e20696. https://doi.org/10.3897/rio.3.e20696
Figure 5 - Flexible syntax, prepositions, adjectives, verb tenses, and other common elements of grammar, all facilitate the human ability to communicate an infinite number of novel images with the use of a finite number of words. The graph shows the number of distinct images that can be transmitted with high fidelity in a communication system with 1,000 nouns as a function of the number of spatial prepositions. In a communication system with no spatial prepositions and other recursive elements, 1000 nouns can communicate 1000 images to a listener. Adding just one spatial preposition allows for the formation of three-word phrases (such as: 'a bowl behind a cup' or 'a cup behind a bowl') and increases the number of distinct images that can be communicated to a listener from 1000 to one million (1000x1x1000). Adding a second spatial preposition and allowing for five-word sentences of the form object-preposition-object-preposition-object (such as: a bowl on a cup behind a plate) increases the number of distinct images that can be communicated to four billion (1000x2x1000x2x1000). The addition of a third spatial preposition increases the number of distinct images to 27 trillion (1000x3x1000x3x1000x3x1000), and so on. In general, the number of distinct images communicated by three-word sentences of the structure object-preposition-object equals the number of object-words times the number of prepositions times the number of object-words. A typical language with 1000 nouns and 100 spatial prepositions can theoretically communicate 1000101 x 100100 distinct images. This number is significantly greater than the total number of atoms in the universe. For all practical purposes, an infinite number of distinct images can be communicated by a syntactic communication system with just 1000 words and a few prepositions. Prepositions, adjectives, and verb tenses dramatically facilitate the capacity of a syntactic communication system with a finite number of words to communicate an infinite number of distinct images. Linguists refer to this property of human languages as recursion. The "infiniteness" of human language has been explicitly recognized by "Galileo, Descartes, and the 17th-century 'philosophical grammarians' and their successors, notably von Humboldt" (Hauser et al. 2002). The infiniteness of all human languages stand in stark contrast to finite homesign communication systems that are lacking spatial prepositions, syntax, and other recursive elements of a formal sign language.
Figure 4 from: Vyshedskiy A, Mahapatra S, Dunn R (2017) Linguistically deprived children: meta-analysis of published research underlines the importance of early syntactic language use for normal brain development. Research Ideas and Outcomes 3: e20696. https://doi.org/10.3897/rio.3.e20696
Figure 4 - Synchronicity has to be understood in terms of synchronicity of the arrival of action potentials to a target neuron rather than absolute equality of action potential conduction times over different paths. Consider the following example: suppose neuron A is receiving excitatory input from neurons B and C via two different pathways (neuron A is the target neuron for both neurons B and C). Suppose that the action potential conduction time is 2ms from neuron B to neuron A and 22ms from neuron C to neuron A (i.e., the axonal pathway B-A has a significantly shorter conduction time than the axonal pathway C-A). Does it mean that the connections B-A and C-A are always asynchronous? No. The answer depends on the predominant neural activity rhythm in this network. At the firing rate of 50Hz (inter-spike interval of 20ms that correspond to Gamma rhythm), neurons B and C can actually be considered synchronous in relationship to neuron A: consider a train of action potentials synchronously fired by neurons B and C. The first action potential from neuron B will reach neuron A in 2ms and the first action potential from neuron C will reach neuron A in 22ms. Obviously, there would be no coincidence in the arrival times of the 1st action potentials from neurons B and C. However the second action potential from neuron B will arrive to neuron A in 22ms, concurrently with the 1st action potential from neuron C. Thus, starting with the second action potential, neuron A will receive synchronous activation from neurons B and C. The synchronous activation has a significantly greater probability of enhancing synaptic connections between neurons A and B, and A and C (Hebbian learning: 'neurons that fire together, wire together' (Hebb 1949). Thus, synchronicity does not need to imply absolute equality in the conduction time over different pathways. Rather synchronicity implies near-zero phase-shift between the two firing trains of action potentials at the postsynaptic cells. This phase-shift depends on conduction times over each pathway and also on the dominant firing frequency in the neural network.
Figure 2 from: Vyshedskiy A, Mahapatra S, Dunn R (2017) Linguistically deprived children: meta-analysis of published research underlines the importance of early syntactic language use for normal brain development. Research Ideas and Outcomes 3: e20696. https://doi.org/10.3897/rio.3.e20696
Figure 2 - A typical question testing subject's ability to mentally rotate an object is shown here as a 2x2 matrix with six answer choices displayed below the problem. The top row of the matrix indicates the rule: "the object in the right column is the result of 45° clockwise rotation." Applying this rule to the bottom row, we arrive at the correct answer depicted on the right.
Figure 1 from: Vyshedskiy A, Mahapatra S, Dunn R (2017) Linguistically deprived children: meta-analysis of published research underlines the importance of early syntactic language use for normal brain development. Research Ideas and Outcomes 3: e20696. https://doi.org/10.3897/rio.3.e20696
Figure 1 - Visual information processing in the cortex. From the primary visual cortex (V1, shown in yellow), the visual information is passed in two streams. The neurons along the ventral stream also known as the ventral visual cortex (shown in purple) are primarily concerned with what the object is. The ventral visual stream runs into the inferior temporal lobe. The neurons along the dorsal stream also known the dorsal visual cortex (shown in green) are primarily concerned with where the object is. The dorsal visual stream runs into the parietal lobe.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.