Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,721
datasets available to search
ShareScore release 0.9.0
Dataset results
1,721 results for “Language”
The Alice Dataset: fMRI Dataset to Study Natural Language Comprehension in the Brain
Open the record for dataset details and reuse information.
The language network reemerges during recovery from severe traumatic brain injury
Open the record for dataset details and reuse information.
Dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area"
<p>This repository archives the dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area". The cognate annotation of Tshangla, Kho-Bwa, Hrusish, Mishmic, and Tani languages were done by us. The cognate decision on the other languages was annotated by Sagart et al. (2019). Please use the following information to cite our work: <br> Wu, M.-S, Bodt, T. A, Tresoldi, T. (2022). Bayesian phylogenetics illuminate shallower relationships Trans-Himalayan languages in the Tibet-Arunachal area. Linguistics of the Tibeto-Burman Area. [forthcoming]</p>
Improving the Developer Experience with a Low-Code ProcessModelling Language: Companion site
<p>This companion site contains additional data to complement the paper:</p> <p><em><strong>Henriques, H., Lourenço, H., Amaral, V., and Goulão, M. (2018). Improving the developer experience with a low-code process </strong></em><em><strong>modelling</strong></em><em><strong> language. In ACM/IEEE 21st International Conference on Model Driven Engineering Languages and Systems (MODELS 2018), Copenhagen, Denmark. ACM. https://doi.org/10.1145/3239372.3239387</strong></em></p> <p><strong>Abstract</strong></p> <p><strong>Context</strong><strong>: </strong>The OutSystems Platform is a development environment composed of several DSLs, used to specify, quickly build and validate web and mobile applications. The DSLs allow users to model different perspectives such as interfaces and data models, define custom business logic and construct process models.</p> <p><strong>Problem</strong><strong>: </strong>TheDSL for process modelling (Business Process Technology (BPT)), has a low adoption rate and is perceived as having usability problems hampering its adoption. This is problematic given the language maintenance costs.</p> <p><strong>Method:</strong> We used a combination of interviews, a critical review of BPT using the “Physics of Notation” and empirical evaluations of BPT using the System Usability Scale (SUS)and the NASA Task Load indeX (TLX), to develop a new version ofBPT, taking these inputs and Outsystems’ engineers culture into account.</p> <p><strong>Results: </strong>Evaluations conducted with 25 professional soft-ware engineers showed an increase of the semantic transparency on the new version, from 31% to 69%, an increase in the correctness of responses, from 51% to 89%, an increase in the SUS score, from 42.25 to 64.78, and a decrease of the TLX score, from 36.50 to 20.78. These differences were statistically significant.</p> <p><strong>Conclusions:</strong> These results suggest the new version of BPT significantly improved the developer experience of the previous version. The end users background with OutSystems had a relevant impact on the final concrete syntax choices and achieved usability indicators.</p> <p> </p> <p><strong>Contents</strong></p> <p>This companion site provides a permanent link for additional data to the supported paper.</p> <p>This repository includes:</p> <ul> <li>Surveys and Questionnaires used in the evaluation reported in the paper <ul> <li>Survey on OutSystems BPT notations (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/survey.pdf">survey.pdf</a>)</li> <li>Prototype Symbol Set Questionnaire (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/PrototypeSymbolSetQuestionnaire.pdf">PrototypeSymbolSetQuestionnaire.pdf</a>)</li> <li>Original BPT Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/languages.png">languages.png</a>)</li> <li>Usability Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/sus.png">sus.png</a>)</li> <li>Cognitive Effort Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/tlx.png">tlx.png</a>)</li> <li>Testing environment screenshot (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/Testing%20Environment%20Screenshot.png">Testing Environment Screenshot</a>)</li> </ul> </li> <li>Statistics <ul> <li>SUS and NASA TLX <ul> <li>Descriptive statistics (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXDescriptiveStats.pdf">SUSTLXDescriptiveStats.pdf</a>)</li> <li>Normality tests (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXNormalityTests.pdf">SUSTLXNormality.pdf</a>)</li> <li>Correlation test (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXCorrelation.pdf">SUSTLXCorrelation.pdf</a>)</li> <li>Scatterplot (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXScatterPlot.pdf">SUSTLXScatterplot.pdf</a>)</li> </ul> </li> </ul> </li> </ul> <p> </p>
functional MRI study on the language stress perception in a foreign language
<p>fMRI dataset of 91 participants during a linguistic task about language stress perception in a foreign language.</p> <p>Participants listened to pairs of words in a foreign language (Spanish) and had to indicate if the words were the same or different. The different pairs differed either by the stress pattern, or by the final vowel.</p> <p>This dataset was divided into two groups: 51 participants with French as native language and 40 with Swiss-German as native language. None of the participants had knowledge of Spanish.</p> <p>This repository respects the BIDS standard (<a href="https://bids.neuroimaging.io/">https://bids.neuroimaging.io/</a>), including all the raw data (func, fmap, anat) and metadata in order to reproduce the processing.</p> <p>These data have been used in two papers:</p> <p>S. Schwab, M. Mouthon, L.B. Jost, J. Salvadori, I. Yakoub, E. Ferreira da Silva, N. Giroud, B. Perriard and J.M. Annoni, Neural correlates of lexical stress processing in a foreign free-stress language; Brain and Behavior (2023)</p> <p>L. Rogenmoser, M. Mouthon, F. Etter, J. Kamber, J.M. Annoni and S. Schwab; The processing of stress in a foreign language modulates functional antagonism between default mode and attention network regions, (submitted)</p>
Adult language learners
Open the record for dataset details and reuse information.
Adolescent language learners
Open the record for dataset details and reuse information.
Collection of spatial information and maps of human past and environment in the Uralic languages speaker area
<p>The collection of spatial information and maps of the past and environment in the Uralic languages speaker area consists excessive amount of multidisciplinary data related to the vast region extending from Eastern Europe to Siberia, encompassing countries like Russia, Finland, and parts of Scandinavia. Uralic speakers are predominantly found in this region, with historical roots in areas around the Ural Mountains and adjacent territories. These datasets can be integrated for multidisciplinary purposes, allowing to explore human-environment interactions, migration patterns, and cultural evolution over time. Datasets are collected initially by the BEDLAN team <a href="https://bedlan.net/">https://bedlan.net/</a> - a research group specialized in various disciplines - linguists, archaeologists, geneticists, and geographers. The data collection and mapmaking have grown beyond the initial stages (publications, applications, exhibitions), hence collaborative effort for data publishing is now crucial. As the data collections and mapmaking continue to evolve dynamically together with ongoing projects, the current repository will be updated accordingly.</p>
CloneCorp: Cross-language clone detection dataset
<p>Data set of mobile apps and code examples to evaluate clone detection algorithms across languages (Kotlin, Swift, and Dart)</p>
EDICTOR 3: Interactive Tool for Computer-Assisted Language Comparison
This software offers the most recent and mostly stable version of the EDICTOR tool, also available for direct usage from <a href="https://edictor.org">edictor.org/</a>.
Coding data to accompany "A quantitative approach to sociotopography in Austronesian languages"
<p>Dataset consists of csv files with sample languages identified by name and Glottocode. Coding for four sociolinguistic variables, as well as an overall "orientation type." Each file corresponds to a different method for coding languages employing multiple spatial orientation strategies, as described in the document coding.pdf.</p> <p><strong>Orientation type</strong></p> <ul> <li>land-sea = axis oriented orthogonal to the coast, based on opposition between landward (inland) and seaward (toward the coast), regardless of whether these terms reflect PAN *daya and *lahud </li> <li>land-sea* = land-sea systems in which the land-sea opposition is indistinguishable from geophysical elevation</li> <li>coastal = axis oriented parallel to the coast, often but not necessarily co-lexified with vertical `up' and `down'</li> <li>elevation = axis that distinguishes global or geophysical elevation with respect to deictic center </li> <li>riverine = axis oriented parallel to the river, typically with secondary axis orientated orthogonal to river</li> <li>cardinal = axis fixed according to conventions which do not vary with local geography (although they may be motivated by environmental factors such as wind and the sun)</li> </ul> <p><strong>Distribution</strong></p> <ul> <li>distributed</li> <li>island</li> <li>village</li> </ul> <p><strong>Economy</strong></p> <ul> <li>diversified</li> <li>agriculture</li> <li>subsistence</li> </ul> <p><strong>Geography</strong></p> <ul> <li>diversified</li> <li>inland</li> <li>coast</li> </ul> <p><strong>Terrain</strong></p> <ul> <li>mountainous</li> <li>non-mountainous</li> </ul>
Language-enhanced cognitive skills model
<p>This is a model of cognitive skills required in the workplace which enhance previous models by including a more detailed measurement of linguistic skills. Linguistic skills are defined as the set of abilities, competencies and knowledge which principally involve the use of linguistic code. More specifically, the linguistic items used are reading and writing competencies, ability to speak, listen or communicate, as well as knowledge of second languages (as a whole). These variables were factorialised together with a list of competencies from previous models of cognitive skills. Principal component analysis (PCA) with equamax rotation was applied to reduce the dimensionality of all items to a few interpretable dimensions according to the correlations between them. The result is nine factors with similar variances among at least three express linguistic-related skills: The first factor expresses the demand for scientific and engineering knowledge. The second refers to a collection of competencies which could be called verbal-reasoning. These include deductive and inductive reasoning skills or those of identifying and solving complex problems. Some linguistic competencies relating to the level of oral and written comprehension and expression are also relevant in this factor. The third factor expresses numerical or quantitative competencies. The fourth expresses the demand for communicative competencies, composed of variables related to efficient communication goals such as clarity of speech, active listening or speaking. The fifth factor expresses creative abilities. The sixth, competencies and knowledge linked to electronics and computers. The seventh expresses managerial competencies. The eighth expresses nurturing competencies and the ninth factor basically expresses knowledge of foreign languages.</p>
Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language
<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>
Polifonia Corpus - Books Module Metadata - French Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Dutch Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - German Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Spanish Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Polifonia Corpus - Books Module Metadata - Italian Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Architectural Languages for the Microservices Architecture: A systematic mapping study [Data set]
<p>This repository contains all artifacts related to the study: Architectural Languages for the Microservices Architecture: A systematic mapping study.</p>
Languages of Glottolog 4.6 proyect
<p><strong>Glottolog is collaborative work. Harald Hammarström collected many individuals' bibliographies and compiled them into a master bibliography. Harald also collected extensive information about the proved genealogical relations of the languages of the world. His top-level classification is merged with low-level (i.e. dialect level) information from multitree. Sebastian Nordhoff designed and programmed the database and the first version of the web application with help from Hagen Jung and Robert Forkel. He also took care of the import of the bibliographies. Robert Forkel programmed the second version of the web application as part of the Cross Linguistic Linked Data project. Martin Haspelmath gave advice at every stage throughout the project, helped with coordination, and is currently responsible for languoid names and dialects. Sebastian Bank made numerous improvements to the code base for the Glottolog website and organized the import of several large bibliographies. A table of source bibliographies for Glottolog is available at References information. The Agglomerated Endangerment Status (AES) is derived from the databases of <a href="http://www.endangeredlanguages.com/"> </a>The Catalogue of Endangered Languages (ELCat), UNESCO Atlas of the World's Languages in Danger and <a href="http://www.ethnologue.com/"> </a>Ethnologue. For more information see GlottoScope.</strong></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.