Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,721

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,721 results for “Language”

Learn how ShareScore rates datasets ↗
OpenNeuro52/100

The Alice Dataset: fMRI Dataset to Study Natural Language Comprehension in the Brain

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
OpenNeuro52/100

The language network reemerges during recovery from severe traumatic brain injury

Open the record for dataset details and reuse information.

openCC0Jan 2020View details →
zenodo52/100

Dataset of the article "Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area"

<p>This repository archives the dataset of the article &quot;Bayesian phylogenetics illuminate shallower relationships in Trans-Himalayan languages in Tibet-Arunachal area&quot;. The cognate annotation of Tshangla, Kho-Bwa, Hrusish, Mishmic, and Tani languages were done by us. The cognate decision on the other languages was annotated by Sagart et al. (2019).&nbsp;&nbsp;Please use the following information to cite our work:&nbsp;<br> Wu, M.-S, Bodt, T. A, Tresoldi, T. (2022). &nbsp;Bayesian phylogenetics illuminate shallower relationships Trans-Himalayan languages in the Tibet-Arunachal area. Linguistics of the Tibeto-Burman Area. [forthcoming]</p>

opencc-by-4.0Dec 2021View details →
zenodo52/100

Improving the Developer Experience with a Low-Code ProcessModelling Language: Companion site

<p>This companion site contains additional data to complement the paper:</p> <p><em><strong>Henriques, H., Louren&ccedil;o, H., Amaral, V., and Goul&atilde;o, M. (2018). Improving the developer experience with a low-code process </strong></em><em><strong>modelling</strong></em><em><strong> language. In ACM/IEEE 21st International Conference on Model Driven Engineering Languages and Systems (MODELS 2018), Copenhagen, Denmark. ACM. https://doi.org/10.1145/3239372.3239387</strong></em></p> <p><strong>Abstract</strong></p> <p><strong>Context</strong><strong>:&nbsp;</strong>The OutSystems Platform is a development environment composed of several DSLs, used to specify, quickly build and validate web and mobile applications. The DSLs allow users to model different perspectives such as interfaces and data models, define custom business logic and construct process models.</p> <p><strong>Problem</strong><strong>:&nbsp;</strong>TheDSL for process modelling (Business Process Technology (BPT)), has a low adoption rate and is perceived as having usability problems hampering its adoption. This is problematic given the language maintenance costs.</p> <p><strong>Method:</strong> We used a combination of interviews, a critical review of BPT using the &ldquo;Physics of Notation&rdquo; and empirical evaluations of BPT using the System Usability Scale (SUS)and the NASA Task Load indeX (TLX), to develop a new version ofBPT, taking these inputs and Outsystems&rsquo; engineers culture into account.</p> <p><strong>Results:&nbsp;</strong>Evaluations conducted with 25 professional soft-ware engineers showed an increase of the semantic transparency on the new version, from 31% to 69%, an increase in the correctness of responses, from 51% to 89%, an increase in the SUS score, from 42.25 to 64.78, and a decrease of the TLX score, from 36.50 to 20.78. These differences were statistically significant.</p> <p><strong>Conclusions:</strong> These results suggest the new version of BPT significantly improved the developer experience of the previous version. The end users background with OutSystems had a relevant impact on the final concrete syntax choices and achieved usability indicators.</p> <p>&nbsp;</p> <p><strong>Contents</strong></p> <p>This companion site provides a permanent link for additional data to the supported paper.</p> <p>This repository includes:</p> <ul> <li>Surveys and Questionnaires used in the evaluation reported in the paper <ul> <li>Survey on OutSystems BPT notations (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/survey.pdf">survey.pdf</a>)</li> <li>Prototype Symbol Set Questionnaire (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/PrototypeSymbolSetQuestionnaire.pdf">PrototypeSymbolSetQuestionnaire.pdf</a>)</li> <li>Original BPT Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/languages.png">languages.png</a>)</li> <li>Usability Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/sus.png">sus.png</a>)</li> <li>Cognitive Effort Evaluation (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/tlx.png">tlx.png</a>)</li> <li>Testing environment screenshot (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/Testing%20Environment%20Screenshot.png">Testing Environment Screenshot</a>)</li> </ul> </li> <li>Statistics <ul> <li>SUS and NASA TLX <ul> <li>Descriptive statistics (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXDescriptiveStats.pdf">SUSTLXDescriptiveStats.pdf</a>)</li> <li>Normality tests (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXNormalityTests.pdf">SUSTLXNormality.pdf</a>)</li> <li>Correlation test (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXCorrelation.pdf">SUSTLXCorrelation.pdf</a>)</li> <li>Scatterplot (<a href="https://zenodo.org/api/files/68bdc7fa-684a-496d-ab63-d956271f1f7d/SUSTLXScatterPlot.pdf">SUSTLXScatterplot.pdf</a>)</li> </ul> </li> </ul> </li> </ul> <p>&nbsp;</p>

opencc-by-4.0Jul 2018View details →
zenodo52/100

functional MRI study on the language stress perception in a foreign language

<p>fMRI dataset of 91 participants during a linguistic task about language stress perception in a foreign language.</p> <p>Participants listened to pairs of words in a foreign language (Spanish) and had to indicate if the words were the same or different. The different pairs differed either by the stress pattern, or by the final vowel.</p> <p>This dataset was divided into two groups: 51 participants with French as native language and 40 with Swiss-German as native language. None of the participants had knowledge of Spanish.</p> <p>This repository respects the BIDS standard (<a href="https://bids.neuroimaging.io/">https://bids.neuroimaging.io/</a>), including all the raw data (func, fmap, anat) and metadata in order to reproduce the processing.</p> <p>These data have been used in two papers:</p> <p>S. Schwab, M. Mouthon, L.B. Jost, J. Salvadori, I. Yakoub, E. Ferreira da Silva, N. Giroud, B. Perriard and J.M. Annoni, Neural correlates of lexical stress processing in a foreign free-stress language; Brain and Behavior (2023)</p> <p>L. Rogenmoser, M. Mouthon, F. Etter, J. Kamber, J.M. Annoni and S. Schwab; The processing of stress in a foreign language modulates functional antagonism between default mode and attention network regions, (submitted)</p>

opencc-by-4.0Dec 2022View details →
OpenNeuro48/100

Adult language learners

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
OpenNeuro48/100

Adolescent language learners

Open the record for dataset details and reuse information.

openCC0Jan 2021View details →
zenodo48/100

Collection of spatial information and maps of human past and environment in the Uralic languages speaker area

<p>The collection of spatial information and maps of the past and environment in the Uralic languages speaker area consists excessive amount of multidisciplinary data related to the vast region extending from Eastern Europe to Siberia, encompassing countries like Russia, Finland, and parts of Scandinavia. Uralic speakers are predominantly found in this region, with historical roots in areas around the Ural Mountains and adjacent territories. These datasets can be integrated for multidisciplinary purposes, allowing to explore human-environment interactions, migration patterns, and cultural evolution over time. Datasets are collected initially by the BEDLAN team <a href="https://bedlan.net/">https://bedlan.net/</a>&nbsp; - a research group specialized in various disciplines - linguists, archaeologists, geneticists, and geographers. The data collection and mapmaking have grown beyond the initial stages (publications, applications, exhibitions), hence collaborative effort for data publishing is now crucial. As the data collections and mapmaking continue to evolve dynamically together with ongoing projects, the current repository will be updated accordingly.</p>

opencc-by-4.0Oct 2023View details →
zenodo48/100

CloneCorp: Cross-language clone detection dataset

<p>Data set of mobile apps and code examples to evaluate clone detection algorithms across languages (Kotlin, Swift, and Dart)</p>

opencc-by-4.0Jan 2024View details →
zenodo48/100

EDICTOR 3: Interactive Tool for Computer-Assisted Language Comparison

This software offers the most recent and mostly stable version of the EDICTOR tool, also available for direct usage from <a href="https://edictor.org">edictor.org/</a>.

openmit-licenseApr 2021View details →
zenodo48/100

Coding data to accompany "A quantitative approach to sociotopography in Austronesian languages"

<p>Dataset consists of csv files with sample languages identified by name and Glottocode. Coding for four sociolinguistic variables, as well as an overall &quot;orientation type.&quot; Each file corresponds to a different method for coding languages employing multiple spatial orientation strategies, as described in the document coding.pdf.</p> <p><strong>Orientation type</strong></p> <ul> <li>land-sea = axis oriented orthogonal to the coast, based on opposition between landward (inland) and seaward (toward the coast), regardless of whether these terms reflect PAN *daya and *lahud&nbsp;</li> <li>land-sea* = land-sea systems in which the land-sea opposition is indistinguishable from &nbsp;geophysical elevation</li> <li>coastal = axis oriented parallel to the coast, often but not necessarily co-lexified with vertical `up&#39; and `down&#39;</li> <li>elevation = axis that &nbsp;distinguishes global or geophysical elevation with respect to deictic center&nbsp;</li> <li>riverine = axis oriented parallel to the river, typically with secondary axis orientated orthogonal to river</li> <li>cardinal = axis fixed according to conventions which do not vary with local geography (although they may be motivated by environmental factors such as wind and the sun)</li> </ul> <p><strong>Distribution</strong></p> <ul> <li>distributed</li> <li>island</li> <li>village</li> </ul> <p><strong>Economy</strong></p> <ul> <li>diversified</li> <li>agriculture</li> <li>subsistence</li> </ul> <p><strong>Geography</strong></p> <ul> <li>diversified</li> <li>inland</li> <li>coast</li> </ul> <p><strong>Terrain</strong></p> <ul> <li>mountainous</li> <li>non-mountainous</li> </ul>

opencc-by-4.0Apr 2021View details →
zenodo48/100

Language-enhanced cognitive skills model

<p>This is a model of cognitive skills required in the workplace which enhance previous models by including a more detailed measurement of linguistic skills. Linguistic skills are defined as the set of abilities, competencies and knowledge which principally involve the use of linguistic code. More specifically, the linguistic items used are reading and writing competencies, ability to speak, listen or communicate, as well as knowledge of second languages (as a whole). These variables were factorialised together with a list of competencies from previous models of cognitive skills. Principal component analysis (PCA) with equamax rotation was applied to reduce the dimensionality of all items to a few interpretable dimensions according to the correlations between them.&nbsp;The result is nine factors with similar variances among at least three express linguistic-related skills: The first factor expresses the demand for scientific and engineering knowledge. The second refers to a collection of competencies which could be called verbal-reasoning. These include deductive and inductive reasoning skills or those of identifying and solving complex problems. Some linguistic competencies relating to the level of oral and written comprehension and expression are also relevant in this factor. The third factor expresses numerical or quantitative competencies. The fourth expresses the demand for communicative competencies, composed of variables related to efficient communication goals such as clarity of speech, active listening or speaking. The fifth factor expresses creative abilities. The sixth, competencies and knowledge linked to electronics and computers. The seventh expresses managerial competencies. The eighth expresses nurturing competencies and the ninth factor basically expresses knowledge of foreign languages.</p>

opencc-by-4.0Jan 2022View details →
zenodo48/100

Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language

<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Polifonia Corpus - Books Module Metadata - French Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Polifonia Corpus - Books Module Metadata - Dutch Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Polifonia Corpus - Books Module Metadata - German Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Polifonia Corpus - Books Module Metadata - Spanish Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Polifonia Corpus - Books Module Metadata - Italian Language (Full)

<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>

opencc-by-4.0Jun 2022View details →
zenodo48/100

Architectural Languages for the Microservices Architecture: A systematic mapping study [Data set]

<p>This repository contains all artifacts related to the study: Architectural Languages for the Microservices Architecture: A systematic mapping study.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Languages of Glottolog 4.6 proyect

<p><strong>Glottolog is collaborative work. Harald Hammarstr&ouml;m collected many individuals&#39; bibliographies and compiled them into a master bibliography. Harald also collected extensive information about the proved genealogical relations of the languages of the world. His top-level classification is merged with low-level (i.e. dialect level) information from multitree. Sebastian Nordhoff designed and programmed the database and the first version of the web application with help from Hagen Jung and Robert Forkel. He also took care of the import of the bibliographies. Robert Forkel programmed the second version of the web application as part of the&nbsp;Cross Linguistic Linked Data&nbsp;project. Martin Haspelmath gave advice at every stage throughout the project, helped with coordination, and is currently responsible for languoid names and dialects. Sebastian Bank made numerous improvements to the code base for the Glottolog website and organized the import of several large bibliographies. A table of source bibliographies for Glottolog is available at&nbsp;References information. The Agglomerated Endangerment Status (AES) is derived from the databases of&nbsp;<a href="http://www.endangeredlanguages.com/">&nbsp;</a>The Catalogue of Endangered Languages (ELCat),&nbsp;UNESCO Atlas of the World&#39;s Languages in Danger&nbsp;and&nbsp;<a href="http://www.ethnologue.com/">&nbsp;</a>Ethnologue. For more information see&nbsp;GlottoScope.</strong></p>

opencc-by-4.0Jul 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record