Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

187

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

187 results for “Language Data”

Learn how ShareScore rates datasets ↗
zenodo48/100

Coding data to accompany "A quantitative approach to sociotopography in Austronesian languages"

<p>Dataset consists of csv files with sample languages identified by name and Glottocode. Coding for four sociolinguistic variables, as well as an overall &quot;orientation type.&quot; Each file corresponds to a different method for coding languages employing multiple spatial orientation strategies, as described in the document coding.pdf.</p> <p><strong>Orientation type</strong></p> <ul> <li>land-sea = axis oriented orthogonal to the coast, based on opposition between landward (inland) and seaward (toward the coast), regardless of whether these terms reflect PAN *daya and *lahud&nbsp;</li> <li>land-sea* = land-sea systems in which the land-sea opposition is indistinguishable from &nbsp;geophysical elevation</li> <li>coastal = axis oriented parallel to the coast, often but not necessarily co-lexified with vertical `up&#39; and `down&#39;</li> <li>elevation = axis that &nbsp;distinguishes global or geophysical elevation with respect to deictic center&nbsp;</li> <li>riverine = axis oriented parallel to the river, typically with secondary axis orientated orthogonal to river</li> <li>cardinal = axis fixed according to conventions which do not vary with local geography (although they may be motivated by environmental factors such as wind and the sun)</li> </ul> <p><strong>Distribution</strong></p> <ul> <li>distributed</li> <li>island</li> <li>village</li> </ul> <p><strong>Economy</strong></p> <ul> <li>diversified</li> <li>agriculture</li> <li>subsistence</li> </ul> <p><strong>Geography</strong></p> <ul> <li>diversified</li> <li>inland</li> <li>coast</li> </ul> <p><strong>Terrain</strong></p> <ul> <li>mountainous</li> <li>non-mountainous</li> </ul>

opencc-by-4.0Apr 2021View details →
zenodo48/100

Architectural Languages for the Microservices Architecture: A systematic mapping study [Data set]

<p>This repository contains all artifacts related to the study: Architectural Languages for the Microservices Architecture: A systematic mapping study.</p>

opencc-by-4.0Aug 2022View details →
zenodo48/100

Data on the typology and stability of evidentiality in language contact situations

<p>This material contains the dataset from the&nbsp;<a href="https://version.helsinki.fi/gramadapt/evidentiality/" target="_blank" rel="noopener">gitlab repository</a> of the following MA thesis. Please cite the thesis when using the data.</p> <p>Hyv&ouml;nen, Anu. 2024. <em>Typology and stability of evidentiality in language contact situations</em>. MA thesis, University of Helsinki. Openly available at <a href="https://helda.helsinki.fi/items/2ee41e80-0a04-4af4-8a90-a83598447b0e">https://helda.helsinki.fi/items/2ee41e80-0a04-4af4-8a90-a83598447b0e</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

TDA4ContextualEmbeddings - Public - Debug Data for the codebase of the publication "Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction"

<p>Debug dataset for testing the <a href="https://gitlab.cs.uni-duesseldorf.de/general/dsml/tda4contextualembeddings-public">codebase</a> of the paper <a href="https://doi.org/10.18653/v1/2024.sigdial-1.31">&ldquo;Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction&rdquo;</a> published at the 25th Meeting of the Special Interest Group on Discourse and Dialogue, Kyoto, Japan (SIGDIAL 2024).</p>

openapache2.0Nov 2024View details →
zenodo48/100

Praxis and language brain hemispheric activity data from Kroliczak, Buchwald, et al., 2021 - Cortex - publication

<p>Hemispheric activity fMRI data for praxis and language dataset from the Kroliczak, Buchwald, et al., article published in 2021 in the Cortex publication.</p> <p><strong>Cite as:</strong></p> <p>Kroliczak, G., Buchwald, M., Kleka, P., Klichowski, M., Potok, W., Nowik, A. M., ... &amp; Piper, B. J. (2021). Manual praxis and language-production networks, and their links to handedness. <em>Cortex</em>,&nbsp;<em>140</em>, 110-127.&nbsp;<a href="https://doi.org/10.1016/j.cortex.2021.03.022">https://doi.org/10.1016/j.cortex.2021.03.022</a></p> <p>Link to the publication:&nbsp;<a href="https://www.sciencedirect.com/science/article/pii/S0010945221001337">https://www.sciencedirect.com/science/article/pii/S0010945221001337</a></p> <p>The complete dataset for this publication was published at OSF.io:&nbsp;<a href="https://osf.io/63hjt/">https://osf.io/63hjt/</a></p>

opencc-by-4.0May 2023View details →
zenodo44/100

Study Data: Is It Time to Reconsider our Current Approaches to Natural Language Understanding?

<p>Participants consisted of 95 traditional, undergraduate students enrolled in multiple undergraduate psychology courses offered at a private, Mid-Atlantic liberal arts college.</p>

openmit-licenseFeb 2021View details →
zenodo44/100

CLDF dataset with data and supplements for Barlow "Loss of colexification of 'hand' and 'five' in Austronesian languages"

CLDF dataset with data and supplements for Barlow "Loss of colexification of 'hand' and 'five' in Austronesian languages"

opencc-by-4.0Oct 2024View details →
zenodo44/100

Data associated with the article 'Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis'

<p>The efficacy of oral language comprehension interventions varies, but the reasons for this variation have received little attention. A meta-analysis was conducted to examine intervention factors associated with the efficacy (as expressed with effect sizes) of oral language comprehension interventions in children under the age of 18 with or at risk for (Developmental) Language Disorder, (D)LD.</p> <p>The meta-analysis article together with this additional material comprise the content needed for a thorough understanding and replication of the results.</p> <p>This dataset is based on two systematic scoping reviews on oral language comprehension interventions (Tarvainen et al., 2020, 2021). Further information from the sourced articles was extracted for this study titled &lsquo;Intervention factors associated with efficacy, when targeting oral language comprehension of children with or at risk for (Developmental) Language Disorder: A meta-analysis&rsquo;.&nbsp;</p> <p>In the future, we hope that this data is used with a growing body of oral language comprehension interventions to conduct further and more detailed examinations of intervention factors associated with efficacy.</p> <p>References:</p> <p>Tarvainen, S., Launonen, K., &amp; Stolt, S. (2021). Oral language comprehension interventions in school-age children and adolescents with developmental language disorder: A systematic scoping review. <em>Autism &amp; Developmental Language Impairments</em>, <em>6</em>, 1&ndash;24. https://doi.org/10.1177/23969415211010423</p> <p>Tarvainen, S., Stolt, S., &amp; Launonen, K. (2020). Oral language comprehension interventions in 1&ndash;8-year-old children with language disorders or difficulties: A systematic scoping review. <em>Autism &amp; Developmental Language Impairments</em>, <em>5</em>, 1&ndash;24. https://doi.org/10.1177/2396941520946</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2024View details →
zenodo44/100

Embedding Evaluation Data for South African Languages

<p><strong>WordSim and Simlex Data for South African Languages</strong></p> <ul> <li>Setswana</li> <li>Sepedi</li> </ul> <p><strong>Embedding Evaluation Data for South African Languages</strong></p> <p><strong>Dataset Information\</strong></p> <p>The datasets(Simlex and WordSim) contain pairs of Setswana and Sepedi words that have been assigned similarity ratings by humans to measure semantic relatedness. The word-pairs(Simlex and WordSim) are manually translated from English to Setswana and Sepedi. The evaluation task aims to find the degree of correlation between the scores provided by the model and the human rating, the score of the model is collected by computing the cosine similarity of corresponding vectors for word pairs.</p> <p>Online Repository link</p> <ul> <li><a href="https://zenodo.org/record/5673974">Zenodo Data Repository</a>&nbsp;- Link to the data repository.</li> </ul> <p>Authors</p> <ul> <li><strong>Vukosi Marivate</strong>&nbsp;-&nbsp;<a href="https://twitter.com/vukosi">@vukosi</a></li> <li><strong>Valencia Wagner</strong></li> <li><strong>Mack Makgatho</strong></li> <li><strong>Tshephisho Sefara</strong></li> </ul> <p>See also the list of&nbsp;<a href="https://github.com/dsfsi/embedding-eval-data//contributors">contributors</a>&nbsp;who participated in this project.</p> <p>Citing the dataset</p> <p>To appear in conference proceedings</p> <blockquote> <p>@article{Makgatho_Marivate_Sefara_Wagner_2022, title={Training Cross-Lingual embeddings for Setswana and Sepedi},&nbsp;<br> volume={3},&nbsp;<br> url={https://upjournals.up.ac.za/index.php/dhasa/article/view/3822},&nbsp;<br> DOI={10.55492/dhasa.v3i03.3822},&nbsp;<br> number={03},<br> journal={Journal of the Digital Humanities Association of Southern Africa },<br> author={Makgatho, Mack and Marivate, Vukosi and Sefara, Tshephisho and Wagner, Valencia},&nbsp;<br> year={2022},&nbsp;<br> month={Feb.}}</p> </blockquote>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Metadata and annotation data for the XSample corpus on German academic language

<p>The XSample corpus war created in the project <em>XSample</em> (https://www.izus.uni-stuttgart.de/fokus/fdm-projekte/xsample/) at Universit&auml;t Stuttgart in 2021 by Melanie Andresen and Axel Pichler. It contains 135 German academic journal articles, 45 each from the disciplines linguistics, literary studies and philosophy. The texts themselves cannot be made public for copyright reasons. However, metadata and some annotation data are published here.</p> <p><strong>xsample-metadata.csv</strong><br> This file contains metadata on the texts in the corpus, like journal, title, authors, text length, and the URL to the original paper. It also contains two analytical metrics, &#39;past-ratio&#39; and &#39;temp-expr-ratio&#39;, that are based on the annotations in the other two files. The variable &#39;past-ratio&#39; expresses the proportion of verbs in past tense relative to all finite verbs in the text. The variable &#39;temp-expr-ratio&#39; gives the number of temporal expressions per 1000 token.</p> <p><strong>xsample-heidel.csv</strong><br> This file contains all temporal expressions found and classified by the annotation tool <em>HeidelTime</em> (https://github.com/HeidelTime/heideltime, V. 2.2.1, Str&ouml;tgen &amp; Gertz 2013 ). The variable &#39;position&#39; expresses the position of the first character of the temporal expression in the text in characters.</p> <p><strong>xsample-sticker2.csv</strong><br> This file contains all finite verbs found and classified by the annotation tool <em>sticker2</em> (https://github.com/stickeritis/sticker2). The variable &#39;position&#39; expresses the position of the first character of the finite verb in the text in characters.</p> <p><strong>References</strong><br> Str&ouml;tgen, Jannik &amp; Michael Gertz. 2013. Multilingual and cross-domain temporal tagging. <em>Language Resources and Evaluation</em>. Springer 47(2). 269&ndash;298. <a href="https://doi.org/10.1007/s10579-012-9179-y">https://doi.org/10.1007/s10579-012-9179-y</a>.</p> <p>&nbsp;</p>

opencc-by-4.0May 2022View details →
zenodo44/100

Dikhyang Bugun language data: photographs

<p>These sound files constitute the elicitation and the triple-repetition of the lexical entries of the Basic Word List in the Dikhyang variety of Bugun. Dikhyang village is a recent (1977) satellite settlement of Wanggo (&lsquo;Wangho&rsquo;) village, via Rambung and Hemrai settlements. Wanggo is located on the other side of the ridge.&nbsp; The Bugun variety of Dikhyang should hence closely match the Bugun variety of Wanggo (e.g. &lsquo;Wangho&rsquo; in Abraham et al. 2005).</p>

opencc-by-4.0Dec 2017View details →
zenodo44/100

Dikhyang Bugun language data: sound files

<p>These sound files constitute the elicitation and the triple-repetition of the lexical entries of the Basic Word List in the Dikhyang variety of Bugun. Dikhyang village is a recent (1977) satellite settlement of Wanggo (&lsquo;Wangho&rsquo;) village, via Rambung and Hemrai settlements. Wanggo is located on the other side of the ridge.&nbsp; The Bugun variety of Dikhyang should hence closely match the Bugun variety of Wanggo (e.g. &lsquo;Wangho&rsquo; in Abraham et al. 2005).</p>

opencc-by-4.0Dec 2017View details →
zenodo44/100

Bangru Language Data: sound files (uncut)

<p>These files form the empirical basis for the following article:</p> <p>Bodt, Timotheus Adrianus and Ismael Lieberherr. 2015. First notes on the phonology and classification of the Bangru language of India. <em>Linguistics of the Tibeto-Burman Area 38:1</em> (2015), 66&ndash;123.</p> <p>doi 10.1075/ltba.38.1.03bod</p> <p>issn 0731&ndash;3500 / e-issn 2214&ndash;5907 &copy; John Benjamins Publishing Company</p> <p>These data were collected in Sarli circle, Kurung Kumey district, Arunachal Pradesh, India.</p> <p>The data collectors were the following faculty, students and associated researchers of the Department of English and Foreign Languages, Tezpur University, Assam, India:</p> <p>Nupur Sinha&nbsp;&nbsp; (Faculty), Ismael Lieberherr&nbsp;&nbsp; (Affiliated PhD scholar), Timotheus A. Bodt (Affiliated PhD scholar), Diksha Konwar, Eshani Baishya, Nawaf Helmi, Pinaz Mirza, Ratul Mahela, Sansuma Brahma (students).</p> <p>This material is made freely available to everyone for informative or scientific purposes as long as the source (this DOI) / the collectors are properly credited. Please note that use of the material for&nbsp;commercial purposes&nbsp;<em><strong>of any kind</strong>, which includes conversion into commercial audio-visual media (documentaries etc.), storage and dissemination through sites that require registration &amp; payment for access, or sites that rely on advertisement (including YouTube)&nbsp;</em>is&nbsp;<strong>not</strong>&nbsp;permitted without&nbsp;<strong>specific written consent</strong>&nbsp;from the speakers and their community, obtained through the collectors of the material. By downloading our material, you agree to these restrictions.</p> <p>This data set falls under the Attribution-NonCommercial-ShareAlike (CC BY-NC-SA) license. This license lets you remix, tweak, and build upon this work non-commercially, as long as you credit us and license your new creations under the identical terms. License Deed on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/">https://creativecommons.org/licenses/by-nc-sa/4.0/</a>. Legal Code on&nbsp;<a href="https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode">https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode</a>.</p>

opencc-by-4.0Dec 2017View details →
zenodo44/100

Historical tone data for Tai languages

<p>Comma-separated values (CSV) of historical tone data for 300+ Tai doculects (languages and dialects). Gives historical categories (Gedney 1972) and tone numerals (Chao 1930). Contact author for source citations. I recommend getting in touch if you&#39;d like to use this dataset&nbsp;There is a good chance I have a newer (bigger, cleaner, better) version of it you could use!</p>

opencc-by-4.0Jan 2018View details →
zenodo44/100

Prediction Experiment for Western Kho-Bwa language data: dataset

<p><strong>Prediction Experiment on Western Kho-Bwa languages</strong></p> <p><em>Timotheus A. Bodt (SOAS, London) and Johann-Mattis List (Max Planck Institute, Jena)</em></p> <p>This database includes all the sound files and the transcriptions of the prediction experiment for Western Kho-Bwa. This experiment was registered as:</p> <p>Bodt, Timotheus A., Nathan W. Hill and Johann-Mattis List. 2018. <em>Prediction experiment for missing words in Kho-Bwa language data. </em>Open Science Framework Preregistrations October 5.&nbsp; <a href="https://osf.io/evcbp/">https://osf.io/evcbp/</a> &nbsp;&nbsp;&nbsp;</p> <p>The data and code can be found on:</p> <p>Timotheus A. Bodt, Nathan W. Hill, &amp; Johann-Mattis List. (2018, October 8). Prediction experiment for missing words in Kho-Bwa language data (Version v1.0.1). Zenodo.&nbsp;<a href="http://doi.org/10.5281/zenodo.1451176">http://doi.org/10.5281/zenodo.1451176</a></p> <p>A paper explaining the experiment is under review:</p> <p>Bodt, Timotheus A. and Johann-Mattis List. 2019 (under review). Testing the predictive force of the comparative method: An ongoing experiment on unattested words in Western Kho-Bwa languages. <em>Papers in Historical Phonology</em> Volume 1: 1&ndash;21.</p> <p>The results of the experiment will be presented at the International Conference on Historical Linguistics 24: 01-Jul-2019 - 05-Jul-2019, Canberra, Australia.</p> <p>The uncut sound files, cut sound files, original field notes and preliminary transcriptions have been saved as:</p> <p>Bodt, Timotheus Adrianus. (2019). <em>&#39;Retrodiction&#39; experiment Western Kho-Bwa languages: data [Data set]</em>. Zenodo. <a href="http://doi.org/10.5281/zenodo.2529727">http://doi.org/10.5281/zenodo.2529727</a></p> <p><strong>How to use these files?</strong></p> <ul> <li>Download the zip folder soundfiles_prediction_experiment.zip</li> <li>Extract the files in a separate folder</li> <li>Search for the required sound file(s)</li> </ul> <p>Searching sound files can best be done using the English CONCEPTS from the predictions_results.csv file. For example, searching for BACK will give all the sound files that contain the English gloss &lsquo;back&rsquo; (including &lsquo;backwards&rsquo;, &lsquo;back&rsquo; as body part, turn &lsquo;back&rsquo; etc.).</p> <p>Another option is the select all the sound files of a given linguistic variety / doculect by searching for the original sound file number.</p> <p>I would advise against using a certain attested form in the predictions_results.csv file and search for that (e.g. p a ŋ + b u &lsquo;chest&rsquo;), because the cut sound files have been saved without spaces and morpheme breaks and because the actual transcriptions of the sound files may have changed after analysis, but were not updated in the name of the cut sound files.</p> <p>If you cannot find a certain sound file, then it may simply not have been recorded or not cut from the main sound file. If you are really interested, please mail me at <a href="mailto:timintibet@hotmail.com">timintibet@hotmail.com</a> and I will attempt to find it or record it.</p>

opencc-by-4.0Apr 2019View details →
zenodo44/100

Data for "SeaMoon: from protein language models to continuous structural heterogeneity"

<p>Datasets used for development of SeaMoon:&nbsp;<br><a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</p> <p>This upload contains the following data:</p> <ul> <li><strong>precomputed_emb.tar.gz</strong> is a compressed archive containing the precomputed data used for training and testing the models of the SeaMoon method, in Torch <strong>.pt </strong>format.&nbsp;<br>The file prefixes consist of two IDs, "ID1_ID2_", identifying the <a href="https://github.com/PhyloSofS-Team/DANCE">DANCE</a> [1] protein conformational collection used for its generation. "ID1" represents the first member of the collection in alphabetical order, while "ID2" is the reference conformation for the structural alignment. The "ESM_data" or "ProstT5_data" suffixes designate the type of embeddings, generated by either ESM2 [2] or ProstT5 [3].<br>The dictionnary contains the following keys: <ul> <li><strong>emb:</strong> The per-residue embedding.</li> <li><strong>data: </strong>A tuple containing "ID2" (the reference), the amino acid sequence, and the coverage of the positions in the original DANCE collection.</li> <li><strong>eigvect:</strong> The eigenvectors of the covariance matrix of the "ID1_ID2" collection, centered on reference conformaton "D2".</li> <li><strong>eigval:&nbsp;</strong>The associated eigenvalues.</li> <li><strong>ref:</strong> The coordinates of the C-alpha atoms of the reference conformaton "ID2".</li> </ul> </li> <li><strong>train_list.txt, train_list_5ref.txt, val_list.txt </strong>and<strong> test_list.txt</strong> contain the identifiers of the samples used for training and evaluating the SeaMoon models. In the "5ref" setting, we used up to 5 reference conformations per collection.&nbsp;</li> </ul> <p>For details on SeaMoon see:</p> <div> <div>SeaMoon: Prediction of molecular motions based on language models</div> </div> <div>Valentin Lombard, Dan Timsit, Sergei Grudinin, Elodie Laine</div> <div>bioRxiv 2024.09.23.614585; doi: https://doi.org/10.1101/2024.09.23.614585</div> <div>&nbsp;</div> <div>For more information on data usage and generation please see <a href="https://github.com/PhyloSofS-Team/seamoon">https://github.com/PhyloSofS-Team/seamoon</a>.</div> <div>&nbsp;</div> <div>Abstract:</div> <p>How protein move and deform determines their interactions with the environment and is thus of utmost importance for cellular functioning. Following the revolution in single protein 3D structure prediction, researchers have focused on repurposing or developing deep learning models for sampling alternative protein conformations. In this work, we explored whether continuous compact representations of protein motions could be predicted directly from protein sequences, without exploiting nor sampling protein structures. Our approach, called SeaMoon, leverages protein Language Model (pLM) embeddings as input to a lightweight (~1M trainable parameters) convolutional neural network. SeaMoon achieves a success rate of up to 40% when assessed against ~1,000 collections of experimental conformations exhibiting a wide range of motions. SeaMoon capture motions not accessible to the normal mode analysis, an unsupervised physics-based method relying solely on a protein structure's 3D geometry, and generalises to proteins that do not have any detectable sequence similarity to the training set. SeaMoon is easily retrainable with novel or updated pLMs.&nbsp;</p> <p>&nbsp;</p> <p>[1] Lombard, V.; Grudinin, S.; Laine, E. Explaining Conformational Diversity in Protein Families through Molecular Motions. Scientific Data 2024, 11, 752.</p> <p>[2] Lin, Z.; Akin, H.; Rao, R.; Hie, B.; Zhu, Z.; Lu, W.; Smetanin, N.; Verkuil, R.; Kabeli, O.; Shmueli, Y.; Dos Santos Costa, A.; Fazel-Zarandi, M.; Sercu, T.; Candido, S.; Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 2023, 379, 1123&ndash;1130.</p> <p>[3] Heinzinger, M.; Weissenow, K.; Sanchez, J. G.; Henkel, A.; Steinegger, M.; Rost, B. ProstT5: Bilingual language model for protein sequence and structure. bioRxiv 2023, 2023&ndash;07.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Wordlist files of lexical data from Papua New Guinea and western Solomons Oceanic languages collated for Ross's 1986 PhD thesis and 1988 publication thereof

<p>It occurs to me that the files containing&nbsp;Western Oceanic lexical data&nbsp;that&nbsp;I collected in the late 70s/early 80s for my PhD (Ross 1988) might be useful to someone. They are also used in the volumes of <em>The lexicon of Proto&nbsp;Oceanic </em>(Ross, Pawley &amp; Osmond 1998, 2003, 2011, 2016, 2023). In any case, it is right that they be made publicly available, something that wasn&#39;t so easy back then. Most of the material is from wordlists that I collected during fieldwork in Papua New Guinea from around 1978 to 1982. The file cor06 is omitted because it contains SE Solomonic data (outside Western Oceanic) drawn from Tryon &amp; Hackman 1983.</p> <p>I keyed the data into text files in a format such that each line was the entry for a single word, and each field within an entry was marked by a backslash code (I adapted this format from SIL&#39;s conventions at the time), then arranged them in cognate sets, each set separated from the next by an empty line. This work was done between 1983 and 1985, when text files were the best way to store data. They were entered on a terminal connected to a mainframe computer at the ANU. I have converted the ASCII symbols used in the original files into UTF-8 here in the interests of readability. The conversion was largely automatic, and I have not done a full check of each file, so there may be glitches.</p> <p>Each file contains languages from a region, as listed below (and the regions sometimes cut across subgroups determined by the comparative method). Three-letter abbreviations are used for language names, and two key files are also provided, one (COR-abbrevs) ordered by regions (determined by the numerals that start each line), the other by alphabetical order of&nbsp;language name (COR-abbrevs-alph). Some three-letter codes are followed by a hyphen and an extra letter. These are dialects. For example, MUM stands for Mumeng&nbsp;and MUM-P for the Patep dialect of Mumeng.</p> <p>Data files are labelled with COR (for &#39;correspondence sets&#39;) plus a numeral. The numerals are: 1-3 New Ireland; 4 Willaumez Peninsula (New Britain) area; 5 NW Solomonic; 7+8 Papuan Tip; 9 Vitiaz Strait area and NG north coast; 10 Huon Gulf and Markham Valley; 11 South and west New Britain. 7+8 are partial only. When I keyed the files,&nbsp;I had to rely on a mainframe&#39;s nightly back-up onto tape spools. One night the system failed, and so did the restore, and I lost some data.</p> <p>The backslash codes in the data files are: \l language; \p protolanguage; \w word; \g gloss; \n note; \s source. The formatting of these files is a little odd, since they served as input to routines I wrote to pull out sound correspondences. Anything after &#39;%&#39; is the elicited form: what immediately precedes &#39;%&#39; has had something &#39;undone&#39;, e.g. metathesis.</p> <p>The orthography of the files is phonemic and largely obvious. The conventions are set out in the introductions to the volumes of&nbsp;<em>The lexicon of Proto&nbsp;Oceanic.</em></p> <p>Finally, the files also contain reconstructions at various interstages at the top of a cognate set. These were inserted for heuristic reasons during my research. Many of them did not survive into my PhD thesis, and they should preferably be ignored. The reader who is interested in current Oceanic reconstructions should turn to the volumes of <em>The lexicon of Proto&nbsp;Oceanic.</em></p>

opencc-by-4.0Apr 2023View details →
zenodo40/100

LRRo: A Lip Reading Data Set for the Under-resourced Romanian Language

<p>Two distinct collections are presented in this repository:</p> <p>(i) wild LRRo data is designed for an Internet in-the-wild, ad-hoc scenario, coming with more than 35 different speakers, 1.1k words, a vocabulary of 21 words, and more than 20 hours;</p> <p>(ii) lab LRRo data, addresses a lab controlled scenario for more accurate data, coming with 19 different speakers, 6.4k words, a vocabulary of 48 words, and more than 5 hours.</p>

opencc-by-4.0Apr 2020View details →
zenodo40/100

Learner Data from a Study on Latin Language Learning

<p>The dataset contains test results from a digital intervention study of the&nbsp;<a href="https://www.projekte.hu-berlin.de/en/callidus-en">CALLIDUS Project</a>&nbsp;in a high school in Berlin.&nbsp;13 Students were randomly sampled in two groups and completed various linguistic tasks. The focus of the study was to find out whether learning Latin vocabulary in authentic contexts leads to higher lexical competence, compared to memorizing traditional vocabulary lists.</p> <p>The data is available in&nbsp;<a href="https://www.json.org">JSON format</a> as provided by the&nbsp;<a href="https://h5p.org">H5P implementation</a>&nbsp;of&nbsp;<a href="https://www.valamis.com/hub/xapi">XAPI</a>. File names indicate the time of test completion, in the concatenated form of &quot;year-month-day-hour-minute-second-millisecond&quot;. This allows us to trace the development of single learners who were fast enough to&nbsp;perform&nbsp;the test twice in a row.</p> <p>Changelog:</p> <p>Version 2.0: Each exercise now has a unique ID that is consistent in the whole dataset, so evaluation/visualization can refer to specific exercises more easily.</p> <p>Version 3.0: A simplified Excel Spreadsheet has been added to enhance the&nbsp;reusability&nbsp;of the dataset. It contains a slightly reduced overview of the data, but the core information (user ID, task statement, correct solution, given answer, score, duration) is still present.</p>

opencc-zeroJan 2020View details →
zenodo40/100

Data set of the article: Language Bias in the Google Scholar Ranking Algorithm

<p>Data of investigation published&nbsp;in the article Crist&ograve;fol Rovira; Llu&iacute;s Codina; Carlos Lopezosa&nbsp;Language Bias in the Google Scholar Ranking Algorithm. Future Internet, 2021, 13.</p> <p><strong>Abstract: </strong>The visibility of academic articles or conference papers depends on their being easily found in academic search engines, above all in Google Scholar. To enhance this visibility, search engine optimization (SEO) has been applied in recent years to academic search engines in order to optimize documents and, thereby, ensure they are better ranked in search pages (i.e., academic search engine optimization or ASEO). To achieve this degree of optimization, we first need to further our understanding of Google Scholar&rsquo;s relevance ranking algorithm, so that, based on this knowledge, we can highlight or improve those characteristics that academic documents already present and which are taken into account by the algorithm. This study seeks to advance our knowledge in this line of research by determining whether the language in which a document is published is a positioning factor in the Google Scholar relevance ranking algorithm. Here, we employ a reverse engineering research methodology based on a statistical analysis that uses Spearman&rsquo;s correlation coefficient. The results obtained point to a bias in multilingual searches conducted in Google Scholar with documents published in languages other than in English being systematically relegated to positions that make them virtually invisible. This finding has important repercussions, both for conducting searches and for optimizing positioning in Google Scholar, being especially critical for articles on subjects that are expressed in the same way in English and other languages, the case, for example, of trademarks, chemical compounds, industrial products, acronyms, drugs, diseases, etc.</p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record