Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,721

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,721 results for “Language”

Learn how ShareScore rates datasets ↗
zenodo48/100

Data on the typology and stability of evidentiality in language contact situations

<p>This material contains the dataset from the&nbsp;<a href="https://version.helsinki.fi/gramadapt/evidentiality/" target="_blank" rel="noopener">gitlab repository</a> of the following MA thesis. Please cite the thesis when using the data.</p> <p>Hyv&ouml;nen, Anu. 2024. <em>Typology and stability of evidentiality in language contact situations</em>. MA thesis, University of Helsinki. Openly available at <a href="https://helda.helsinki.fi/items/2ee41e80-0a04-4af4-8a90-a83598447b0e">https://helda.helsinki.fi/items/2ee41e80-0a04-4af4-8a90-a83598447b0e</a>.</p>

opencc-by-4.0Apr 2024View details →
zenodo48/100

Dataset for "Large Language Models as molecular design engines"

<ol> <li><strong>claude-gpt-paper.zip :</strong><br><br>This dataset contains data and results associated with the paper "Large Language Models as molecular design<br>engines" The paper investigates the use of large language models, specifically Claude 3 Opus, for generating and analyzing chemical structures based on various prompts from A-H (as mentioned in the manuscript), and guided design related to electron-withdrawing groups (EWG), electron-donating groups (EDG).</li> </ol> <p>The dataset includes:</p> <ol> <li>PM7 MOPAC energy calculations for generated molecules, along with their SMILES representations and molecule IDs.</li> <li>PM7-calculated charges for the generated molecules.</li> <li>Output files from the Claude 3 Opus language model for each prompt category along.</li> <li>Original dataset (subset of ZINC database) used to build common keys and the initial design space.</li> <li>JSON file containing common keys for featurizing unknown SMILES.</li> <li>PCA object to convert molecule embeddings to 3-dimensional embeddings.</li> </ol> <p>The data is organized into the following folders:</p> <ul> <li><code>pm7_charge_results</code>: Contains HOMO-LUMO energy differences for plotting.</li> <li><code>pm7_charge_calculation</code>: Contains PM7 MOPAC energy calculations and charges.</li> <li><code>out</code>: Contains output files from the Claude 3 Opus language model.</li> <li><code>fact-dropbox</code>: Contains the original dataset, common keys, and PCA object file.</li> </ul> <p>The data can be used to reproduce the results presented in the paper and serve as a foundation for further research in this area.</p> <p>For a detailed description of the folder structure and contents, please refer to the File_descriptions.md file included in the dataset.<br><br><br>2. llm-visulizer-dashapp.zip<br><br>This is the code for the visualizer app for viewing the molecules generated by the LLM. The README.md file has details about running the app.</p> <p>3. claude-gpt-paper-codes.zip&nbsp;</p> <p>This contains the notebook GPT_modification_just_plots.ipynb for plotting, and other codes. The README.md file has details about running the main notebook for getting the plots.</p>

opencc-by-4.0May 2024View details →
zenodo48/100

TDA4ContextualEmbeddings - Public - Debug Data for the codebase of the publication "Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction"

<p>Debug dataset for testing the <a href="https://gitlab.cs.uni-duesseldorf.de/general/dsml/tda4contextualembeddings-public">codebase</a> of the paper <a href="https://doi.org/10.18653/v1/2024.sigdial-1.31">&ldquo;Local Topology Measures of Contextual Language Model Latent Spaces With Applications to Dialogue Term Extraction&rdquo;</a> published at the 25th Meeting of the Special Interest Group on Discourse and Dialogue, Kyoto, Japan (SIGDIAL 2024).</p>

openapache2.0Nov 2024View details →
zenodo48/100

Sentiment polarity lexicon of Bosnian language

<p>First sentiment annotated lexicon of the Bosnian language.</p> <p>The lexicon is divided into two files: positive and negative polarity.</p> <p>The lists are prepared in separate files, each file populated by words of the aforementioned polarity listed each word in a separate row.&nbsp;</p> <p>The positive polarity list&nbsp;BOSNIAN_POSITIVE.txt holds 1219 words.</p> <p>The negative polarity list&nbsp;BOSNIAN_NEGATIVE.txt holds 3935 words.</p>

opencc-by-4.0Jan 2023View details →
zenodo48/100

CLDF dataset accompanying Miller and List's "Borrowing in South American Languages" from 2023

<p>Cite the source of the dataset as:</p> <blockquote> <p>Miller, John and List, Johann-Mattis (2023): Detecting Lexical Borrowings from Dominant Languages in Multilingual Wordlists. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics. Short Papers.</p> </blockquote>

opencc-by-4.0Jan 2023View details →
zenodo48/100

Cross-language corpora of privacy policies

<p>The dataset consists of three different privacy policy corpora (in English and Italian) composed of 81 unique privacy policy texts spanning the period 2018-2021. This dataset makes available an example of three corpora of privacy policies. The first corpus is the English-language corpus, the original used in the study by Tang et al. [2]. The other two are cross-language corpora built (one, the source corpus, in English, and the other, the replication corpus, in Italian, which is the language of a potential replication study) from the first corpus.</p> <p>The policies were collected from:</p> <ol> <li>the Alexa top 10 Italy and U.S. websites rank;</li> <li>the Play Store apps rank in the &quot;most profitable games&quot; category of the Play Store for Italy and the U.S.</li> </ol> <p>We manually analyzed the Alexa top 10 Italy websites as of November 2021. Analogously, we analyzed selected apps that, in the same period, had ranked better in the &quot;most profitable games&quot; category of the Play Store for Italy.</p> <p>All the privacy policies are ANSI-encoded text files and have been manually read and verified.<br> The dataset is helpful as a starting point for building comparable cross-language privacy policies corpora. The availability of these comparable cross-language privacy policies corpora helps replicate studies in different languages.&nbsp;<br> Details on the methodology can be found in the accompanying paper.</p> <p>The available files are as follows:</p> <ul> <li><strong>policies-texts.zip</strong> --&gt;&nbsp;contains a directory of text files with the policy texts. File names are the SHA1 hashes of the policy text.</li> <li><strong>policy-metadata.csv</strong> --&gt;&nbsp;Contains a CSV file&nbsp;with the metadata&nbsp;for each privacy policy.</li> </ul> <p>This dataset is the original dataset used in the publication [1]. The original English U.S. corpus is described in the publication [2].</p> <p>[1] F. Ciclosi, S. Vidor and F. Massacci. &quot;Building cross-language corpora for human&nbsp;understanding of privacy policies.&quot; Workshop on Digital Sovereignty in Cyber Security: New Challenges in Future Vision. Communications in Computer and Information Science. Springer International Publishing, 2023, In press.</p> <p>[2] J. Tang, H. Shoemaker, A. Lerner, and E. Birrell. Defining Privacy: How Users&nbsp;Interpret Technical Terms in Privacy Policies. Proceedings on Privacy Enhancing&nbsp;Technologies, 3:70&ndash;94, 2021.</p>

opencc-by-4.0Mar 2023View details →
zenodo48/100

Praxis and language brain hemispheric activity data from Kroliczak, Buchwald, et al., 2021 - Cortex - publication

<p>Hemispheric activity fMRI data for praxis and language dataset from the Kroliczak, Buchwald, et al., article published in 2021 in the Cortex publication.</p> <p><strong>Cite as:</strong></p> <p>Kroliczak, G., Buchwald, M., Kleka, P., Klichowski, M., Potok, W., Nowik, A. M., ... &amp; Piper, B. J. (2021). Manual praxis and language-production networks, and their links to handedness. <em>Cortex</em>,&nbsp;<em>140</em>, 110-127.&nbsp;<a href="https://doi.org/10.1016/j.cortex.2021.03.022">https://doi.org/10.1016/j.cortex.2021.03.022</a></p> <p>Link to the publication:&nbsp;<a href="https://www.sciencedirect.com/science/article/pii/S0010945221001337">https://www.sciencedirect.com/science/article/pii/S0010945221001337</a></p> <p>The complete dataset for this publication was published at OSF.io:&nbsp;<a href="https://osf.io/63hjt/">https://osf.io/63hjt/</a></p>

opencc-by-4.0May 2023View details →
zenodo48/100

Dataset for : A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification

<p>We present&nbsp;a novel solution combining Large Language Model (LLM) capabilities with Formal Verification strategies to falsify and automatically repair software vulnerabilities. Initially, we employ Bounded Model Checking (BMC) to locate the software vulnerability and derive a counterexample. Relying on mathematical proofs, counterexamples provide evidence that the system behaves incorrectly or contains a vulnerability, thereby preventing the generation of false positive alerts. The counterexample that has been detected, along with the source code, are provided to the LLM engine. Our approach involves establishing a specialized prompt language for conducting code debugging and generation to understand the vulnerability&#39;s root cause and repair the code. Finally, we use BMC to verify the corrected version of the code generated by the LLM. As a proof of concept, we create \esbmcai based on the Efficient SMT-based Context-Bounded Model Checker (ESBMC) and a pre-trained Transformer model, specifically gpt-3.5-turbo, to detect and fix errors in C programs. We generated a dataset comprising $1{,}000$ C code samples, each consisting of $20$ to $50$ lines of C code. Experimental results show that our proposed method achieved an impressive success rate of up to $80$\% in repairing vulnerable code, encompassing buffer overflow, arithmetic overflow, and pointer dereference failures. To our knowledge, \esbmcai represents the first proposal for a pioneering initiative to integrate a Large Language Model (LLM) with software model checking. We advocate that this automated approach has the potential to incorporate into the software development lifecycle&#39;s continuous integration and deployment (CI/CD) process.&nbsp;</p> <p>&nbsp;</p> <p>The uploaded&nbsp;dataset contains 1000 codes,&nbsp; each comprising 20&nbsp;to 50&nbsp;lines of C code generated with gpt-3.5-turbo. The material also consists of a version of ESBMC statically compiled with all dependencies, a classifier script, and the output file.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2023View details →
OpenNeuro44/100

Language Production fMRI

Open the record for dataset details and reuse information.

openCC0Jan 2019View details →
zenodo44/100

cldf-datasets/apics: The "Atlas of Pidgin and Creole Language Structures Online" as CLDF dataset

<p>Cite as</p> <blockquote> <p>Michaelis, Susanne Maria &amp; Maurer, Philippe &amp; Haspelmath, Martin &amp; Huber, Magnus (eds.) 2013. Atlas of Pidgin and Creole Language Structures Online. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at <a href="https://apics-online.info">https://apics-online.info</a>)</p> </blockquote>

opencc-by-4.0Nov 2013View details →
zenodo44/100

The Dublin Language Garden Perceptual Dialectology of Irish English Collection

<blockquote> <p><strong>Recommended citation for this dataset:</strong><br> Garnett, Vicky, &amp; Lucek, Stephen. (2020). The Dublin Language Garden Perceptual Dialectology of Irish English Collection (Version 1.0.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.4247829</p> </blockquote> <p>&nbsp;</p> <p><strong>About this Dataset</strong><br> The field of Perceptual Dialectology is&nbsp;an area of sociolinguistic study that investigates how non-linguists view different varieties of language.&nbsp; It often includes hand-drawn map exercises in which participants indicate where they believe various varieties are spoken, and their attitudes towards them.&nbsp;</p> <p>In 2015, as part of a public linguistics outreach event (the Dublin Language Garden) held at Trinity College Dublin, the authors created an activity for members of the public and collected hand-drawn maps from them that gave responses to the following tasks:</p> <p>a. Indicate where you come from on the map (using a red dot sticker)<br> b. Draw where you think the Dublin dialect occurs<br> c. Draw the boundaries of any other dialects you believe occur in Ireland<br> d. Tell us what you think are the features of those dialects<br> e. Tell us what you think are the characteristics of the people who speak those dialects.</p> <p>Participants of all ages were encouraged to take part, but only data from those over 18 were retained after the event and used in this data collection. &nbsp;Participants were all given information on how the data was to be anonymised, processed and published on a clearly displayed poster to read before they were given a map to complete the 5 tasks (listed above). &nbsp;No additional information about the participants, aside from that acquired through Task a, was collected.</p> <p>&nbsp;</p> <p><strong>File List:</strong></p> <ul> <li>_READ_ME - Dublin Language Garden Perceptual Dialectology of Irish English data.txt<br> Contains a detailed description of this dataset.<br> &nbsp;</li> <li>DLG_PDIE_KML_data_by_location.zip<br> This zipped folder contains the .kml data of multiple hand-drawn maps organised into folders by their location<br> &nbsp;</li> <li>DLG_PDIE_KML_data_by_part.zip<br> This zipped folder contains the .kml data of multiple hand-drawn maps organised into folders according to the participants.</li> </ul> <p>These folders have been organised in this way in order to make discoverability easier between the data. &nbsp;Users may wish to analyse the data only by the locations of the varieties identified by the participants. &nbsp;Other users may only be interested in the data given by specific participants, and therefore the folder that organises the data in this way may be of better use to them. &nbsp;Both folders, however, contain the same data, it is simply how they are organised.</p> <ul> <li>Garnett and Lucek DLG_PD_IE Qualitative Data (Nov 2020).xlsx<br> Spreadsheet featuring tabulated qualitative data taken from all maps<br> &nbsp;</li> <li>Sample Hand-drawn Maps.zip<br> Folder containing 2 sample hand-drawn maps from the participants to help contextualise the data presented here.</li> </ul> <p>&nbsp;</p> <p><strong>Any questions?</strong><br> Any enquiries regarding this dataset should be directed to either Vicky Garnett (<a href="mailto:garnetv@tcd.ie">garnetv@tcd.ie</a>) or Stephen Lucek (<a href="mailto:stephen.lucek@ucd.ie">stephen.lucek@ucd.ie</a>).</p>

opencc-by-4.0Nov 2020View details →
zenodo44/100

A studyforrest extension, an annotation of spoken language in the German dubbed movie ``Forrest Gump'' and its audio-description (validation analysis)

<p>This component contains the data of the analysis that we ran as a validation of the annotation of speech spoken in the research cut (Hanke et al., 2016) of the movie &quot;Forrest Gump&quot; (Zemeckis, 1994) and its audio-description. The corresponding paper is hosted on github (https://github.com/psychoinformatics-de/studyforrest-paper-speechannotation)&nbsp;and published in f1000research (https://doi.org/10.12688/f1000research.27621.1).</p>

opencc-by-4.0Dec 2020View details →
zenodo44/100

Modified Swadesh-100 list of 23 Cariban languages

<p>Cite the source of the dataset as:</p> <blockquote> <p>Matter, Florian. 2020. Comparative Cariban Database. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at https://cariban.clld.org/)</p> </blockquote>

opencc-by-4.0Jan 2021View details →
zenodo44/100

Linguistic praxeological organization of French as the schooling language

<p>Elements of the&nbsp;linguistic praxeological organization of French as the schooling language used to produce&nbsp;the framework (https://zenodo.org/deposit/4462850) used for the PEAPL project (https://blog.hepfr.ch/create/peapl/) and the framework used for the COMPER project (https://comper.fr/accueil).</p> <p>The following&nbsp;article will give explanations about the elaboration of the&nbsp;linguistic praxeological organizations:&nbsp;https://www.researchgate.net/publication/333244876_Francais_langue_de_scolarisation_Reflexions_sur_les_referentiels_de_competences_et_l&#39;adaptive_learning</p>

opencc-by-4.0Aug 2020View details →
zenodo44/100

Study Data: Is It Time to Reconsider our Current Approaches to Natural Language Understanding?

<p>Participants consisted of 95 traditional, undergraduate students enrolled in multiple undergraduate psychology courses offered at a private, Mid-Atlantic liberal arts college.</p>

openmit-licenseFeb 2021View details →
zenodo44/100

Personal pronoun systems in the languages of the Greater Burma Zone

<p>A collection of 51 languages of Myanmar and surrounding areas (the <em>"Greater Burma Zone"</em>), listing the systems of personal pronouns with some additional information and metadata about the language, including the source. This dataset was used in Müller &amp; Weymuth (2017) to show the correlation between the societal structure (hierarchical vs. non-hierarchical) and the types of personal pronoun systems (hierarchical vs. grammatical).</p> <p>The *.zip file contains 53 files, most *.docx, and some *.odt, and a template file. It also contains an R script that we used to produce the map shown in the paper. This R file is rather crude and many things were entered manually instead of reading it automatically from the dataset.</p> <p><strong>Source:</strong><br> Müller, André &amp; Rachel Weymuth. 2017. "How Society Shapes Language: Personal Pronouns in the Greater Burma Zone." In: <em>Asiatische Studien – Études Asiatiques </em>71(1), 409–432. DOI: 10.1515/asia-2016-0021 (URL: https://www.degruyter.com/downloadpdf/j/asia.2017.71.issue-1/asia-2016-0021/asia-2016-0021.pdf)</p>

opencc-by-4.0May 2017View details →
zenodo44/100

CLDF dataset with data and supplements for Barlow "Loss of colexification of 'hand' and 'five' in Austronesian languages"

CLDF dataset with data and supplements for Barlow "Loss of colexification of 'hand' and 'five' in Austronesian languages"

opencc-by-4.0Oct 2024View details →
zenodo44/100

CLDF dataset derived from Chan's "Numeral Systems of the World's Languages." from 2019

<p>Cite the source of the dataset as:</p> <blockquote> <p>Chan, Eugene (eds.) 2019. Numeral Systems of the World&#x27;s Languages. Accessed: 2019-09-30</p> </blockquote>

opencc-by-4.0Oct 2019View details →
zenodo44/100

A German Language Labeled Dataset of Tweets

<p>Our dataset contains 8,048 German language tweets related to Jewish life from a four-year timespan.&nbsp;</p><p>The dataset consists of 18 samples of tweets with the keyword "Juden" or "Israel." The samples are representative samples of all live tweets (at the time of sampling) with these keywords respectively over the indicated time period. Each sample was annotated by two expert annotators using an Annotation Portal that visualizes the live tweets in context. We provide the annotation results based on the agreement of two annotators, after discussing discrepancies (Jikeli et al. 2022: 3-6).&nbsp;</p><p>&nbsp;Overall, 335 tweets (4%) were labelled as antisemitic following the IHRA Working Definition of Antisemitism. 1345 tweets (17 %) come from 2019, 1364 tweets (17 %) from 2020, 2639 tweets (33 %) from 2021 and 2700 tweets (34 %) from 2022.&nbsp;</p><p>About half of the tweets, a total of 4,493 tweets (56 %) come from queries with the keyword "Juden," which is representative of a continuous time period from January 2019 to December 2022: 864 tweets (19 %) come from 2019, 891 tweets (20 %) from 2020, 1364 tweets (30 %) from 2021 and 1374 (31 %). 148 out of the 4493 tweets, so 3% from the query with "Juden" are antisemitic.&nbsp;</p><p>The other part of the tweets, a total of 3,555 (44 %)&nbsp; results of queries with the keyword "Israel". 481 tweets (14 %) of the keywords containing Israel stem from 2019, 473 (13 %) come from 2020, 1275 tweets (36 %) from 2021 and 1326 tweets (37 %) are from 2022. Out of all tweets from the "Israel" query, 187 (5 %)&nbsp; are antisemitic.&nbsp;</p><p>The csv file contains diacritics and special characters of the German language (e.g., "ä", "ü", "ö", "ß"), which should be taken into account when opening it with anything other than a text editor.&nbsp;</p><p><strong>Acknowledgements</strong>&nbsp;</p><p>This work used Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services &amp; Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.&nbsp;&nbsp;</p><p>We are grateful for the support of Indiana University's Observatory on Social Media (OSoMe) (Davis et al. 2016) and the contributions and annotations of all team members in our Social Media &amp; Hate Research Lab at Indiana University's Institute for the Study of Contemporary Antisemitism, especially Grace Bland, Elisha S. Breton, Kathryn Cooper, Robin Forstenhäusler, Sophie von Máriássy, Mabel Poindexter, Jenna Solomon, Clara Schilling, Emma Shriberg and Victor Tschiskale.&nbsp;</p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

Phlorest phylogeny derived from Sagart et al. 2019 'Dated language phylogenies shed light on the ancestry of Sino-Tibetan'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Sagart L, Jacques G, Lai Y, Ryder RJ, Thouzeau V, Greenhill SJ, List J- M. 2019 Dated language phylogenies shed light on the ancestry of Sino-Tibetan. Proceedings of the National Academy of Sciences, 201817972.</p> </blockquote>

opencc-by-4.0Aug 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record